smaddrellmander
- Karma
- 0
- Created
- ()
- Submissions
- 0
- Fast weights and sparse attention in GLM-5.3-Flash (idlemachines.co.uk)
- Tokenisation: Who decides what a token is anyway? (idlemachines.co.uk)
- Adam and AdamW: adaptive optimisation and weight decay (idlemachines.co.uk)
- Thinking Machines dropped RoPE, and it's a good idea (idlemachines.co.uk)
- Nesso-1: Accelerating Open-Source Binding Affinity Predictions [pdf] (valencelabs.com)
- The annotated PyTorch training loop (idlemachines.co.uk)
- Mae vs. MSE: more than just the mean vs. median debate (idlemachines.co.uk)
- DiffusionGemma: Discrete diffusion in a large language model (idlemachines.co.uk)
- Heaven knows I'm perplexed now (idlemachines.co.uk)
- Reading MAI's efficiency gain. How to pick architectures like serious people (idlemachines.co.uk)
- MAI-Thinking-1: Building a Hill-Climbing Machine [pdf] (microsoft.ai)
- Are contrastive losses just cross entropy all along? (idlemachines.co.uk)
- Every token, everywhere, all at once (idlemachines.co.uk)
- The cut in the Mixture of Experts compute graph (idlemachines.co.uk)
- DeepSeek V4 from the Inside (idlemachines.co.uk)
- Softmax, can you derive the Jacobian? And should you care? (idlemachines.co.uk)
- Gemma 4 is not your standard transformer (idlemachines.co.uk)