somnial
- Karma
- 0
- Created
- ()
- Submissions
- 0
- What happens when a GPU writes memory (blog.doubleword.ai)
- What happens when a GPU reads memory? (blog.doubleword.ai)
- The case for disaggregated LLM serving (blog.doubleword.ai)
- On-the-fly snapshot compression for elastic inference at scale (blog.doubleword.ai)
- NVLink, NVSwitch, and All That (blog.doubleword.ai)
- The Anatomy of an Instruction Pipeline Hazard (hiraditya.github.io)
- Width vs. Depth: Speculating on the Margin (blog.doubleword.ai)
- 70x faster cold(ish) starts for SGLang (fergusfinn.com)
- LLM powered data structures: A lock-free binary search tree (fergusfinn.com)
- Parallel Primitives for Multi-Agent Workflows (fergusfinn.com)
- Scheduling in LLM Inference (fergusfinn.com)
- How fast can an LLM go? (fergusfinn.com)