philipkiely
- Karma
- 0
- Created
- ()
- Submissions
- 0
- The efficient frontier of LLM inference (baseten.co)
- We built a day-0 API for Kimi K3 (baseten.co)
- We built the new fastest API for GLM-5.2 (baseten.co)
- We built the fastest API for GLM-5.2 (280 TPS) (baseten.co)
- The Math Behind TurboQuant (baseten.co)
- Show HN: Inference Engineering (baseten.com)
There is a ton of demand for inference, but there are relatively few engineers working in the space. This leaves novel, interesting, and deeply technical challenges left to solve at every level of the stack. To make it…
- Baseten raises $150M Series D at $2.15B (fortune.com)
- Three techniques to adapt LLMs for any use case (baseten.co)
- Serving four million Riffusion requests in two days (baseten.co)
- Show HN: Free Stable Diffusion 2.0 hosted interface (app.baseten.co)
- Try it yourself: Speech to text with Whisper (app.baseten.co)
- Deploying Stable Diffusion in Production Using Truss (baseten.co)