zagwdt
- Karma
- 0
- Created
- ()
- Submissions
- 0
- A/B testing LLMs in production (together.ai)
- Inference Optimization for MiniMax Sparse Attention (together.ai)
- DeepSeek V4 in vLLM: Efficient Long-Context Attention (vllm-website-pdzeaspbm-inferact-inc.vercel.app)
- Introspective Diffusion Language Models (introspective-diffusion.github.io)
- EinsteinArena: Harnessing the collective intelligence of agents in the wild (einsteinarena.com)
- RL Meets Adaptive Speculative Training (together.ai)
- Weak models excel at long context tasks (together.ai)
- TorchSpec: Speculative Decoding Training at Scale (pytorch.org)
- Flash Attention 4 (together.ai)