OsamaJaber
- Karma
- 0
- Created
- ()
- Submissions
- 0
- Agent loops stop re-reading your codebase on every turn (runinfra.ai)
- GLM 5.3 Flash faster and cheaper (runinfra.ai)
- The fastest and cheapest GLM 5.3 Flash endpoint (runinfra.ai)
- Fastest Inference in MENA (runinfra.ai)
- Kimi K3 2.78T on One CPU with 8GB RAM (github.com)
- Mixture-of-Kittens: An MoE training megakernel for NVL72 (twitter.com)
- Lossless Inference (runinfra.ai)
- DeepSeek V4 Flash 2.98x faster, lossless (runinfra.ai)
- DeepSeek-V4-Flash 2.98x faster on 4x B200, lossless (twitter.com)
- What LLM Inference Costs (twitter.com)
- Kimi k3 run on RTX 5090 (github.com)
- Kimi k3 now runs on one consumer GPU (twitter.com)
- The AGI Compiler "Auto" (github.com)
- Compiler for LLMs, world models, and AGI (arxiv.org)
- Auto: The AGI Compiler (github.com)
- Build You Own Model (runinfra.ai)
- Mega Kernels, Written by Agents (arxiv.org)