gpjt
- Karma
- 0
- Created
- ()
- Submissions
- 0
- Extending Raschka's GPT-2: an MoE trained from scratch on an RTX 3090 (gilesthomas.com)
- Putting my Jax-trained models on the Hugging Face Hub (gilesthomas.com)
- Why do OpenAI's GPT-2 weights beat mine? Part four: digging into dropout (gilesthomas.com)
- Adding diagrams to my static site generator with D2 (gilesthomas.com)
- Use the built-in GELU, don't roll your own (gilesthomas.com)
- A Quick(ish) Chinchilla Check (gilesthomas.com)
- I use AI on this blog (gilesthomas.com)
- Why do OpenAI's GPT-2 weights beat mine? Part two: the bugfix (gilesthomas.com)
- Why do OpenAI's GPT-2 weights beat mine? (gilesthomas.com)
- Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090 (gilesthomas.com)
- Building intuition about LLM parameter counts (gilesthomas.com)
- Poppy the training box, part 1: the beginnings (gilesthomas.com)
- From bigrams to GPT-2, one component at a time (in Jax) (gilesthomas.com)
- Building a Jax training loop for an LLM training run (gilesthomas.com)
- Thoughts on Role Confusion (gilesthomas.com)
- Flax debugging: making a hash of things (gilesthomas.com)
- 10Gb/s Ethernet: switching to a Broadcom SFP+ module (gilesthomas.com)
- Jax: Commitment Issues (gilesthomas.com)
- Jax Back Ends and Devices (gilesthomas.com)
- Using Safetensors with Flax (gilesthomas.com)
- First Looking into Jax (gilesthomas.com)
- 10Gb/s Ethernet: using mini-heatsinks with a 10GBASE-T SFP+ module (gilesthomas.com)
- 10Gb/s Ethernet: what I did to get it working in my home (gilesthomas.com)