19 comments

[ 2.3 ms ] story [ 27.2 ms ] thread
extremely cool. I’ve had Claude walk model architectures to debug Loras and fine tunes, but this is delightful
I'd love to hear more about that. Are you able to fine tune specific layers?

Specifically, one particular otherwise excellent model I use has an alignment problem (sycophancy) that I've isolated to a specific layer. I can nuke the layer with lora and the behaviour stops - but I'm not sure what else I'm nuking in the process. I'm quite new at this so I'd love any advice. Thank you!

Thanks for good vibes. This is awesome
Love it, I thought it only shows the "guts" of a model, turns out it also estimates the cost of serving that model.
How are you generating the architecture graph from model configs, and do you plan to surface tensor shapes or layer-level parameter counts as well?
Visually inspecting how a particular model is architected is quite interesting. Thank you for coming up with this!
[dead]
(comment deleted)
when you mean interactive and animated i had high hopes