I think the idea is interesting though, although I wonder if training time for LoRA is such a bottleneck to deserve its own, extremely narrowly scoped, leaderboard. Maybe if it was more tasks or more models we could hope that it transfers? With a single task, and a single model, I’d be afraid of this overfitting pretty heavily.
For NanoGPT, I think the idea always was that the ideas can be transferred to much larger models, or serve as stepping stones for investigations on larger models.
I think there is real value in going smaller/limiting resources. The trend is 'just make the weights bigger and throw more data at it'. It is a MBA's view of winning. We have a knob, keep turning it. It does work but it may not drive as much creativity as resource limits can drive. It is like urban growth boundaries in city planning. If you aren't allowed to 'just expand' you are forced to build more intelligently inside the city and those creative solutions often lead to major improvements.
The AI slop "why" section in the readme is completely useless. I struggle with this too. I start every project doc with a "why" then "overview" section, and AI is still horrendously bad at the "why", filled with vapid corpo/tech startup speak. I usually write the "why" myself, or at least put down the outline to be completed.
Inspired by parameter golf and speedrun approaches I make the case for picking loss functions like a wallclock for LoRA on AI safety targets. The result when I tried it was a functional distillation of an Sparse AutoEncoder into a 5.3MB probe. I have a technical writeup below about it if anyone is interested.
Reading the title, for a moment I thought this would be something about a LoRA serial data bridge over RF and an actual physical wall clock, possibly some novel new homebuilt piece of hardware to display perfectly accurate NTP synchronized time or something.
15 comments
[ 3.2 ms ] story [ 6.1 ms ] threadI think the idea is interesting though, although I wonder if training time for LoRA is such a bottleneck to deserve its own, extremely narrowly scoped, leaderboard. Maybe if it was more tasks or more models we could hope that it transfers? With a single task, and a single model, I’d be afraid of this overfitting pretty heavily.
For NanoGPT, I think the idea always was that the ideas can be transferred to much larger models, or serve as stepping stones for investigations on larger models.
What is LoRA in this context? the communication protocol? Or another term appropriated by LLMs? Why the speedrun?
The AI slop "why" section in the readme is completely useless. I struggle with this too. I start every project doc with a "why" then "overview" section, and AI is still horrendously bad at the "why", filled with vapid corpo/tech startup speak. I usually write the "why" myself, or at least put down the outline to be completed.
https://www.lesswrong.com/posts/PagGF8roBJmjLunsX/competitiv...