This small Klingon speaking language model was trained completely on ESP32. The training took 2 days
Number of parameters: 319K
(Disclaimer) The models is tiny and is not a pocket chatbot. It mistakes and is not capable to support conversation, but that's not the goal of the project
The goals of the project is to bring training to edge devices and it worked out
How can one use it? By using solar panels such device could be turned into autonomous meteorological station
Very cool project! Sounds like it was a fun challenge :)
I wonder when we'll start seeing clusters of ESP32-S3s... not sure how interconnects would go though, but I guess the interconnect wouldn't be the bottleneck anyway.
Would probably be better to demonstrate by exsmple how this approach is used to train on sensor data and then use it (as is hinted by the author) instead of acknowledging that the klingon poc is useless.
I love the cool tech demo! I also would have loved to see a offline sensor calibration package - or whatnot - and analysis of model precision instead of Klingon. Just trying to think up of an actual use case where you would actually truly need an LLM instead of one of the other well known data analytics methods.
12 comments
[ 3.9 ms ] story [ 37.4 ms ] threadNumber of parameters: 319K
(Disclaimer) The models is tiny and is not a pocket chatbot. It mistakes and is not capable to support conversation, but that's not the goal of the project
The goals of the project is to bring training to edge devices and it worked out
How can one use it? By using solar panels such device could be turned into autonomous meteorological station
I wonder when we'll start seeing clusters of ESP32-S3s... not sure how interconnects would go though, but I guess the interconnect wouldn't be the bottleneck anyway.
> Backpropagation (gradients derived by hand)
What does by hand mean in this context?
Also how did you write the readme? It's a curious blend of human and AI writing.
I think it is a small thing.