18 comments

[ 3.1 ms ] story [ 89.5 ms ] thread
Not gonna work, will be starved by inter node bandwidth.
Hm, looks like this is an old project (2022) associated with HuggingFace (https://huggingface.co/bigscience).

Doesn't look like it's very active nowadays: https://github.com/bigscience-workshop/petals

Article with more details: https://techcrunch.com/2022/12/20/petals-is-creating-a-free-...

> [...] volunteers can donate their hardware power to tackle a portion of a text-generating workload and team up others to complete larger tasks, similar to Folding@home and other distributed compute setups.

It's an interesting concept but the timing is probably too early. If more people had reliable low latency gigabit or ideally 10gbit throughput, then it might start to approach feasibility. There's other blockers too, but that springs to mind immediately.

It's cool to imagine a planet wide neural network interconnected with fiber - the nervous system of a planetary intelligence. But perhaps mushrooms do that already (Alpha Centauri ever relevant).

What? LLMs are perhaps the least bandwidth-constrained thing you can do over the network.

If you run a local model at a good pace, you're emitting tokens at, what, 60 tokens/sec? That's, very roughly, 300 bytes/sec. You can get that speed with dialup from the 80's! And when it takes a half second just to ingest the input tokens (to get to the part where you start generating the reply), you're not gonna care about even a crazy 250ms latency.

Unless you think that the model should be downloaded on-the-fly in response to a user request, there's zero reason for an LLM to need any sort of fast connection.

I have been thinking of something on these lines but with much smaller models. The entire model has to fit on a single computer. Host owner would choose the model they prefer, perhaps because they already use it. Then it is more about utilizing the GPU for LLM requests.

Peer to peer, consumers get to route their request to a host with compatible model. Consumers have to also contribute GPU but it does not have to be equal - I have not thought through the fairness part. Perhaps initially it starts with "create your friends group and have access to all the host nodes".

Petals is from 2022. Nowadays intelligence of smaller models, quantization techs, and optimizations to run models faster on consumer GPUs have improved a lot.

For distributed inference of smaller LLMs and diffusion models that fits in one consumer GPU rather than splits on multiple machines, there are already pretty good solutions such as AI Horde (formerly Stable Horde) [0]. Notably, it's the default provider that powers SillyTavern. It also has an interesting economy model of kudos.

[0] https://stablehorde.net/

Imagine if there was some kind of way to cryptographically 'prove' your GPU is doing some kind of 'work' here and contributing to the network.

You could even hand out some kind of digital 'currency' to the people proportional to the amount of work their doing!

The recently discussed https://meshllm.cloud/ is the one I've been playing with but don't have the hardware to try with a model split between nodes, which apparently is supported and just not part of the public demo.
Any updates here to make this work on more modern llms?
This is awesome. Now we also need distributed training of models!
This is super cool technology, but will be abused to death
I did a triple take. Many years old but an early ground breaking project. Has some exploratoy ideas a year or two ago last I looked. Crazy to see Petals at 16 frontage (at time of writing)
[flagged]
The reason i want to run models at home is to be able to experiment on them, change things, see how they get better or worse.
OP this project has been abandoned for years, you didn't even check the last commit date?