7 comments

[ 0.25 ms ] story [ 7.7 ms ] thread
This looks useful for somebody with a 16-32GB Mac Mini interested in running larger MoE models.

I've been working on a mesh environment that relies on an explicit prefix hash in the request and enables constructing a new session with a cached pre-filled system prompt specific KV-cache beyond what OpenAI-compatible APIs offer. Can you see a feature like that being supported?

(comment deleted)
does it need M1 Pro specifically, or any M1 could work?