2 comments

[ 3.0 ms ] story [ 14.5 ms ] thread
(comment deleted)
hey guys, we've optimised our inference engine to run Meta Muse Glimmer 30B faster than any other engine (llama.cpp etc) on Mac. Best performance on M5 Pro.