[–] darkolorin 1y ago ↗ I made it! 90 t/s on my iPhone with llama1b fp16We completely rewrite the inference engine and did some tricks. This is a summarization with llama 3.2 1b float16. So most of the times we do much faster than MLX. lmk in comments if you wanna test the inference and I’ll post a link.
1 comment
[ 79.5 ms ] story [ 424 ms ] threadWe completely rewrite the inference engine and did some tricks. This is a summarization with llama 3.2 1b float16. So most of the times we do much faster than MLX. lmk in comments if you wanna test the inference and I’ll post a link.