I tried Qwen MoE a while back. Using my 8GB RX470, somehow got 10+ token/sec, lot's of trial and error with llama.cpp config, and it's still slow to be used for my usecase. Even at 12 to 16 IMO it's slow. For chat,…
LLM skew the time estimate tho. Now everybody expect stuff based on LLM work instead of normal human work. I/we can choose to solve problem normally, but the expectations have changed.
At first I thought you meant QR for payment, which is weird because most people (at least in SEA, where I lived) consider less friction and more convenient than cash or cards. But it turns out you meant QR for menu,…
Tried it just now. The onboarding process could be better, for example guide user to pick the available models if providers is setup but it's not anthropic. Wasted a bit of time foguring out that the provider is…
Same, which is why I just go with my own pace and stop being FOMO. Fortunately my org is not that gung-ho about pushing AI.
Emulator?
I tried Qwen MoE a while back. Using my 8GB RX470, somehow got 10+ token/sec, lot's of trial and error with llama.cpp config, and it's still slow to be used for my usecase. Even at 12 to 16 IMO it's slow. For chat,…
LLM skew the time estimate tho. Now everybody expect stuff based on LLM work instead of normal human work. I/we can choose to solve problem normally, but the expectations have changed.
At first I thought you meant QR for payment, which is weird because most people (at least in SEA, where I lived) consider less friction and more convenient than cash or cards. But it turns out you meant QR for menu,…
Tried it just now. The onboarding process could be better, for example guide user to pick the available models if providers is setup but it's not anthropic. Wasted a bit of time foguring out that the provider is…
Same, which is why I just go with my own pace and stop being FOMO. Fortunately my org is not that gung-ho about pushing AI.
Emulator?