No. The 996 working culture in this page is actually *996.ICU*. "Work by '996', sick in ICU", an ironic saying among Chinese developers, which means that by following the "996" work schedule (9 am to 9 pm, six days a…
> I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over…
Good to see new open source/weight LLM families. Waiting for the final release of https://huggingface.co/IFM/K2-Horizon-32B .
For people who want to quantize kv caches for longer context, basically, if the model checkpoint you use didn't trained to use quantized kv caches (and all most all of model checkpoints you can access didn't), you…
In contrast to memory-bounded autoregressive decoding, diffusion decoding can be computation-bounded. Devices with a decent amount of compute but very small memory, e.g. nvidia consumer-grade GPUs, could benefit from…
Thank you Qwen team for this release. Compared to closed weight (especially unreleased and access-limited) and open weight/source large sparse MoE LLMs/VLMs, open weight/source small dense models benefits public the…
Thanks Bonsai team. Now open weight LLMs/VLMs/LMMs are becoming even larger to the extent that consumer-grade hardware are no longer able to run these models. In contrast, quantization and pruning make the model better…
The criminal complaint missing in this article: https://www.justice.gov/usao-ndil/media/1450651/dl Basically, Microsoft logs things from windows users, including and not limiting to, the machine GDID, IPs that come with…
Yeah, searxng instance reliability problem goes deeper than simple uptime. The google response time function on public instance page [1], a good measure of the search engine availability of one instance, is broken for…
I've been using searxng for several years now. I don't run my own instances because the inhumane network censorship imposed by GFW, and proxy detection enforced by search engines. Instead, I rely on public instances on…
Android developer verification program, together with recent reCAPTCHA push [1], and Manifest v2 force depreciation on chrome [2], make one thing crystal clear. When companies like GOOGLE talks about things in the name…
Glad to see more open models. However, where are the 31b models?
I've been using FUTO voice input for several month. It's not perfect as I have to manually do the correction like always. However, it definitely saves me some time and effort. It also helps me to degoogle, starting from…
We paid for your things, AMD. If you want to strip some features from things we bought after the purchasing, you must ask me and every other customers for consents explicitly, with a reasonable explanation, and before…
The github issue that never made it in this news: https://github.com/AMDESE/AMDSEV/issues/292 Silent enshittification in the name of updates is getting out of hand. There are several evidences that downgrading…
Seems like the affected version of chrome is 150 and 151, and the timewindow is near. Because of ungoogled-chromium, I'm used to manually update the browser. This time I won't update until MV2 problem is solved…
Please do not claim you trained a new model, only to got caught red-handed by others. There are already several people or groups did that, got caught, and vanished in no time. Check how the "authors" of "this model"…
I don't know how open source AI wins. The description is too vague for serious discussions. What I do know is that, once closed source AI groups become anti-you, you should punish them, or help open source groups, or…
Hilarious read. I laugh out loud multiple times during the read. In the end, I think amd should pay the author simply for the will of debugging for a broken software written by amd, as well as the sheer amounts of loose…
Thanks gemma team for this release. Compared to autoregressive decoding, diffusion is huge for local MoE inference because of the improved token generation efficiency, especially for normal GPU + ram offload setting.…
More rants about local inference, consider yourself warned. Together with bf16 related deliberate hardward degrades on consumer-level nvidia gpus, i.e., gtx 10, rtx 20, 30, 40, 50 series, things gets sour really quickly.
From the perspective of a local llm user, I think the qat doesn't solve the major problem of the gemma models. Gemma family (gen 1 to gen 4) is consistent with extreme range of activations, i.e., 600000, essentially…
A small dense multimodal model with audio support, interesting. Wait, *Excluding Chinese language. This is ... curious. P.S. Where is gemma 4 124b?
This website brings me some good chuckles. Now I really know how powerful an on-demand bullsh*t generator is.
Like the recent copilot silent signing incident, the without consent part is blatant foul move. If you don't like be treated like anything but human, you should seriously consider replacing chrome with ungoogled…
No. The 996 working culture in this page is actually *996.ICU*. "Work by '996', sick in ICU", an ironic saying among Chinese developers, which means that by following the "996" work schedule (9 am to 9 pm, six days a…
> I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over…
Good to see new open source/weight LLM families. Waiting for the final release of https://huggingface.co/IFM/K2-Horizon-32B .
For people who want to quantize kv caches for longer context, basically, if the model checkpoint you use didn't trained to use quantized kv caches (and all most all of model checkpoints you can access didn't), you…
In contrast to memory-bounded autoregressive decoding, diffusion decoding can be computation-bounded. Devices with a decent amount of compute but very small memory, e.g. nvidia consumer-grade GPUs, could benefit from…
Thank you Qwen team for this release. Compared to closed weight (especially unreleased and access-limited) and open weight/source large sparse MoE LLMs/VLMs, open weight/source small dense models benefits public the…
Thanks Bonsai team. Now open weight LLMs/VLMs/LMMs are becoming even larger to the extent that consumer-grade hardware are no longer able to run these models. In contrast, quantization and pruning make the model better…
The criminal complaint missing in this article: https://www.justice.gov/usao-ndil/media/1450651/dl Basically, Microsoft logs things from windows users, including and not limiting to, the machine GDID, IPs that come with…
Yeah, searxng instance reliability problem goes deeper than simple uptime. The google response time function on public instance page [1], a good measure of the search engine availability of one instance, is broken for…
I've been using searxng for several years now. I don't run my own instances because the inhumane network censorship imposed by GFW, and proxy detection enforced by search engines. Instead, I rely on public instances on…
Android developer verification program, together with recent reCAPTCHA push [1], and Manifest v2 force depreciation on chrome [2], make one thing crystal clear. When companies like GOOGLE talks about things in the name…
Glad to see more open models. However, where are the 31b models?
I've been using FUTO voice input for several month. It's not perfect as I have to manually do the correction like always. However, it definitely saves me some time and effort. It also helps me to degoogle, starting from…
We paid for your things, AMD. If you want to strip some features from things we bought after the purchasing, you must ask me and every other customers for consents explicitly, with a reasonable explanation, and before…
The github issue that never made it in this news: https://github.com/AMDESE/AMDSEV/issues/292 Silent enshittification in the name of updates is getting out of hand. There are several evidences that downgrading…
Seems like the affected version of chrome is 150 and 151, and the timewindow is near. Because of ungoogled-chromium, I'm used to manually update the browser. This time I won't update until MV2 problem is solved…
Please do not claim you trained a new model, only to got caught red-handed by others. There are already several people or groups did that, got caught, and vanished in no time. Check how the "authors" of "this model"…
I don't know how open source AI wins. The description is too vague for serious discussions. What I do know is that, once closed source AI groups become anti-you, you should punish them, or help open source groups, or…
Hilarious read. I laugh out loud multiple times during the read. In the end, I think amd should pay the author simply for the will of debugging for a broken software written by amd, as well as the sheer amounts of loose…
Thanks gemma team for this release. Compared to autoregressive decoding, diffusion is huge for local MoE inference because of the improved token generation efficiency, especially for normal GPU + ram offload setting.…
More rants about local inference, consider yourself warned. Together with bf16 related deliberate hardward degrades on consumer-level nvidia gpus, i.e., gtx 10, rtx 20, 30, 40, 50 series, things gets sour really quickly.
From the perspective of a local llm user, I think the qat doesn't solve the major problem of the gemma models. Gemma family (gen 1 to gen 4) is consistent with extreme range of activations, i.e., 600000, essentially…
A small dense multimodal model with audio support, interesting. Wait, *Excluding Chinese language. This is ... curious. P.S. Where is gemma 4 124b?
This website brings me some good chuckles. Now I really know how powerful an on-demand bullsh*t generator is.
Like the recent copilot silent signing incident, the without consent part is blatant foul move. If you don't like be treated like anything but human, you should seriously consider replacing chrome with ungoogled…