At the current pace of iteration from Deepseek I think this is a moot point, they'll keep building their own RL environments, and just keep RL on top of whatever flavor of model they can train/host on their Huawei…
NATO v Yugoslavia 1999 is not like the others
Depends on what software/hardware you'll run it. GGUFs from Unsloth can run on pretty much every single potato; full weights need beefy gpus
benchmark where gemini flash is better than fable btw.
Crazy calling sovereign states "US Puppets".
Minimax has been great for super high speed web/js/ts related work. It compares in my experience to Claude Sonnet, and at times gets stuff similar to Opus. Design wise it produces some of the most beautiful AI generated…
It is arguable that the new Minimax M2.1 and GLM4.7 are drastically above Sonnet 3.7 in capabilities.
cline is used by a lot of devs
The longer "it" reasons, the more attention sinks are used to come to a "better" final output.
My biggest gripe with Ollama is the badly named models, e.g. under deepseek-r1, it defaults to the distill models.
I'm pretty sure that Neosync[0] does this to a pretty good degree, it is open source and YC funded too. [0] https://www.neosync.dev/
If GPT4 is 220B/8 experts, that would be in-line with 3.5 Turbo being a 20B model, and GPT4 being a 55B activation out of a total 220B parameters. It is ultimately all speculation, until Deepseek releases their own 145B…
I personally was affected by this fire, although I've always kept 3 month backups of production data, encrypted, on-site, just in case of emergencies like this. Haven't touched their services for anything production…
It's buried deep in the Gemini report, but goddamn are these incredible stats.
The Albanian takeover of AI continues. It's incredibly exciting!
I can't wait to see this open sourced, there's a lot of sampling strategies that help coding. And I also can't wait to see how much Phind will improve further if the Glaive dataset is added onto it. Edit: Contrastive…
It was known by Polynesians for at least 1000 years before Columbus. See sweet potatoes.
It's missing a lot of crucial details. Nothing on the dataset used, nothing on the data mix, nothing on their data cleaning procedures, nothing on the tokens trained.
Not all of them per se, take a look at something like Mistral. It's a 7B model displaying incredible performance. IMO, we still haven't even scratched the surface of what is possible with small LLMs. Especially not with…
Added, and reached out on Twitter.
I would like to get in touch with you related to books4. Do you happen to have discord? or would twitter be ok? There's currently multiple attempts at creating what you describe as books4.
I've had some success using vast.ai[0] with the Oobabooga LLM WebUI (LLaMA2) instances. One click to start up, minimal editing in the interface settings to enable OpenAI compatible interface. [0] https://cloud.vast.ai/
Yeah wildly inaccurate.
This reads like a Serbian owned business from the North of Kosovo. But anybody following the local politics would know that: 1) Corruption as an issue is disappearing in Kosovo, especially compared to Albania,…
Yes, but the acquisition of that data itself is illegal in almost all jurisdictions, since libgen is treated as a piracy website. Now if there were a pipeline to access books from Amazon or the Google Books project for…
At the current pace of iteration from Deepseek I think this is a moot point, they'll keep building their own RL environments, and just keep RL on top of whatever flavor of model they can train/host on their Huawei…
NATO v Yugoslavia 1999 is not like the others
Depends on what software/hardware you'll run it. GGUFs from Unsloth can run on pretty much every single potato; full weights need beefy gpus
benchmark where gemini flash is better than fable btw.
Crazy calling sovereign states "US Puppets".
Minimax has been great for super high speed web/js/ts related work. It compares in my experience to Claude Sonnet, and at times gets stuff similar to Opus. Design wise it produces some of the most beautiful AI generated…
It is arguable that the new Minimax M2.1 and GLM4.7 are drastically above Sonnet 3.7 in capabilities.
cline is used by a lot of devs
The longer "it" reasons, the more attention sinks are used to come to a "better" final output.
My biggest gripe with Ollama is the badly named models, e.g. under deepseek-r1, it defaults to the distill models.
I'm pretty sure that Neosync[0] does this to a pretty good degree, it is open source and YC funded too. [0] https://www.neosync.dev/
If GPT4 is 220B/8 experts, that would be in-line with 3.5 Turbo being a 20B model, and GPT4 being a 55B activation out of a total 220B parameters. It is ultimately all speculation, until Deepseek releases their own 145B…
I personally was affected by this fire, although I've always kept 3 month backups of production data, encrypted, on-site, just in case of emergencies like this. Haven't touched their services for anything production…
It's buried deep in the Gemini report, but goddamn are these incredible stats.
The Albanian takeover of AI continues. It's incredibly exciting!
I can't wait to see this open sourced, there's a lot of sampling strategies that help coding. And I also can't wait to see how much Phind will improve further if the Glaive dataset is added onto it. Edit: Contrastive…
It was known by Polynesians for at least 1000 years before Columbus. See sweet potatoes.
It's missing a lot of crucial details. Nothing on the dataset used, nothing on the data mix, nothing on their data cleaning procedures, nothing on the tokens trained.
Not all of them per se, take a look at something like Mistral. It's a 7B model displaying incredible performance. IMO, we still haven't even scratched the surface of what is possible with small LLMs. Especially not with…
Added, and reached out on Twitter.
I would like to get in touch with you related to books4. Do you happen to have discord? or would twitter be ok? There's currently multiple attempts at creating what you describe as books4.
I've had some success using vast.ai[0] with the Oobabooga LLM WebUI (LLaMA2) instances. One click to start up, minimal editing in the interface settings to enable OpenAI compatible interface. [0] https://cloud.vast.ai/
Yeah wildly inaccurate.
This reads like a Serbian owned business from the North of Kosovo. But anybody following the local politics would know that: 1) Corruption as an issue is disappearing in Kosovo, especially compared to Albania,…
Yes, but the acquisition of that data itself is illegal in almost all jurisdictions, since libgen is treated as a piracy website. Now if there were a pipeline to access books from Amazon or the Google Books project for…