Distillation requires you to have the actual logits of each token from the teacher model, which in practice means having the model itself. What you're describing is just synthetic data. Note Anthropic misused the term…
I don't think you know what distill means
This and their newer Aura A1 I've wanted to get cause of the compact size, but they seem to lack support for US 5g bands
As someone who's done a lot of llm fiction, that reads as pretty typical slop, and very human centered, nothing like a museum
I highly doubt it's behind in practice, except for Anthropic
Cache hit % on openrouter is not a good metric, it's mainly driven by openrouter's own provider juggling than the providers themselves
LLMs overuse em-dashes and use them in specific ways that are annoying to read. Stringing together clauses unnecessarily, spacing words out instead of using a comma etc. I don't kind em-dashes in decent human writing,…
Glm had made vision models in the past. Look up GLM 5v. The only question now is if it's 5.3v, 5.4/5.5 or a dedicated flash/vision model
They already do, extensively. Its not Shakespeare, in fact, it sucks at prose and creativity, like most newer llms. But people are not as alike as you think. I doubt I share your unique preferences. That said, I don't…
It's an opt in feature in openrouter to get a 1% discount.
Love the X logo that goes to bluesky
Might be in part because multiple times I've seen signs/poster etc only good the agent to contradict it once you get there, adding or removing requirements. So you stop trusting it. A whiteboard I might trust because it…
It doesn't. It's called preserved reasoning and every recent reasoning model does it
More than 8 primary schools in a small town seems a lot no?
Yeah and it's degraded significantly since then. Older llms were still mostly language focused and had a lot of latent knowledge about things like writing styles. Now it's crowded out in favor of agenetic work,…
> Muse Spark 1.2 is available today in Muse Code and in Meta Model API with expanded global access Wasn't the previous one us only? This is probably the biggest part of the post Anyone know if muse code is open source?
Two things 1. If we're using native harnesses, I'd have preferred you use kimi code, not opencode 2. The variation in the two kimi providers just shows how you can't trust n = 1 trials
What about your examples has an llm tell? I don't trust pangram 4 much. There was a post here earlier confirming it fails for many others.
It feels like that should be a golden opportunity for competitors, but every time a competitor makes a decent replacement, the big tech company either buys it or briefly invests in their product again to make it good…
Why is the title "Africa" and not Morocco
Screams it in fact
Jeez way to ruin of the few remaining joys of flying. Lock everyone in a tin can. You will experience reality through screens only and you will enjoy it
Llms use them a lot more than humans, including this blog post. Like all slop. There's a reason it's called slop and it's not because of restraint
For a fair comparison, you should compare to K3 (which AA has not tested yet unfortunately) and GPT 5.6 Sol also on medium or the closest equivalent
This is a lot of words to say "switch models instead of writing a plan file and starting a new session" which I do anyway. That said, it takes me a while to reads plans and often I'll take a break, by when the cache has…
Distillation requires you to have the actual logits of each token from the teacher model, which in practice means having the model itself. What you're describing is just synthetic data. Note Anthropic misused the term…
I don't think you know what distill means
This and their newer Aura A1 I've wanted to get cause of the compact size, but they seem to lack support for US 5g bands
As someone who's done a lot of llm fiction, that reads as pretty typical slop, and very human centered, nothing like a museum
I highly doubt it's behind in practice, except for Anthropic
Cache hit % on openrouter is not a good metric, it's mainly driven by openrouter's own provider juggling than the providers themselves
LLMs overuse em-dashes and use them in specific ways that are annoying to read. Stringing together clauses unnecessarily, spacing words out instead of using a comma etc. I don't kind em-dashes in decent human writing,…
Glm had made vision models in the past. Look up GLM 5v. The only question now is if it's 5.3v, 5.4/5.5 or a dedicated flash/vision model
They already do, extensively. Its not Shakespeare, in fact, it sucks at prose and creativity, like most newer llms. But people are not as alike as you think. I doubt I share your unique preferences. That said, I don't…
It's an opt in feature in openrouter to get a 1% discount.
Love the X logo that goes to bluesky
Might be in part because multiple times I've seen signs/poster etc only good the agent to contradict it once you get there, adding or removing requirements. So you stop trusting it. A whiteboard I might trust because it…
It doesn't. It's called preserved reasoning and every recent reasoning model does it
More than 8 primary schools in a small town seems a lot no?
Yeah and it's degraded significantly since then. Older llms were still mostly language focused and had a lot of latent knowledge about things like writing styles. Now it's crowded out in favor of agenetic work,…
> Muse Spark 1.2 is available today in Muse Code and in Meta Model API with expanded global access Wasn't the previous one us only? This is probably the biggest part of the post Anyone know if muse code is open source?
Two things 1. If we're using native harnesses, I'd have preferred you use kimi code, not opencode 2. The variation in the two kimi providers just shows how you can't trust n = 1 trials
What about your examples has an llm tell? I don't trust pangram 4 much. There was a post here earlier confirming it fails for many others.
It feels like that should be a golden opportunity for competitors, but every time a competitor makes a decent replacement, the big tech company either buys it or briefly invests in their product again to make it good…
Why is the title "Africa" and not Morocco
Screams it in fact
Jeez way to ruin of the few remaining joys of flying. Lock everyone in a tin can. You will experience reality through screens only and you will enjoy it
Llms use them a lot more than humans, including this blog post. Like all slop. There's a reason it's called slop and it's not because of restraint
For a fair comparison, you should compare to K3 (which AA has not tested yet unfortunately) and GPT 5.6 Sol also on medium or the closest equivalent
This is a lot of words to say "switch models instead of writing a plan file and starting a new session" which I do anyway. That said, it takes me a while to reads plans and often I'll take a break, by when the cache has…