Pretty cool someone is still doing this. Training in house LLMs was extremely popular in 2023-2024, back when domain-specific LLMs could easily top GPT in their field. In my field alone (tax/HR tech) I remember that…
Bluesky has <100 employees so it could be possible
I used this a lot over the weekend, and it's a really intelligent and strange model. It gets really agentic at the higher reasoning levels. It does the basics like goal tracking and tool calling well, but more than…
This is the full proposal: https://www.whitehouse.gov/wp-content/uploads/2026/07/Scienc... It's pretty thoughtful about diagnosing the problems of the current system, but I don't know about the solutions. Disbursing…
Following the citations, the original source is this 2013 paper: https://pmc.ncbi.nlm.nih.gov/articles/PMC4279242/pdf/nihms58... The paper compares women in STEM to women outside STEM (as the baseline). However, the…
Referring to ownership, not renting
For those who don't know what's going on in Korea, KOSPI is up 3x in the last year and a large amount of HBM employees have made huge amounts of bonus pay. This has led to an insane FOMO frenzy in a society that's…
The newest generation of LLMs have a very high obsession level with autonomous problem solving. For example, I'm often working with Codex in a WSL terminal. GPT-5.6 often does things autonomously that I thought would…
This study found that between 2022-2024, there was a negative correlation between "jobs with high AI exposure" (i.e., tech jobs) and % change in wages. According to them, software engineering has both the biggest wage…
The entire safety evals industry is essentially funded and controlled by OpenAI/Anthropic. Notice that on recent models, they exclusively use internal testing or black box external vendors (e.g., Gray Swan) whose entire…
The numbers from Anthropic seem heavily cherry-picked, Artificial Analysis has Opus 5 at 1.25x the cost of Sonnet and 2x the cost of GPT 5.6 and K3. https://artificialanalysis.ai/?cost=cost-per-task
In 2026, people do already prefer to use an LLM for coding help rather than Stackoverflow. The reasons (people on SO can be rude, interactions are stressful, replies are slow, etc) are all risks associated with human…
They get money from subscriptions and tokens, same as for closed-weight providers. Yes they'll lose some traffic to hosting services, but many users prefer to use the original training company since they have a…
Similar story here. DS models are absurdly good value for mid-end tasks. I've found DSv4 Flash to be ~10% the cost of GPT-5.4-mini/Claude Haiku at similar performance. We used to pay OpenAI >1m$/month for fraud…
With all due respect, there's zero chance that humans with relevant knowledge scored these themselves. Reading through the winners, every single one is classic vibe-research, with the usual pure-LLM-research patterns: -…
This is jaw-droppingly lazy slop. The authors really didn't put in even an ounce of thought or effort.
If there was some grand strategy for all Chinese labs, surely it'd have leaked by now. I think its more likely that: - Companies can still make money from commodities - Chinese labs only have 5-10% the valuation of…
Strictly dominates both Sonnet 5 and Opus 4.8 in both cost and performance: https://artificialanalysis.ai/models/comparisons/kimi-k3-vs-... https://artificialanalysis.ai/models/comparisons/kimi-k3-vs-...
This study only looks at one specific vendor algorithmn (a job assesment given by a company called pymetrics)
Pretty cool someone is still doing this. Training in house LLMs was extremely popular in 2023-2024, back when domain-specific LLMs could easily top GPT in their field. In my field alone (tax/HR tech) I remember that…
Bluesky has <100 employees so it could be possible
I used this a lot over the weekend, and it's a really intelligent and strange model. It gets really agentic at the higher reasoning levels. It does the basics like goal tracking and tool calling well, but more than…
This is the full proposal: https://www.whitehouse.gov/wp-content/uploads/2026/07/Scienc... It's pretty thoughtful about diagnosing the problems of the current system, but I don't know about the solutions. Disbursing…
Following the citations, the original source is this 2013 paper: https://pmc.ncbi.nlm.nih.gov/articles/PMC4279242/pdf/nihms58... The paper compares women in STEM to women outside STEM (as the baseline). However, the…
Referring to ownership, not renting
For those who don't know what's going on in Korea, KOSPI is up 3x in the last year and a large amount of HBM employees have made huge amounts of bonus pay. This has led to an insane FOMO frenzy in a society that's…
The newest generation of LLMs have a very high obsession level with autonomous problem solving. For example, I'm often working with Codex in a WSL terminal. GPT-5.6 often does things autonomously that I thought would…
This study found that between 2022-2024, there was a negative correlation between "jobs with high AI exposure" (i.e., tech jobs) and % change in wages. According to them, software engineering has both the biggest wage…
The entire safety evals industry is essentially funded and controlled by OpenAI/Anthropic. Notice that on recent models, they exclusively use internal testing or black box external vendors (e.g., Gray Swan) whose entire…
The numbers from Anthropic seem heavily cherry-picked, Artificial Analysis has Opus 5 at 1.25x the cost of Sonnet and 2x the cost of GPT 5.6 and K3. https://artificialanalysis.ai/?cost=cost-per-task
In 2026, people do already prefer to use an LLM for coding help rather than Stackoverflow. The reasons (people on SO can be rude, interactions are stressful, replies are slow, etc) are all risks associated with human…
They get money from subscriptions and tokens, same as for closed-weight providers. Yes they'll lose some traffic to hosting services, but many users prefer to use the original training company since they have a…
Similar story here. DS models are absurdly good value for mid-end tasks. I've found DSv4 Flash to be ~10% the cost of GPT-5.4-mini/Claude Haiku at similar performance. We used to pay OpenAI >1m$/month for fraud…
With all due respect, there's zero chance that humans with relevant knowledge scored these themselves. Reading through the winners, every single one is classic vibe-research, with the usual pure-LLM-research patterns: -…
This is jaw-droppingly lazy slop. The authors really didn't put in even an ounce of thought or effort.
If there was some grand strategy for all Chinese labs, surely it'd have leaked by now. I think its more likely that: - Companies can still make money from commodities - Chinese labs only have 5-10% the valuation of…
Strictly dominates both Sonnet 5 and Opus 4.8 in both cost and performance: https://artificialanalysis.ai/models/comparisons/kimi-k3-vs-... https://artificialanalysis.ai/models/comparisons/kimi-k3-vs-...
This study only looks at one specific vendor algorithmn (a job assesment given by a company called pymetrics)