Qwen is trained off of gpt’s outputs. This is both a positive and negative
They are not training a whole model in a matter of days
How can a vision model have taste?
Claude’s native harness is pretty bad though relatively. Pretty sure at one point it was scoring last on harness related benchmarks using Opus at the time
Google’s pro models are almost certainly bigger than Openai’s lol
Going to assume you didnt capture the data but could you add time taken to completion for each if you have it?
Does anyone know how they turn html to enable powerpoints so seamlessly?
The benchmarks they released
LLM comment spotted
I mean looking antigravity, jules & gemini cli, they have have no problem with their developers fighting for resources
Qwen is trained off of gpt’s outputs. This is both a positive and negative
They are not training a whole model in a matter of days
How can a vision model have taste?
Claude’s native harness is pretty bad though relatively. Pretty sure at one point it was scoring last on harness related benchmarks using Opus at the time
Google’s pro models are almost certainly bigger than Openai’s lol
Going to assume you didnt capture the data but could you add time taken to completion for each if you have it?
Does anyone know how they turn html to enable powerpoints so seamlessly?
The benchmarks they released
LLM comment spotted
I mean looking antigravity, jules & gemini cli, they have have no problem with their developers fighting for resources