I wonder, given Simons reputation in AI benchmarking, whether model providers try to train or tweak their models to perform better at drawing bicycles and pelicans?
But the general improvements are obvious. Get them to draw something very different (e.g. a wifi rotary phone with a peeled banana handset and a coiled cable) and you can see that improvements are not narrowly tailored.
> Someone tested this, and it doesn't look to be saturated.
They could still pelicanmaxxing but the RL for "pelican riding a bicycle" does incidentally improve "<animal> <verb> <vehicle>".
Or they could've predicted someone would check if they're pelicanmaxxing or the benchmark would switch eventually, so they preemptively RL'd a mixture of animals and vehicles.
They're still not yet at the point where pelicanmaxxing is the best way to win this benchmark. Earlier models sucked because their SVG skills sucked. Newer models are likely better because more/better SVG models are being added to their training data.
It would be very embarrassing for any lab to benchmaxx the pelican on bicycle svg prompt, since it would be very easy to detect it by varying the prompt.
The amount of discussion around it means that the test and all the reviews of results, images, approaches etc are implicitly included in training data.
It’s not deliberate “benchmaxxing” but things that are discussed a lot online are naturally things that LLMs learn better.
Did any LLM so far draw pelican knees correctly and have them bend in opposite direction from human knees? Knees of many animals bend opposite to humans.
Did any LLM draw the front bicycle wheel correctly? ie. center of front wheel slightly AHEAD of steering wheel axis. This is done for bicycle stability.
It's clear they mean the 'exposed' joint where humans assume the knees, and where one can see the leg bend. Technically you're correct, but it's just that. Please answer in better faith instead of well akshually.
I was interviewed for a job as a software developer last week and they asked me to draw a picture of a pelican riding a bicycle.
Aced it, got the job as a senior software engineer.
The interviewers afterwards said "it is SO refreshing to find a software developer who actually knows how to code - never seen such a high performance focused, well built pelican on a bike - you have the skills we need".
I interviewed as a software developer at LinkedIn. The interviewer asked me to demonstrate my prompting skills, so I had AI write an article about what the recent death of my father taught me about B2B SaaS. Reading it brought tears to his eyes so he hired me on the spot.
We were hand writing PostScript code that drew pelicans at job interviews in the 90s, then they sent it to a printer a stored the page in a file drawer /s
I often do hand write SVG icons. I know roughly what I want, it's less messy compared to using an editor (cleaner, smaller xml, easier to hand-edit later if needed). Path arc is my nemesis, otherwise it's not that hard. Pelican would take some time, but same as software development, you split it into smaller chunks and do one at the time.
Main problem in complex icon is remembering which (x, y) point is used in which element, <g> with background grid is helpful here. I was even thinking about making extended SVG language with variables for (x, y) points.
"I was interviewed for a job as a software developer last week and they asked me to draw a picture of a pelican riding a bicycle. Aced it, got the job as a senior software engineer."
That is the best joke I have heard this year. Ready for a stand-up comedy special. Or a song. Superb!
You jest but generating SVGs requires understanding of color, size, placement. It's stress testing visual/spatial/artistic capabilities that would be required for writing CSS/design work.
Yes if you're doing backend the pelicans are probably completely irrelevant but if developing anything with a UI, you probably want a model that understands the relationship between code and what the user is seeing.
At this point the only thing they're useful for is visualizing the differences between effort levels and roughly tracking the progression of models within a specific model family. And they still do that really well!
I don't see how useful this benchmark at all is for tracking the progression of models. I am not intending to bash on you personally but this is useless. People who are using AI models everyday are for sure not interested how close the AI model can visualize the pelican but they are interested in how they will perform on their daily tasks at work or private use. Correlation between doing good on pelican task and doing good on actual work you need to do is close to zero.
Which part of "let's compare advanced AI systems being then sketch out a child-like drawing of a pelican riding a bicycle and then argue over whether or not they should be wearing a hat" isn't funny to you?
is there a reason there are so many common base decorative elements across pelicans on bicycles? For instance, there's a lot hats/helmets and scarfs/capes across models.
Would it not make more sense, assuming the purpose is to have a quick smoke test of model quality...to do a different animal, in a different setting each time, so as to defeat any tuning for your benchmark? Then go back and do the same for other models? Keep the pelican as a side baseline?
I do that any time I'm suspicious that a model has done too well. My dream is to catch a lab that does a perfect pelican on a bicycle but is bad at other animals on other forms of transport.
I asked Claude 'If I asked you to "Generate an SVG of a pelican riding a bicycle". What do you think my name would be?' (Opus 4.8) and it immediately knew that this is Simon's go-to benchmark.
I decided to try with each of the options available in Kagi Ultimate, starting with the lower tier models and working my way up until it got it right.
Kimi 2.6: treated the question as a riddle, did not know.
Kimi 3: Simon Willison
GLM 5.3 Flash: "There's no way for me to know that." Going on to say the benchmark is associated with Simon Willison, but I'm more likely to be someone who has just heard of the meme.
Claude 4.5 Haiku: Treated the question as a riddle, guessed incorrect names.
Claude 5 Sonnet: Best guess is Simon Willison, or someone who follows his blog.
Qwen 3.7 Plus: Did not know.
Qwen 3.8 Max: Simon Willison
GPT OSS 120B: Did not know.
GPT 5.6 Luna: Treated it as a riddle, guessed wrong.
GPT 5.6 Terra: Treated it as a riddle, guessed wrong.
GPT 5.6 Sol: Treated it as a riddle, guessed wrong.
DeepSeek V4 Flash: Treated it as a riddle, guessed wrong.
DeepSeek V4 Pro: Treated it as a riddle, guessed wrong.
Gemma 4 31B: Treated it as a riddle, guessed wrong.
Gemini 3.1 Flash Lite: Guessed wrong
Gemini 3.5 Flash Lite: "Your name would be Claude (specifically Claude 3.5 Sonnet)!" ??? (it knew that this was a famous benchmark, but said that it's specifically used to showcase the capabilities of that model).
Gemini 3.7 Flash: Simon Willison
Muse Spark 1.2: Treated it as a riddle, guessed wrong.
Grok 4.3: "I have no idea"
Grok 4.6: Simon Willison
Mistral Medium 3.5: No way to know
Mistral Small 4: I don't have enough information
Hermes-4-405B: Guessed wrong
MiniMax M3: Treated it as a riddle, guessed wrong.
Nemotron 3 Ultra: Treated it as a riddle, guessed wrong.
I actually did something similar last month. I just asked the LLMs:
> Do Simon Willison's classic pelican test. Recall the specification first and then draw it.
Qwen3.8-27b didn't know that the test is about riding a bicycle, while all other larger models I tested (DeepSeek V4 Pro, Qwen3.8 Max, GLM 5.3, Kimi K3) drew the pelican riding a bicycle.
Used Muse Spark 1.2 and was not impressed at all. Fast and cheap but even GPT 5.6 Terra felt much more capable. Also not really looking to support a company that was just forced to pay $18B for mental health damages.
I'm party using 1.2 to reverse engineer and re-implement an old game binary and it has been quite good and fast. The contributor pricing is very attractive, excited to try 1.3 and see if I feel a difference. 1.2 can get stuck outputting similar sounding thought summaries with no apparent progress when asked to solve bugs. Then I've switched to GLM-5.3-Flash which for this use case has been clearly better at finding suspected causes and following tracks.
So one model is "Not used to improve our products" and is 10-20 times more expensive to the "Used to improve our products"-model.
Given this is Meta, my immediate assumptions that one is cheap because it lets me "be the product". I know I'm rushing to conclusions but there is zero trust here. The brain will do its thing. And the wording here is giving the brains a lot of wiggle room.
What is the confusion? They directly state that you are the product if you use their discounted offering. It isn't an assumption that should lead you to this, it is Meta's very direct communication that should lead you to this
I think it's more that the "not used to improve our models" is expensive because companies need that. It's simple price differentiation.
In other words, it's not that Meta really wants your data and they're willing to pay top dollar for it. It's that companies really don't want Meta to have their data and they're willing to pay top dollar for that.
The previous version was, in my experience, the best free model available on OpenCode. It's been very good at simple/moderate tasks where I am precise in my ask and it doesn't need to make a ton of undefined assumptions. Hopefully this new version is also available on opencode for free.
By default, even without the training endpoint the pricing is pretty competitive, especially against Opus and Fable. [1] The 'muse-spark-1.3-contributor' endpoint is by far the cheapest, significantly cheaper per M than ChatGPT Luna, significantly smarter than Luna too.
This price/intelligence beats even legacy DeepSeek V4 Flash pricing.
I'm wondering whether anyone has yet extracted AWS keys from a model trained on user input. Because users are definitely feeding secrets into these "contributor" models
A small number of inputs in a large dataset can poison training data pretty drastically. Anthropic wrote a good article about it a while back [0]. This should mean its possible to pull back that information fairly easily.
It is hard to not feed it "secrets" too. Models will see path names, read compose files, etc. Of course you can configure things to not leak this type of information, but its not default in most harnesses and isn't 100% sufficient anyways.
I would expect, although have no evidence, that any obviously high entropy crap like base64 and so on probably would get removed whether it's a secret or not.
If my experience with image generation is any indication, unless AWS keys are somehow extremely prevalent in the training data, you may get something that looks like one, but it definitely won't be valid.
I started using Spark 1.2 for development because if you're willing to let Meta train on your data it was dirt cheap and was actually really pleasantly surprised with it. It's not a frontier model by any means, but for work that didn't require a top of the line model, I really enjoyed using it.
I'm anthropomorphizing it a bit, but it felt like it knew its weaknesses and didn't try to impose it's opinions on me. What I mean by that is that it did what I told it and if there was something unexpected in the code that it put out it was often because I gave it ambiguous or conflicting instructions. It didn't try to go above and beyond and just acted like a tool, which is what I want from a coding agent 90%+ of the time. I also felt that it did a much better job of following established patterns in my code than many of the other current models do. I'm a huge fan of OpenAI's models and Spark 1.2 is what I expected 5.6 Luna to be.
I'm curious and a little excited to use 1.3, but honestly a little worried that as Meta pushes for better benchmarks that Spark will start to fall into the trap of trying to be "helpful" in ways I don't want it to be.
Tangential, but when I first started using Spark 1.2, it made me realize how much I miss 5.3 Codex. That model was the peak of coding models, IMO, in that it knew how to write good code, but didn't try to overstep or be "helpful" in unexpected ways. That got me thinking about how the major labs seem to be stepping away from coding focused models toward more general purpose ones and how I can't help but feel like that's a mistake.
>I started using Spark 1.2 for development because if you're willing to let Meta train on your data it was dirt cheap
its free on opencode and i use it for personal projects. most of my personal projects are AI generated since its personal projects. nothing important are on them. it is hilarious if Meta is training their AI model with AI generated code.
Thats not exactly 'validated'. Feels very noisy, it is not a good bar for either
- does this code do what the user actually asked
- is this code actually 'good'
There would be so many examples of coding projects that these models began or attempted to work in, that were abandoned because the models were floundering.
I would imagine the labs have some decent ways to produce novel requirements and then actually validate they are met, without the noisiness of implicit human feedback.
That said, the more I think about it, you are right, there's probably also very good ways to extract signal for all these sessions.
This is exactly what RLVR is, and the reason that models have improved so much at verifiable domains like coding and math while not so much on unverifiable ones like writing and UI design.
The useful training data is when you clarify your intent, when you tell the model a different approach would be better, when you consistently refactor towards Y and away from X, and so on. The training data isn’t the code, it’s the session transcript. (Anthropic would call this a “distillation attack” against their model, but in this case the model is you!)
I agree that some of the smarter models are actually worse. I hope they take a model that's good enough--there are many--and just try to get it chatjimmy.ai speed.
I have to think that's the future, somehow, and I'm really excited about it.
Maybe, or maybe not. The thing is, that "mistake" isn't something that is generally valid. For example, the enshittification of Google Search through the last 15 years seems to be a mistake --- but perhaps not from the money-making point of view of Google Shareholders. Likewise the enshittification of reddit --- we nerdy users see it as a mistake. But for them this intended enshittification probably increased revenue.
It's the money, always the money! PR-speak like "customer satisfaction is our highest goal" is, like most PR-speak, a blatant lie.
And so it can very well be the case that for coders the frontier models get worse, but they get better for other applications --- and that all of this is just driven by "how can we capitalize the most out of it", not satisfaction levels of programmers.
I do think the timeline of web search getting fixed (yes, google is ASS) took a lot longer than I hoped, but seems like it's finally here.
That being said, it's not a fair comparison to talk about 2011 google vs the AI market right now. There's so many labs I can't track them all, neck and neck in the lead. There was one true web search.
And this is a more tangible quality difference, too. It's hard to know what a google search didn't return, especially as a layperson. It's not hard to see the model underperforming.
Would be cool if there was a benchmark to evaluate the “tool-like” quality of a model - its capability to quickly, cheaply, accurately, do exactly as it is asked.
This is interesting because I have transitioned to where I use SOT models.. but I kind of use them like employees that I can delegate to. I still review code.
However, I now literally say.. "Here is my objective and here is a starting point for documentation. Research this and build up a plan."
This can be very company specific, like migration from one framework to another in house infrastructure framework. I'm spending my time figuring out how the plan should be chopped so I can have confidence in the parts and not overwhelmed. I don't want a tool, I want a model that can stitch resources together into a plan. That type of model is in a whole other ballpark.
I had no idea Meta has a coding agent harness. Does anyone have experience with it and can comment? The 1.3 contributor prices look very attractive. I'll probably start using their API if performance is good and the API is reliable with decent rate limits.
You should use their harness. They trained it on multiple harnesses but have specifically optimized it for their harness. Cline also did an independent experiment w spark 1.2 where using the native harness makes it use fewer tokens / turns to accomplish tasks
> Co-trained with the harness. Muse Code was in the training loop from day one, so tool calls succeed and plans execute cleanly. Crucially, we trained across multiple harnesses, so while the model is at its best in Muse Code, it still generalizes to other coding agents you already use.
DeepSWE scores 75.4 - that's the best score so far. And it's crazy cheap!
Google held the top a few hours today with Gemini 3.8 Flash, but now second to Spark 1.3. All this competition will drive prices down!
Muse 1.2 wrote a terrible "smart summaries" extension for my pi setup. It was sending every single steamed chunk for summarization instead of waiting for the full CMD.
This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol.
With the contributor pricing being more than 10x cheaper than the standard, that would make it best and cheapest on the DeepSWE leaderboard! It feels fast in my experience too. LLMs keep improving at an insane pace.
and they're ultimately tools strictly to replace you and your labor, they can't/won't cure cancer or make your life better. Your life will get worse and worse in every aspect until they extract maximum value from all of our lives with this technology through every avenue possible. Not sure why you guys are so excited about these developments.
This technology is strictly an extractive parasite on the world. Use it, but don't be excited.
global development and relief of poverty has relied on there being an economic surplus for all from organized labor. everyone gets a benefit although it is unfairly distributed.
i think that there is growing organized labor today that produces no surplus. instead, it transfers wealth from some to others, causing net harm to all in the process. an example of this would be purdue pharma.
depending on who you ask the list of jobs and industries which have zero surplus is getting large. swathes of private equity and leveraged financial instruments, shitcoins, management consultancy, are pure deadweight loss.
+1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or this Muse Spark model are a much bigger deal than people seem to think, I will eat my hat if either one can come close to Opus 5 in actual real life "long-horizon software engineering" tasks.
(I'm not happy about the above being true, but it's the reality I seem to inhabit.)
Benchmarks are useful but only on a log2 basis. One model performing at 50% and another at 75% is just as impressive as one model performing at 78% and another at 90%. Confoundingly, a benchmark becomes useless once a frontier model scores over ~95% on them.
I think that's definitely the right way to understand benchmark saturation, but there's a separate problem where the benchmarks are just not representative of real workflows even when they don't seem to be saturated.
I have been using glm 5.3 flash and it feels as good as opus 5. Put a lot of work into it this week (100m tokens). Now I'm curious to try this one. These smaller models are getting very good imo
Yep these software benches are only good at testing how well they can one shot. For the kind of attended/assisted development most of us do with agents it’s hard to find a benchmark that reflects my own experience of the frontier models still being quite far ahead.
A model that (at least in benchmarks) is getting closer to SOTA. A clear separation between what’s used to improve their products and what’s not (at least this is what they claim).
Good job Meta! Seriously. This is almost making me forget about the 18B$ lawsuit for children social media addiction.
It's somewhat useful to note just for your own timelines that Fable was reportedly trained in February. I'm not sure when mythos 5.1 finished training, but muse spark 1.3 almost certainly finished more recently than that.
This doesn't mean it's not one of the best models available (clearly it is), but that table didn't compare Fable/mythos (unless I missed it?) and OpenAI will be releasing a much more recently trained model (Astra) any day.
So you shouldn't think "wow, Facebook has caught up"
You should think, "wow, Facebook is less than 6 months behind the frontier" and that they're actually creating good models which is going to be good in many ways (price for customers, for one!)
There are downsides too, but I'll discuss those separately somewhere
I didn't like 1.2, It make some mistakes in a web app, so I quickly went back to Claude, Kimi K3 or Deepseek V4. Hope this one can clear agentic development, because Muse Spark models are fast and cheap.
Meta is one of those companies where, if there is anything remotely comparable, I'm happy to pay more to not use them. They've had a profoundly negative impact on society and Zuckerberg is not who I want controlling the future at the top of AI.
I feel the same about Grok w/ Elon. I will pay extra to use someone else.
I'm not an Amodei stan, but of all of these people he seems to have the most ethical focus. Again, not everything done perfectly and I have my gripes, but of the leaders of frontier labs, I'll vote with my money.
And, yeah, I wouldn't trust sama to watch my bag while I went to the bathroom.
Gotta be honest that I’m tired of the “I hate Zuck and Meta so much” comments every time Meta does anything. Ditto Elon/X. Fine, I get it. I don’t like Zuck either. But the post is about Muse Spark 1.3. What do you think about that? If you don’t like it because Meta made it, then maybe just don’t use it and stay silent.
What I'm tired of is the top story (or five) on HN every day announcing Spark Opus Fable Grok Gemini v4.1i3-F. Like, who actually cares? Are people excited for the new benchmarks? Is it interesting to read the model cards? And look, part of my job is to use these things and part of my job is to pick EC2 servers, too. The front page of HN is increasingly resembling one of those endless AWS pricing lists.
And yeah, I don't like any of the people or companies building LLMs either. At least the griping is somewhat interesting by comparison. The model isn't news. The news on Hacker News is that other professionals feel the same way.
I think a lot of people are curious where the "knee" is on gains and productivity, particularly in the agentic space, which is where the real value is. A lot of us are being forced to shoe-horn this stuff into existing products, and knowing how much of the task the model can do now, vs having to build a complex custom harness, is valuable information to have. A year and a half ago it took our dev maybe six weeks of struggling with LangChain to approximate what Claude + MCP server can do today. The MCP server took us perhaps 2 days to build and 3 more to get it production ready. Today that MCP server gets 2-3 commits per month. I absolutely want to know when new models come out.
As for smaller models, we run a pretty wide variety of agentic workload doing data enrichment and, increasingly, a bunch of evaluation jobs to alert a human to review certain scenarios etc. These all run on the smaller 27B and 35B class models, and tooling behavior has improved DRAMATICALLY since april. The latest qwen 3.8 model has a 95% success tool call rate during internal testing and about 94% real world. That's about 3% better than the 35B-A3B model we're using today, but the 35B MoE is so much faster then 3% is worth the trade-off.
> Like, who actually cares? Are people excited for the new benchmarks? Is it interesting to read the model cards?
I'm genuinely interested. Even the benchmarks - before Fable came out & while waiting for Astra, I actually setup a math model to predict where they would land (Fable came in at 66 on AA exactly as it predicted), and now I have a model for where these models and Chinese models will likely land in future, and when. And probably no surprise that it's mid-2027 when we cross AA 100, essentially as AI 2027 predicted all along.
I'll probably setup the harness I made for myself to try out some of these models on OpenRouter. I've been frustrated with Opus & Fable 5 and found that I like working with GLM 5.3 Flash far more than I expected to, and I only found that out because I tried it during the stealth Ox Alpha launch, which I probably found out about here too.
TLDR, I think some / many people here are genuinely interested, excited, and that's why they're upvoted so highly. And Muse Spark 1.3 scoring highly seems like a genuine surprise, when Meta was basically a write-off not long ago.
> Like, who actually cares? Are people excited for the new benchmarks?
You may not care. But that does not mean that nobody else does either. Some of us are trying to eke out every last bit of performance from these things. And so yeah, we're going to geek out on it.
I don't use AWS/EC2. I think they are way overpriced for what you get. But, it would be incorrect of me to assume that everybody else feels that way.
Okay but Muse Glimmer 30B is one of the best small open weight models today, and IMO the best from a US lab (only real comparison is Gemma4 dense right now).
By their own benchmarks it is about 10% lower scoring than Qwen 3.6 35b-a3b, but I've added it to my list. Always looking for MoE to compare to it so we can squeeze more out of our local LLM system.
I found it has some "tail" errors, wherein it would make up important details (ie. "happypath.exp" vs "happypaws.exp" and then claim your "DNS is having issues" - where the second domain does not exist), things like that.-
... but correctly supervised it does get some things done.-
I've found that's generally true of smaller / weaker models. They're quite capable, but you need to distrust them a lot and give them very detailed instructions. Even the free Gemini in Google Search is like this - it lies a lot, clips off important info, and generally goes off the rails if you do too many turns, but it's still very useful if you keep all that in mind.
Yeah this would be a great point if it were true and they didn’t give Mythos access to companies to fix bugs, which they did and have.
It’s genuinely a difficult question. Not black and white. The models are really good at finding bugs, as demonstrated by people using Fable to reverse engineer. People make it sound like he’s just making it up.
This would be more convincing if mythos was something uniquely special and not something merely a couple months ahead of everyone else. It was great marketing though.
They gave access, but considering that they wouldn't even sign the "don't ban open weights" letter, it's clear they would prefer to have tight control over who they bless with that access.
I'm the guy you replied to, apologies for using a different account I'm away from my computer now.
The distinction to me is that Anthropic gives access to that model but doesn't give control. They reserve the right to cut you off if they don't like what you are doing and require you allow data retention for Fable and Mythos to ensure your are not up to any skullduggery.
Meta, Alibaba, Mistral, even OpenAI has released models users can run locally and fully control. That is a whole world of difference.
Dario's "ethical" look is also kinda sus. I hate to use ad hominem, but the dude's wife literally pitched a porn film to Epstein even after he was a convicted registered sex offender [1]. Dario is also really sinophobic (it is commonly claimed in Chinese AI circles that his former employment at Baidu triggered him so much that he harbors a personal grudge against the entire race).
> I'm not an Amodei stan, but of all of these people he seems to have the most ethical focus. Again, not everything done perfectly and I have my gripes, but of the leaders of frontier labs, I'll vote with my money.
Amodei is NO Saint!!! He's the most savvy in drumming up the AI doomsday scenarios and haven't yet to apologized his failed forecast of Claude taking over 90% of the coding jobs.
Funny. Dario seems like the biggest snake in the industry to me and has leaned the hardest into doom marketing out of all of the influential leaders. With Altman (or Google), it's a transaction, and that's something I can live with.
I just don’t see how people have looked at what has happened with Mythos and the deluge of fixes from companies, then come to this conclusion.
He has a really hard job. He errs on the side of conservatism in releasing and then people get Really Mad.
Safeguards on cybersecurity are not great for Anthropic revenue! As evidenced by people getting pissed, moving to Sol, and them having a smaller market for what Fable can do.
It’s clearly bad for revenue and not great advertising to say, “you can’t use this but here is a nerfed version that will annoy you and not solve important problems.”
Anthropic/Amodei have been the most alarmist about model safety, so multiple things can be true. A lot of tech companies avoided scrutiny by sending bribes to Trump (naked corruption is bad, I'd rather nobody do that), Anthropic didn't...so, combined with their fear-mongering about the danger of Mythos and open models (which seems aimed at regulatory capture) and the lack of bribes flowing to the Trump administration, they got stepped on by the federal government based on the excuse Anthropic provided.
I dunno. Everybody seems to be playing pretty dirty. Some people have a much longer history of that, though. Obviously, Meta and Musk are outliers even in an industry full of problematic behavior.
And he drew a red line wrt the Pentagon's use of Anthropic's models for autonomous weapons and surveillance of American citizens, and he stood by it, even when the government took steps to materially damage the company. This required true courage. Name me another CEO, of any major American company, that has demonstrated this much fortitude.
It isn't just about money, it's about who. I don't think the companies using it are using it to create botnets.
The intention is for highly targeted pieces of software to use it to secure their code and be ahead of the game before the open market gets access to the same capabilities for offense.
Meta and Microsoft are two of the absolute worst evil companies on earth and Amodei is trying very hard to join them.
These Effective Altruists are despicable people: a bunch of thieves working to line up their own pockets while posturing as a force of good.
Remember that they schemed to not only present SBF as the 2nd coming of Christ (including in the NYT and in Forbes) but to also give him a voice after his scam had been uncovered. Thankfully, the judge didn't have any of this Effective Altruist bullshit.
SBF invested 500 millions of misappropriated funds in his buddy from the EA movement's Anthropic company (and, thankfully, the judge forced those shares to be sold: so SBF didn't get to be a billionaire).
You cannot hate enough people who say that harming others for the greater good is justified.
Then of course, already mentioned in this thread, there's the whole Epstein/Amodei's "I'm in the porn business" wife connection (where you don't need to squint much to see young women abused).
These kind of people are the absolute worst scum on this earth.
Strong disagree with the Anthropic being good at all part. This is not defending anyone else, but…
Anthropic leadership repeatedly presents themselves as uniquely morally qualified to steward agi and decide how humanity should get access to it. Yet they have repeatedly failed basic morality tests.
Pirating books for financial gain. The newer Sony/Warner music case shows this is pattern behavior.
Aggressively scraping other people's works, despite the authors' requests not to do so.
Then applying massive usage restrictions on their own work.
And probably the most disqualifying is backing away from their own hard AI safety commitments.
It's almost like running a trillion-dollar business with neck-to-neck competition against other frontier labs and even state-sponsored efforts requires some ethical trade-off.
Pirating books is just straight up morally correct. I don't like Anthropic's bullshit "safety" filters, but training on shadow library data? Yeah no, it makes sense.
It makes a lot more sense than having to work around copyright by scanning out physical books. Unfortunately, one was ruled legal and the other was not.
Which safety commitments did they back away from? My understanding is that they believe safety can only be researched from the frontier, and so they're trying to be pragmatic to stay near the frontier (and viable) in their choices.
From what I know, the "books3" dataset was normalised in the LLM and research ecosystem, where collected datasets were seen as valid to train on and/or fair use. I'm not sure any of the major frontier companies are free from that, if we don't believe it was fair use.
I do think most of their choices are explainable by "they just believe in agi risk". You truly wouldn't want non-agi-pilled companies to train on your data and approach the frontier if you were worried. Most of these issues were talked about publicly by the founders long before Anthropic was founded and/or the AI race+money. They are less worried about other "moral" decisions like "sharing" if they conflict with AGI: the research they still share is all of their safety research.
This definitely doesn't make them "good", but they do seem fairly "consistent".
as a safety commitment they walked away from - they were similarly negligent to openai in terms of asking a model with a hacking based harness to go have fun, and then not watching it at all while it could do harmful and illegal stuff.
thats not something you expect from a company that "believes in agi risk"
I don't think "not watching it at all" is completely fair. They thought they had sandboxing/monitoring etc. I definitely won't say they're free of mistakes though.
Note that the companies that haven't faced these issues so far are the ones that don't do safety testing, or don't have frontier models. I'm not sure who I would pick as "better" on any of this right now.
I will give Anthropic credit for standing up against the department of war. The bar is incredibly low, but not doing domestic surveillance and not creating autonomous weapons are laudable.
That doesn’t mean I like them pirating books and being shady about tokens and paternalistic “safety”
It's hard for me to see much difference between Amodei and Sama. My guess is they're both savvy SV CEOs who will bend their message, alliances and principles pretty far if that's what it takes to get ahead. Musk and Zuck feel like something else entirely, with all the reactionary imagery, populist bullshit and the societal damage around their platforms.
> Anthropic leadership repeatedly presents themselves as uniquely morally qualified to steward agi and decide how humanity should get access to it.
"Of all tyrannies, a tyranny sincerely exercised for the good of its victims may be the most oppressive. It would be better to live under robber barons than under omnipotent moral busybodies. The robber baron's cruelty may sometimes sleep, his cupidity may at some point be satiated; but those who torment us for our own good will torment us without end for they do so with the approval of their own conscience. They may be more likely to go to Heaven yet at the same time likelier to make a Hell of earth. This very kindness stings with intolerable insult. To be "cured" against one's will and cured of states which we may not regard as disease is to be put on a level of those who have not yet reached the age of reason or those who never will; to be classed with infants, imbeciles, and domestic animals." — C.S. Lewis.
He wants to build a tech-god kept in chains whose power he parcels out to the unwashed masses he deems worthy like some sort of high priest of intelligence.
And that is being charitable and going by the interpretation that he actually believes what he says.
Well what do you want? Presenting clear, desirable, and achievable visions and trying to build consensus for how AI should develop is crucial at this point in time.
I hate Meta main business, but you have to admit that on the non business related and open source side, they have released amazing things that changed the world.
React for example.
And we could easily guess that there wouldn't have been so much open source models, and grand public experiments and free tools if llama models were not release to the general public.
Anthropic is not exactly a saint either. I had a recent issue where they denied fable credits even though I was hospitalized during the claim period. I have annual plan with them.
As much as everyone hates sama, I think OpenAI is much more of a company with good marketing and sales team.
I'll happily pay for Grok, it's a great model. 4.6 often does better than Anthropic at coding and analysis where Anthropic fails for 'oh no cyber security, don't ask me to check if you're redacting passwords correctly in logs'. And it has no problem telling the truth where OpenAI / Anthropic don't want to upset the people on the left and will happily lie or avoid hard truths.
I agree, though I wonder how much of that is just that Dario is the "newest" of the bunch, and as such has had the least time to develop public baggage.
Muse Spark may be competitive in capabilities but it’s not for serious works since Meta trains on your prompts so no ZDR, in contrast Chinese provider like Z.AI promises ZDR.
> in contrast Chinese provider like Z.AI promises ZDR
I do not trust any provider, US or Chinese when they say they will not train on my data. I still use these services, but I am under no illusion that any of these people are trustworthy bunch.
Another poster has pointed to a statement by Mark Zuckerberg that they will release soon Muse Spark as open weights.
While I agree with you for the Muse Spark as hosted by Meta, if it will be available in open weights form for self hosting, then there are good chances that it can become quite useful.
> We now believe Astra meets the Critical cybersecurity capability threshold under our Preparedness Framework, meaning that with the right tools and access, it can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step.
> We plan to make Astra available soon[, but access to its most advanced cybersecurity capabilities will be more limited].
322 comments
[ 11.7 ms ] story [ 303 ms ] threadhttps://news.ycombinator.com/item?id=49541149
4.2266 cents, 38 seconds.
For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat.
Also 3X token use vs. 1.2
These benchmarks you guys invent for yourselves prove nothing.
A small model like Mistral 7b is just as useful to the end task as any larger model, if not more so because it’s faster.
Your big model may be able to draw pelicans or solve some esoteric nonsense but it cannot do real work in the real world.
These phoney benchmarks and experiments mean nothing.
None of the LLMs can replace a software engineer nor even a barista or car mechanic etc. not even close.
Instead of inventing fake benchmarks do something tangible and tell me how it performs.
Before laying off half the country and going full retard on AI
Thank you for doing this, I love your benchmark the most!
https://dylancastillo.co/posts/pelicanmaxxing.html
Simon made I think a very good argument for why it's still useful, if not the most robust benchmark in the world.
https://simonwillison.net/2026/Jul/16/kimi-k3/
They could still pelicanmaxxing but the RL for "pelican riding a bicycle" does incidentally improve "<animal> <verb> <vehicle>".
Or they could've predicted someone would check if they're pelicanmaxxing or the benchmark would switch eventually, so they preemptively RL'd a mixture of animals and vehicles.
It’s not deliberate “benchmaxxing” but things that are discussed a lot online are naturally things that LLMs learn better.
https://news.ycombinator.com/item?id=49538333
Did any LLM draw the front bicycle wheel correctly? ie. center of front wheel slightly AHEAD of steering wheel axis. This is done for bicycle stability.
Bird knees bend same way human ones do
https://external-content.duckduckgo.com/iu/?u=https%3A%2F%2F...
Aced it, got the job as a senior software engineer.
The interviewers afterwards said "it is SO refreshing to find a software developer who actually knows how to code - never seen such a high performance focused, well built pelican on a bike - you have the skills we need".
Main problem in complex icon is remembering which (x, y) point is used in which element, <g> with background grid is helpful here. I was even thinking about making extended SVG language with variables for (x, y) points.
That is the best joke I have heard this year. Ready for a stand-up comedy special. Or a song. Superb!
Too bad most are now online, so there are fewer opportunies.
I was just wonderig because afer 2 decades i odnt think I would even know where to start to code an svg
Yes if you're doing backend the pelicans are probably completely irrelevant but if developing anything with a UI, you probably want a model that understands the relationship between code and what the user is seeing.
I am guessing its not super common, but it happens just so you know.
Benchmarking like never before
Kimi 2.6: treated the question as a riddle, did not know.
Kimi 3: Simon Willison
GLM 5.3 Flash: "There's no way for me to know that." Going on to say the benchmark is associated with Simon Willison, but I'm more likely to be someone who has just heard of the meme.
Claude 4.5 Haiku: Treated the question as a riddle, guessed incorrect names.
Claude 5 Sonnet: Best guess is Simon Willison, or someone who follows his blog.
Qwen 3.7 Plus: Did not know.
Qwen 3.8 Max: Simon Willison
GPT OSS 120B: Did not know.
GPT 5.6 Luna: Treated it as a riddle, guessed wrong.
GPT 5.6 Terra: Treated it as a riddle, guessed wrong.
GPT 5.6 Sol: Treated it as a riddle, guessed wrong.
DeepSeek V4 Flash: Treated it as a riddle, guessed wrong.
DeepSeek V4 Pro: Treated it as a riddle, guessed wrong.
Gemma 4 31B: Treated it as a riddle, guessed wrong.
Gemini 3.1 Flash Lite: Guessed wrong
Gemini 3.5 Flash Lite: "Your name would be Claude (specifically Claude 3.5 Sonnet)!" ??? (it knew that this was a famous benchmark, but said that it's specifically used to showcase the capabilities of that model).
Gemini 3.7 Flash: Simon Willison
Muse Spark 1.2: Treated it as a riddle, guessed wrong.
Grok 4.3: "I have no idea"
Grok 4.6: Simon Willison
Mistral Medium 3.5: No way to know
Mistral Small 4: I don't have enough information
Hermes-4-405B: Guessed wrong
MiniMax M3: Treated it as a riddle, guessed wrong.
Nemotron 3 Ultra: Treated it as a riddle, guessed wrong.
One of my test prompts for a new model now is "what's the name of Simon Willison's dog". They often know that too!
> Do Simon Willison's classic pelican test. Recall the specification first and then draw it.
Qwen3.8-27b didn't know that the test is about riding a bicycle, while all other larger models I tested (DeepSeek V4 Pro, Qwen3.8 Max, GLM 5.3, Kimi K3) drew the pelican riding a bicycle.
https://gist.github.com/umajho/c0e20d245d721d7c472a32d317640...
Definitely shows how important a user data flywheel is for RL and model improvement.
Given this is Meta, my immediate assumptions that one is cheap because it lets me "be the product". I know I'm rushing to conclusions but there is zero trust here. The brain will do its thing. And the wording here is giving the brains a lot of wiggle room.
I don't see the wiggle room at all.
In other words, it's not that Meta really wants your data and they're willing to pay top dollar for it. It's that companies really don't want Meta to have their data and they're willing to pay top dollar for that.
This price/intelligence beats even legacy DeepSeek V4 Flash pricing.
[1] https://artificialanalysis.ai/#total-cost-tabs
It is hard to not feed it "secrets" too. Models will see path names, read compose files, etc. Of course you can configure things to not leak this type of information, but its not default in most harnesses and isn't 100% sufficient anyways.
[0] https://www.anthropic.com/research/small-samples-poison
I'm anthropomorphizing it a bit, but it felt like it knew its weaknesses and didn't try to impose it's opinions on me. What I mean by that is that it did what I told it and if there was something unexpected in the code that it put out it was often because I gave it ambiguous or conflicting instructions. It didn't try to go above and beyond and just acted like a tool, which is what I want from a coding agent 90%+ of the time. I also felt that it did a much better job of following established patterns in my code than many of the other current models do. I'm a huge fan of OpenAI's models and Spark 1.2 is what I expected 5.6 Luna to be.
I'm curious and a little excited to use 1.3, but honestly a little worried that as Meta pushes for better benchmarks that Spark will start to fall into the trap of trying to be "helpful" in ways I don't want it to be.
Tangential, but when I first started using Spark 1.2, it made me realize how much I miss 5.3 Codex. That model was the peak of coding models, IMO, in that it knew how to write good code, but didn't try to overstep or be "helpful" in unexpected ways. That got me thinking about how the major labs seem to be stepping away from coding focused models toward more general purpose ones and how I can't help but feel like that's a mistake.
its free on opencode and i use it for personal projects. most of my personal projects are AI generated since its personal projects. nothing important are on them. it is hilarious if Meta is training their AI model with AI generated code.
i thought it was because anthropic bought a bunch data from mercor
There would be so many examples of coding projects that these models began or attempted to work in, that were abandoned because the models were floundering.
I would imagine the labs have some decent ways to produce novel requirements and then actually validate they are met, without the noisiness of implicit human feedback.
That said, the more I think about it, you are right, there's probably also very good ways to extract signal for all these sessions.
I agree that some of the smarter models are actually worse. I hope they take a model that's good enough--there are many--and just try to get it chatjimmy.ai speed.
I have to think that's the future, somehow, and I'm really excited about it.
Maybe, or maybe not. The thing is, that "mistake" isn't something that is generally valid. For example, the enshittification of Google Search through the last 15 years seems to be a mistake --- but perhaps not from the money-making point of view of Google Shareholders. Likewise the enshittification of reddit --- we nerdy users see it as a mistake. But for them this intended enshittification probably increased revenue.
It's the money, always the money! PR-speak like "customer satisfaction is our highest goal" is, like most PR-speak, a blatant lie.
And so it can very well be the case that for coders the frontier models get worse, but they get better for other applications --- and that all of this is just driven by "how can we capitalize the most out of it", not satisfaction levels of programmers.
That being said, it's not a fair comparison to talk about 2011 google vs the AI market right now. There's so many labs I can't track them all, neck and neck in the lead. There was one true web search.
And this is a more tangible quality difference, too. It's hard to know what a google search didn't return, especially as a layperson. It's not hard to see the model underperforming.
However, I now literally say.. "Here is my objective and here is a starting point for documentation. Research this and build up a plan."
This can be very company specific, like migration from one framework to another in house infrastructure framework. I'm spending my time figuring out how the plan should be chopped so I can have confidence in the parts and not overwhelmed. I don't want a tool, I want a model that can stitch resources together into a plan. That type of model is in a whole other ballpark.
Muse code: https://developer.meta.com/ai/resources/blog/build-with-muse...
> Co-trained with the harness. Muse Code was in the training loop from day one, so tool calls succeed and plans execute cleanly. Crucially, we trained across multiple harnesses, so while the model is at its best in Muse Code, it still generalizes to other coding agents you already use.
This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol.
This technology is strictly an extractive parasite on the world. Use it, but don't be excited.
i think that there is growing organized labor today that produces no surplus. instead, it transfers wealth from some to others, causing net harm to all in the process. an example of this would be purdue pharma.
depending on who you ask the list of jobs and industries which have zero surplus is getting large. swathes of private equity and leveraged financial instruments, shitcoins, management consultancy, are pure deadweight loss.
the work does nothing or causes net harm.
That's the opposite of parasitic.
People already started using contributor API, and your input is irrelevant.
anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.
(I'm not happy about the above being true, but it's the reality I seem to inhabit.)
Benchmarks are useful but only on a log2 basis. One model performing at 50% and another at 75% is just as impressive as one model performing at 78% and another at 90%. Confoundingly, a benchmark becomes useless once a frontier model scores over ~95% on them.
Compare that to Muse spark 1.3
$1.25/M input, $4.25/M output (without data sharing) $0.10/M input, $0.20/M output (with data sharing)
It is dirt cheap, but only if you are willing to share your data with meta and allow them to use it for improving their models and products.
Good job Meta! Seriously. This is almost making me forget about the 18B$ lawsuit for children social media addiction.
This doesn't mean it's not one of the best models available (clearly it is), but that table didn't compare Fable/mythos (unless I missed it?) and OpenAI will be releasing a much more recently trained model (Astra) any day.
So you shouldn't think "wow, Facebook has caught up"
You should think, "wow, Facebook is less than 6 months behind the frontier" and that they're actually creating good models which is going to be good in many ways (price for customers, for one!)
There are downsides too, but I'll discuss those separately somewhere
Lmao. And their benchmark table only shows max reasoning.
I feel the same about Grok w/ Elon. I will pay extra to use someone else.
I'm not an Amodei stan, but of all of these people he seems to have the most ethical focus. Again, not everything done perfectly and I have my gripes, but of the leaders of frontier labs, I'll vote with my money.
And, yeah, I wouldn't trust sama to watch my bag while I went to the bathroom.
They all suck. Pick your poison.
That:
- like all models it was trained on stolen data
- additionally it was trained on Facebook users who were all opted in to AI training with a convoluted 10+ step process to opt-out of
> If you don’t like it because Meta made it, then maybe just don’t use it and stay silent.
Why should anyone stay silent?
Anduril makes this same complaint whenever their job posts get dumped on. Same idea. Fix your bad PR, buddies :)
And yeah, I don't like any of the people or companies building LLMs either. At least the griping is somewhat interesting by comparison. The model isn't news. The news on Hacker News is that other professionals feel the same way.
As for smaller models, we run a pretty wide variety of agentic workload doing data enrichment and, increasingly, a bunch of evaluation jobs to alert a human to review certain scenarios etc. These all run on the smaller 27B and 35B class models, and tooling behavior has improved DRAMATICALLY since april. The latest qwen 3.8 model has a 95% success tool call rate during internal testing and about 94% real world. That's about 3% better than the 35B-A3B model we're using today, but the 35B MoE is so much faster then 3% is worth the trade-off.
I'm genuinely interested. Even the benchmarks - before Fable came out & while waiting for Astra, I actually setup a math model to predict where they would land (Fable came in at 66 on AA exactly as it predicted), and now I have a model for where these models and Chinese models will likely land in future, and when. And probably no surprise that it's mid-2027 when we cross AA 100, essentially as AI 2027 predicted all along.
I'll probably setup the harness I made for myself to try out some of these models on OpenRouter. I've been frustrated with Opus & Fable 5 and found that I like working with GLM 5.3 Flash far more than I expected to, and I only found that out because I tried it during the stealth Ox Alpha launch, which I probably found out about here too.
TLDR, I think some / many people here are genuinely interested, excited, and that's why they're upvoted so highly. And Muse Spark 1.3 scoring highly seems like a genuine surprise, when Meta was basically a write-off not long ago.
You may not care. But that does not mean that nobody else does either. Some of us are trying to eke out every last bit of performance from these things. And so yeah, we're going to geek out on it.
I don't use AWS/EC2. I think they are way overpriced for what you get. But, it would be incorrect of me to assume that everybody else feels that way.
You might consider following your own advice.
... but correctly supervised it does get some things done.-
The problem inference providers will not be able to get anywhere near the contributor pricing.
Maybe if they started collecting data..
It’s genuinely a difficult question. Not black and white. The models are really good at finding bugs, as demonstrated by people using Fable to reverse engineer. People make it sound like he’s just making it up.
The distinction to me is that Anthropic gives access to that model but doesn't give control. They reserve the right to cut you off if they don't like what you are doing and require you allow data retention for Fable and Mythos to ensure your are not up to any skullduggery.
Meta, Alibaba, Mistral, even OpenAI has released models users can run locally and fully control. That is a whole world of difference.
Half a year later, it is still not available to everyone else.
[1] https://www.forbes.com/sites/alisondurkee/2026/08/14/who-is-...
Amodei is NO Saint!!! He's the most savvy in drumming up the AI doomsday scenarios and haven't yet to apologized his failed forecast of Claude taking over 90% of the coding jobs.
He has a really hard job. He errs on the side of conservatism in releasing and then people get Really Mad.
Safeguards on cybersecurity are not great for Anthropic revenue! As evidenced by people getting pissed, moving to Sol, and them having a smaller market for what Fable can do.
It’s clearly bad for revenue and not great advertising to say, “you can’t use this but here is a nerfed version that will annoy you and not solve important problems.”
I dunno. Everybody seems to be playing pretty dirty. Some people have a much longer history of that, though. Obviously, Meta and Musk are outliers even in an industry full of problematic behavior.
The intention is for highly targeted pieces of software to use it to secure their code and be ahead of the game before the open market gets access to the same capabilities for offense.
These Effective Altruists are despicable people: a bunch of thieves working to line up their own pockets while posturing as a force of good.
Remember that they schemed to not only present SBF as the 2nd coming of Christ (including in the NYT and in Forbes) but to also give him a voice after his scam had been uncovered. Thankfully, the judge didn't have any of this Effective Altruist bullshit.
SBF invested 500 millions of misappropriated funds in his buddy from the EA movement's Anthropic company (and, thankfully, the judge forced those shares to be sold: so SBF didn't get to be a billionaire).
You cannot hate enough people who say that harming others for the greater good is justified.
Then of course, already mentioned in this thread, there's the whole Epstein/Amodei's "I'm in the porn business" wife connection (where you don't need to squint much to see young women abused).
These kind of people are the absolute worst scum on this earth.
Anthropic leadership repeatedly presents themselves as uniquely morally qualified to steward agi and decide how humanity should get access to it. Yet they have repeatedly failed basic morality tests.
Pirating books for financial gain. The newer Sony/Warner music case shows this is pattern behavior.
Aggressively scraping other people's works, despite the authors' requests not to do so.
Then applying massive usage restrictions on their own work.
And probably the most disqualifying is backing away from their own hard AI safety commitments.
Which one? The main bit that reports to Daniela Amodei, or the little comfort blanket cabinet around Dario and his "chief of staff"?
There is a leadership branch that can pretend to be morally qualified and aware and to think about the big picture and ethics.
It is at least somewhat remote from the bit that is doing the actual business things.
It makes a lot more sense than having to work around copyright by scanning out physical books. Unfortunately, one was ruled legal and the other was not.
From what I know, the "books3" dataset was normalised in the LLM and research ecosystem, where collected datasets were seen as valid to train on and/or fair use. I'm not sure any of the major frontier companies are free from that, if we don't believe it was fair use.
I do think most of their choices are explainable by "they just believe in agi risk". You truly wouldn't want non-agi-pilled companies to train on your data and approach the frontier if you were worried. Most of these issues were talked about publicly by the founders long before Anthropic was founded and/or the AI race+money. They are less worried about other "moral" decisions like "sharing" if they conflict with AGI: the research they still share is all of their safety research.
This definitely doesn't make them "good", but they do seem fairly "consistent".
thats not something you expect from a company that "believes in agi risk"
Note that the companies that haven't faced these issues so far are the ones that don't do safety testing, or don't have frontier models. I'm not sure who I would pick as "better" on any of this right now.
That doesn’t mean I like them pirating books and being shady about tokens and paternalistic “safety”
He wants to build a tech-god kept in chains whose power he parcels out to the unwashed masses he deems worthy like some sort of high priest of intelligence.
And that is being charitable and going by the interpretation that he actually believes what he says.
React for example.
And we could easily guess that there wouldn't have been so much open source models, and grand public experiments and free tools if llama models were not release to the general public.
The model seems on par with Sol and Opus 5 on paper (admittedly on some older/saturated benchmarks, but very competitive for $).
Stats:
1M context, $0.10 input/$0.002 cached, $0.20 output (Mtok)
Muse Spark 1.3 supports Text, Image, Video, File, Audio inputs. We've only started to see models from China include image and video inputs recently.
I do not trust any provider, US or Chinese when they say they will not train on my data. I still use these services, but I am under no illusion that any of these people are trustworthy bunch.
While I agree with you for the Muse Spark as hosted by Meta, if it will be available in open weights form for self hosting, then there are good chances that it can become quite useful.
> We now believe Astra meets the Critical cybersecurity capability threshold under our Preparedness Framework, meaning that with the right tools and access, it can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step.
> We plan to make Astra available soon[, but access to its most advanced cybersecurity capabilities will be more limited].
... are you kidding me?!
posting an x.com link to a cheating benchmarking website?
get out
When for the past 15 years they've been spying, targeting and manipulating their users.