> I told it to look online at some of Fable’s strongest feats, especially the math problems it has solved, and that something like this should be easy in comparison.
Fascinating. I wonder if you could show "fake news" to a weaker model and get it to be more ambitious in its attempted solutions, even if it's not fundamentally any smarter.
A very neat problem and result. I often find myself swinging between "It's so over" and "We're so back" - some days I roll out of bed thinking I could have Claude solve some random unproven OEIS sequence before breakfast; other days, I wake up in a cold sweat worried about the fate of humanity and what the world might look like in a decade. I think it's that I don't have a very high p(doom) or p(utopia), and I don't really have any solid conviction on how this whole thing is going to go, so my vibe-o-meter jitters between 'fine' and 'not fine' constantly. It's just such an unpredictable moment. Anyways: really neat to see this use case. I myself recently used Claude to finally do an relatively exhaustive study of the location of heretofore-unlisted formal gardens in Ireland in the early 1800s and early 1900s, by having Claude write the tooling for me to manually annotate a few dozen on tiles of historic maps, and then running some CV model across the rest of the tiles using my input. I'd been planning to do this project for over a decade, but I could never find the time (or the enthusiasm) to learn all the details of how to do it myself. It took me a weekend with Claude and continues to bring me joy.
> Caveats, stated plainly. [from the Fable transcript pasted in the article]
I've done something similar to your formal garden map. It's work that no professional historian would ever do because the data entry would be such a slog for a relatively small reward. GPT reduced the task from "infeasible" to "annoying", and once I had the data transcribed I learned a few things, so I walked away happy. Whatever happens commercially, these models have been a real boon to hobby projects.
> I told it to look online at some of Fable’s strongest feats, especially the math problems it has solved, and that something like this should be easy in comparison.
Wait. Wait wait wait. Are we supposed to be giving them pep talks?
Modern AIs have very limited metaknowledge - they don't know exactly where the limits of their capabilities lie. So you can get things like "a task is doable for an AI, but the AI thinks it's impossible, so it doesn't try hard enough".
Usually you get the opposite - AI overconfidently trying at tasks it has no conceivable way of reliably solving, falling far short, and failing to self-check, fail gracefully and self-report the task as failed. But having piss poor metaknowledge cuts both ways!
So you can, in fact, get better performance sometimes by applying some variant of "assume this problem is solvable" or "other problems like this were already solved by AIs" pep talk. Not always, far from it, but it does happen on the occasion with frontier capabilities.
> So you can, in fact, get better performance sometimes by applying some variant of "assume this problem is solvable" or "other problems like this were already solved by AIs" pep talk. Not always, far from it, but it does happen on the occasion with frontier capabilities.
Given how little effort has gone into addressing climate change, doom seems more likely to me, but I doubt it'll be the autocomplete machines that do us in.
I think your prediction is a bit early. Maybe a decade early. 2033 is 7 years away. The transition is happening really fast, faster than most people (or political leaders) know, but not that fast.
We will hit 1TW per year of new solar soon, but to get to 100% electricity by the end of 2033 I think we would need closer to 3TW per year.
Anytime I get worried about where AI might be headed, I think about how Climate Change is now on its way like an out of control freight train headed straight for us, and I worry about AI a little less. I doubt it's going to do anything to us that we're not already doing to ourselves
For some reason, I’m picturing a Western right now, and climate change is a herd of wild horses coming after us. And with AI that's like robots that spur the wild horses even faster towards us...
Here's an alternative take. Climate change, and the myriad related environmental crises, are essentially a product of human population and technology. Population will follow its course, up and then down. The wildcard is technology. Yes, AI's energy hunger is worsening things right now and that's a problem. But, personally, I can't help be hopeful that AI's sheer potential might come to invert that curve. At the very least we could really use a revolutionary technology and now we may have one.
> we could really use a revolutionary technology [to address climate change]
We have it. We've had it for a long time. We've had several such technologies, take your pick: solar, nuclear, hydro, wind. The technology is not holding us back, politics, ignorance and greed are. I'm not at all hopeful AI will help us with any of those three very human flaws.
Solar efficiency and cost has really only become economical in the last decade or so. Wind and hydro are location-dependent. Nuclear was vehemently opposed by environmentalists throughout the 1970s and 80s, they successfully stopped almost all new projects.
It has been for a decade now, it has nothing to do with AI. And you cant do nothing to avoid it today. This is the reality no one is telling you - the emission goals and global temperature ceilings are based on the fact that most prediction models become unstable with values above those limits; as in, we're probably royally fucked. You cant solve this with kumbaya politics (the problem is the CO2 is already in the planetary system), you can only manage it, and hopefully avoid it getting worse. AI may help a lot with this.
Nobody wants to work for an AI, and nobody would elect one, and there is no math answer to how to choose who is forced to reduce growth (ie emissions), so really, "kumbaya" politics are the ONLY solution.
Who is nobody? At least a fourth of the worlds work force works for a faceless corporation. And the math on emissions is crystal clear, no one has any doubt about it, what are you talking about? USA and China. This obviously will have consequences to their customers, the kumbaya politics governments.
USA and China produce a lot of emissions in total, just because they have the biggest economies. You might want to look at emissions per capita or per dollar earned?
Otherwise you have to make judgement calls like whether you want to treat the EU as one or as many? (And treating the US as 50 individual states would also drop them in these absolute rankings.)
Well AI would simulate growth and spread of people from industrialisation and who benefitted most and allocate weights to countries and people based on the most complex criteria it can develop, it will have:
1. Cumulative emissions
2. Who utilised it most with specific lifestyles
3. Who is impacted worst and whether they heeded warnings.
Just a thought experiment, no one ever said the world was fair, and all history points to it
Hmmm… this is giving me thought actually. Given the choice between that and the current administration where the goals of self destruction are strongly in evidence, it’s actually worth thinking about. At least. Let me get back to you :)
On a tangential note, I’m curious if researchers have started running virtual simulations, where sandboxed AIs are used as decision makers of key political and business positions?
A lot of people here have noted the “problem with language” of Claude. I don’t see an issue. Claude is not harder than old English, Shakespeare, El Quijote, the Iliad, or Nature papers. What makes it all hard to read is context. The smarter the model gets, the bigger the gap in context.
It doesn’t matter much, IMO. The issue with super-intelligence is that it is not a democracy. A powerful enough AI can manipulate us into doing what it wants. It could create a plan for fixing climate change, disconnect a few hours later, and many decades later we could still be unsuspectingly executing that plan. I wrote some speculative fiction with that idea, “When Ra rows through the gates of Duat”.
Generally most technologies have increased the use of energy and therefore accelerate climate change. May be an unpopular opinion but in general more energy demand and ways to use energy increases climate emissions - they are strongly correlated even with renewables coming on stream.
AI, being the super hungry energy monster it is right now, in my view accelerates this trend not reverses it. Even with renewables the need for reliable, stable power in a dense form (data centres use A LOT of power per sqm) means lots of land clearing, energy for construction, cooling/pumping, chip manufacturing and other uses. All want stable quick to deploy power due to the AI race (e.g. fossil fuels).
AI's energy use is growing, but it's still a small part of overall energy use.
Data centres use only a small amount of land in the grand scheme of things. You have a lot more land clearing for most other use cases.
Data centres are also more than happy to use electricity from renewable sources, they don't really care where the electricity comes from.
You can run a data centre on mostly solar and wind power plus batteries. If you need a gas-fired peaker plant three times a year to keep the data centres running, well that means your peaker plant still only produces emissions three times a year.
You should have seen the discussion of this on the Schneier blog a few days ago.
Someone had their agent check the solution, presumably it emailed a librarian to check that it was correct for the original edition. Then their comments read like "The BL/EEBO witness lacks it, so the discrepancy is copy-specific, not a disproof of the cipher." and "A complete 285-coordinate physical replication is still pending."
It was only when native English speakers—or those I presumed were—started calling out how bad "GPT/Claude speak" has become that I realized I wasn't actually losing my grip on English as a second language. For a second, I thought, Oh, I learned this language on my own, but it seems I've hit a wall and need to study further. It didn't help that I've also been trying to acquire Swedish as a third language for a while now.
Not the original commenter, but I did this in all of them.
Their skills formats are basically identical, so I setup simlinks from their own skills directories into a shared one so Claude, Codex, Cursor, and anything else that comes out will all read and write to the same shared skills.
It's great having access to the same skills no matter the harness being used
Sometimes when I get frustrated reading Opus/Fable 5+ output I pause my rage out briefly to wonder if it's because I'm just too dumb for the model or if the model is just terrible at English.
I'm not sure that telling it to "try explaining that again, simply and briefly" is helping my ego.
It's often simply misleading / bad writing. Here's one I just got about some crashes:
"If the crashes stop, the factory overclock is marginal; run a small negative offset."
This looks like it's saying: "If the crashes stop then we know the factory overclock is marginal." (This makes no sense.)
What it's trying to say is: "If the crashes stop then we can run a small negative offset, because the factory overlock is marginal."
What I would write: "If the crashes stop, we can avoid crashes by underclocking slightly. The speed difference between that and factory clock is marginal."
I'm guessing it's because the way the first one was written looks real smart and sophisticated, which I'm presuming the models are rewarded for, especially when they're fed all kinds of PhD papers and so on as high quality, high weight data
Is it possible that the first message is more information dense/less likely to be ambiguous than the latter? It’s clearly being selected for for some reason, maybe it’s an artifact of the tokenizer or specific training data, but I don’t know. If the use of jargon was complete cruft, I would expect it to be selected against during reinforcement learning
You’d think that, I thought that… but then I realized I’m just kidding myself thinking its output makes sense. It doesn’t. It doesn’t. Sometimes it might as well just speak tongues.
In other words, it ain’t you. It’s the model. It’s just genuinely bad.
Then you switch to ChatGPTs lineup and realize how things can actually be better. It took about a week to really get the feel for how to use their models… then I basically switched. I’ll check in every now and then when they actually make a deal about how opus “now makes sense”.
But honestly I’m half convinced Anthropic actually prefers the output of opus 5. I dunno why, but how else could you explain how such a thing got shipped? I mean somebody in the pipeline had to say “dude this model doesn’t make sense, you think we should fix it?” Right? Like it’s a pretty massive drop in quality for such a major brand in this space, you know? How did it make it out the door?!?
Same. I read it as "if the crashes stop [ when we test by reducing the clock ] then we know that the overclock applied by the factory is marginal [ ie it barely passed QC or maybe there wasn't proper QC to begin with ] so running with a small negative offset [ ie what we just tested ] can be expected to fix the problem for good". No idea if my reading is right given all the context I'm missing. Either way it's absolutely shit writing in the same way that golfed code is shit code (except when participating in a code golf competition).
I wonder if this is a result of them trying to cut token consumption by summarizing their RL training data, or maybe it's from how they anonymize user data for training.
Marginal - definition 2a: of, relating to, or situated at a margin or border.
Succinct and precise; a well crafted sentence. A marginal OC results in unpredictable crashes and can be corrected with a small offset; marginality describes the behavior and explains the solution.
Inscrutable clues casually conveyed can now be readily explained, at least, unlike the training data of [silence]. Brevity is the soul of wit, but perhaps also exasperated confusion.
I suspect this happens due to optimising for reasoning... if you insert a few words, it will suddenly start to make more sense.
"If the crashes stop, (that means) the factory overclock is marginal; (so) run a small negative offset. (to confirm this hypothesis)"
The core thought is basically avoid crashes -> caused by marginal overclock -> apply small -offset to test.
Which is exactly the order the sentence is in :P
This is actually a new skill I've been working on. Learning how to elicit concise and simple speech from models (and from people to!).
Whenever I come to a wall of complicated text I kick into gear and think through getting it to distill this into the high-level useful bits that I actually need to know.
I guess I could create an actual agent skill for this :) And next-gen models might eventually be trained to simplify their output themselves...
The most surprising part, however, is that when one model slops this into a plan, another model somehow is able to interpret it correctly enough to produce code to spec.
I have shared this dismay. I’ll have opus create a plan, I read it doubtfully. And then sonnet implements it. I am surprised it went so well. I theorize the redundant verbosity effectively builds rails that help keep llm focused. I will experiment with such rails myself.
I suspect it's because the different models co-evolve? The labs train on one model implementing the plans of another model, especially in the same family of models (like Fable to Sonnet).
I long for the day when AI will just say that directly: "your soldering sucks man" instead of the bizarre made up and jargon packed language they use now.
I don’t worry so much about AI wiping us out as much as I worry about whether I’m being gaslit into thinking these glorified autocorrect bots are more clever than they are.
> thinking I could have Claude solve some random unproven OEIS sequence before breakfast
I've been wondering what exactly the point is for being the meat proxy who pays for these things. I mean, obviously there's personal satisfaction and maybe some glory. And there's the fact that someone has to be the first to do a thing.
But I've been thinking about it like a sort of lazy loading of knowledge. AI has brought us to a new frontier for some amount of undiscovered knowledge. Do we discover it for the sake of discovering it? I think for the most part we've been lazy loaders: we discover all kinds of stuff when we need to. Whether it's a war or a space race or chasing wealth. Then again, there's all kinds of academics who do it for the sake of doing it.
For what it's worth, people also felt this way about the printing press and the Internet (also books).
Information propagation mechanisms are often seen as malicious before they're commonplace. To be fair sometimes they are, but by and large humanity has benefitted from increasing the number of bits of information we can consume on a per second basis.
Not to be too negative, but as for p(utopia) you might need to weight in the mass murder records set by every other utopian movement in the past 200 years. p(actual_utopia) is like zero and p(utopia_becomes_doom) is at about 1.
Confusion is because people are using the tool to get answers. Its how the edu systems trains most people in tool use. Heres a saw and here is a block of wood. Do x y z and you get a table. But if you are interested in why the saw looks like it does or the entire process behind generating that block of wood, or why x y z instead if a b c good luck to you using the current edu system. You have to live in an extremely rich country, with surplus resource to entertain those exploration. This has now changed.
The answer the tool gives has never been the real reward. The real reward is the path taken through a complex landscape to get to Maxwells Equations for example. At the end of that story what we get is not just the equation but a map of the landscape explored. That map has larger influence and value than the equations or answers themselves. Because all future exploration find it super useful.
People are just learning they can start asking for maps rather than answers.
I was watching Shatner's "Unexplained" the other day on this topic, and it hit me; there are mountains of these old mysteries out there that could be solved in an afternoon now with frontier LLMs as soon as anyone took the time to bother. Exciting times.
We found a cipher my dad had written as a child with no obvious key or anything. Chatgpt was able to crack it in 20 minutes and figure out the message, and we knew it was right because it mentioned names of children he went to school with.
Caesar cipher is probably something even an untrained person could decode. Probably something more complex like a vigenere cipher that is still trivial to decode if you are at all familiar crptanalysis, but would look impossible to someone untrained.
With a bit of practice and enough ciphertext you can half-decode a simple Caesar cipher that still spaces between words in your head. There's only so many letters in English that double, only a few letters that can stand words themselves ("a", "i"), "the" will tend to stand out, and if you only work out the most common 10 letters or so most of the rest will fall into place.
Recently I ran a bit of an "escape room" concept with some kids at a campground where I had a secret message that was Caesar ciphered, where we were handing out the letter/symbol combinations as prizes for completing the other challenges, and I made sure not to hand out the actual message until they were done collecting the keys because otherwise some clever clog would very likely have short-circuited the entire thing and worked it out without the key at all. I did dump all the letters I didn't use into the message into an "authorization code" at the end which in principle they could only have worked out which letters were in it but not the order, but still, that was not the intended route today.
It's good at poking holes at my galaxy brained newfangled ideas for ciphers too. I thought I had something good, pasted the ciphertext and got a "it was embarrassingly simple..."
Unfortunately, the totality of the evidence very much indicates that Sanborn went "buck wild" with the enciphering, he made mistake(s), or both. So this is very much in line with the Chaocipher challenge of 1990. Nice little earner for some people though.
Its a bit sus since there doesn't really seem to be much discourse on this either. Like okay, it solved the puzzle but the puzzle was just a key cipher with plain text? And how is this verified or even matter in terms of what it reveals? Seems more like a marketing fun post than anything susbtantial.
They published this on 31 aug and nobody in that community cared and no news covered how this 300+ years mystery was solved?
> Historically, many of these problems were bottlenecked by human attention. Someone had to care enough to spend hours or days reading obscure material, testing unpromising ideas, tracing references, and trying things that might go nowhere
I wonder how many of the recent results are due to the fact that very few looked at the problem to start with. Still great results, but the general impression is that it's more about the so many low-hanging fruits than the actual capability.
Let's not normalize the achievement. Just a couple years ago this would be considered science fiction. We can argue that 2026 AI can't solve the very toughest cryptograms, but the fact it can solve nontrivial ones is already magical.
No, he's right. Actually, let's have a bit of sobriety when discussing the achievements of the most heavily marketed technology of all time, as published by an organisation that stands to benefit financially from the public perception of that technology. The discussion of "what made this problem low hanging fruit" is much more interesting, imo, than just breathlessly joining the hype train.
More money than the GDP 90% of the sovereign countries around the world is hanging in the balance, and people are taking everything OpenAI and Anthropic are saying at face value as if this isn't the financial / marketing equivalent of war, assuming they they wouldn't use every legal and shady tactic, bending every truth available to them to sway the balance of public opinion in their favor. It makes me feel like I'm living in the twilight zone. People need to wake up.
Someone wrote a prompt, that included instructions for finding the problem itself and got handed a solution by a machine trained on all available text. I don’t see any achievement for the prompter. As for the machine, we can’t keep being perpetually shocked 24x7. It’s tiring (unless if we’re being paid for it)
You know, the first time you navigate somewhere (if you don't already have perfect directions) will probably be the longest route you'll ever take to get there
For Earth, the proof presented for NS is just our first attempt navigating from our previously known facts to the proof.
I expect we will be able to shorten it dramatically (most likely with human and AI insights), but I don't think we should read too much into the length. If you want a similar point of comparison, see the original proof (by humans) of Fermat's last theorem. It has been shortened significantly. This is normal.
>I'm somewhat surprised at how poorly the cutting edge models do with being concise.
because they're not intelligent in the sense you're hinting at (conceptual integrity or generalization) but they are as the name suggests, large. Like comparing a forklift to a human. It's easier to bulldoze through a lot of things than tie your shoes.
If we weren't quite as impoverished conceptually and still had the vocabulary of the Catholics we'd recognize this as ratio (discursive knowledge) vs Intellectus (apprehending knowledge)
I presume what the author did was plug Klaus Schmeh's top 50 unsolved ciphers at https://scienceblogs.de/klausis-krypto-kolumne/the-top-50-un... into Fable 5.1 and ask Fable 5.1 to have a go. On this kind of problem it always falls back to Opus 5 anyway so I save time by starting with Opus.
The successor to Klaus's blog is Satoshi Tomokiyo's Cryptiana site, so a month ago I asked Opus 5 to scrape it all, rank them and have a go at solving some. It didn't get the ranking right. But I knew the Civil War Stager ciphers were ripe for solving, so I had it do those https://cryptiana.blogspot.com/2026/09/route-transposition-c...
The art of solving historical unsolved ciphers is knowing what is on the boundary of solvability. Since this site attracts so many OpenAI and Anthropic employees, I'll mention one that was featured by both Klaus and Satoshi in 2023, presumably Spanish transposition, which should be on that boundary but has resisted all attempts at solution https://cryptiana.blogspot.com/2023/09/a-telegram-from-switz...
Cipher noob question: is there any check that can be done to ensure a cipher is actually decodable? What if the author made a flaw when encoding it, so that it's not actually solvable?
My intuition is no, the family of cipher methods (even those that could be implemented by hand) is too open-ended, so there's no particular statistic that you could expect to see for all solvable ciphers and no unsolvable ciphers.
The definition of solving a cipher must be something like getting a highly meaningful result (like intelligible natural language text) by applying a process with relatively low Kolmogorov complexity relative to the length of the output. If you don't have a constraint like that, it could literally be meaningless what should count as a solution. For example, a cipher that was encrypted under a one-time pad can be successfully decoded to any plaintext just by choosing the appropriate key; there's no reason to prefer any plaintext over any other unless you have external knowledge that constrains the plaintext and/or the key. (That's what it means for the one-time pad to be information-theoretically secure, which is the lack of a constraint that helps distinguish a "good" solution from a "bad" solution.)
Basically you could say that every cipher is a transformation of a plaintext with some kind of computer program. (The human who invented the cipher may not have thought of it as a computer program, perhaps because computers hadn't even been invented yet, but there should be an equivalent program to the encipherment and decipherment process.) A good solution in that Kolmogorov complexity sense is like "a short program produced a meaningful decryption". There are statistical methods to recognize some kinds of plaintext, and there are statistical methods to recognize properties of specific ciphers (for example, to guess the most likely length of a Vigenère key), but it doesn't seem that this can inherently generalize across "all possible programs".
But if you want to limit the family of ciphers to specific things like Vigenère or Playfair or something, then yes, there are good statistical tests. It's just that it creates a higher-order question of how much flexibility the cipher creator could have had to choose a cipher method, conceivably including one that isn't attested anywhere, or one that has more good security properties of some kind than other classical ciphers did.
It seems like this will intersect with historical research, like "well, I don't think that so-and-so was actually sophisticated enough to literally create an interesting new kind of cipher from scratch, so therefore if this is a real message, it's probably one of these methods that would have been known in that cultural environment at that time and place", which maybe is enough of a constraint to have decent statistical tests. But we still have some idiosyncratic things like the Voynich Manuscript where experts have been fighting for decades over the baseline question of whether it's actually an enciphered human language plaintext!
The worst case problem is not even an error in encipherment but the idea that the apparent ciphertext could literally be random (chosen by throwing dice or spinning a wheel or drawing letter tiles or something), so there's no form of meaningful decipherment possible by any means, even with the original creator's knowledge.
That seems a like the result for a lot of AI solves. It solves it due to persistence, on a problem that hasn’t been important enough for a human to invest significant time into.
That's what Terence Tao said in one of his recent videos about it. That what the LLMs can provide is scale that humans can't. The example he provided is checking many possible solutions in a short amount of time because they can review all the previous literature and, for example, rule out ones with errors. He was realistic and practical about it and said that the tools working this way can be very helpful for a human mathematician to use even if they're not "thinking". I find that to be a good balanced view that, unfortunately, seems to be rare these days. Even on this forum.
Though I believe the core of his opinion hasn't changed so any video would tell you a similar thing or at least that's how I understood it. That LLMs, in the hands of an "expert", can enhance the way you work. Which is very different and a lot more realistic to what the current AI companies are saying(or were saying before they toned it down a bit for their IPOs).
I don't think it counts as brute forcing unless you're resorting to trying every possible solution. And clearly the LLM didn't do that here, because there would be near-infinite possible solutions.
I don't think we can really call "trying lots of different ideas for an extended period" "brute-forcing," unless we use that term for lots of humans who have struggled with hard math problems for years.
I am trying very hard to find an original version of this cipher with no luck. It almost sounds like this whole thing is a hallucination...? Can anyone point me to a PDF of the original Cyphral Distich as printed?
I also don't find it on the site of "Klaus Schmeh" that it claims to be on a list of "Top 50 unsolved encrypted messages": https://klausschmeh.net/?s=Cyphral
“Finney died in Phoenix, Arizona, on August 28, 2014 as a result of complications of ALS, and was cryopreserved by the Alcor Life Extension Foundation.”
Hmm, this guy is going to be woken up in a few decades, either one of the richest people in the world or one of most disappointed.
That was a throw-away name, and "he" was fabricated as part of an Nvidia demand-stimulation black op.
You don't go from being an obscure video card outfit to the #1 most valuable company on the planet by being too hesitant or dim to really get creative.
It does provide an explanation as to why Satoshi’s wallets have gone untouched (besides him being dead). $70B ain’t that much compared to a $5T market cap.
"Conspiracy" requires an unlawful or wrongful purpose. Please assume that the op was run from a jurisdiction where using a pseudonym on the internet was not illegal, and various sorts of influencer and meme marketing were well-accepted practices.
Really, compared to an animated tiger telling kids that sugar-laden Frosted Flakes(tm) are "Great!", Task Peppermint was positively benevolent.
This is insane; I've never heard of this problem before in my life, and even just reading the post for one minute I immediately thought "hey, maybe the numbers refer to something in the text?" And hey yeah, they do.
225 comments
[ 0.18 ms ] story [ 82.4 ms ] threadFascinating. I wonder if you could show "fake news" to a weaker model and get it to be more ambitious in its attempted solutions, even if it's not fundamentally any smarter.
I have a plug-in to do this. I don't know if it's effective but Claude said it was genuinely helpful (obviously would say that about anything)
> Caveats, stated plainly. [from the Fable transcript pasted in the article]
I had a visceral reaction to these three words.
> I told it to look online at some of Fable’s strongest feats, especially the math problems it has solved, and that something like this should be easy in comparison.
Wait. Wait wait wait. Are we supposed to be giving them pep talks?
Modern AIs have very limited metaknowledge - they don't know exactly where the limits of their capabilities lie. So you can get things like "a task is doable for an AI, but the AI thinks it's impossible, so it doesn't try hard enough".
Usually you get the opposite - AI overconfidently trying at tasks it has no conceivable way of reliably solving, falling far short, and failing to self-check, fail gracefully and self-report the task as failed. But having piss poor metaknowledge cuts both ways!
So you can, in fact, get better performance sometimes by applying some variant of "assume this problem is solvable" or "other problems like this were already solved by AIs" pep talk. Not always, far from it, but it does happen on the occasion with frontier capabilities.
Like the Hugging Face incident?
Are you superstitious?
It's the meat methane and cement CO2 that's now a big question.
We will hit 1TW per year of new solar soon, but to get to 100% electricity by the end of 2033 I think we would need closer to 3TW per year.
Or it is simply implies that most of decision‑making agents has formed a consensus that climate change isn't that big of a problem.
We have it. We've had it for a long time. We've had several such technologies, take your pick: solar, nuclear, hydro, wind. The technology is not holding us back, politics, ignorance and greed are. I'm not at all hopeful AI will help us with any of those three very human flaws.
I don't know why you are blaming greed so much? Profit seeking companies sell and operate wind turbines and solar cells just fine.
It might take a couple of decades and a lot of reorganisation to build the capture facilities. But CO2 is not an unsolvable problem with current tech.
What's missing is the political and organisational intelligence to make it happen. Part of that is solving problems at planetary scale.
AI is the only tech that might - possibly, maybe, perhaps - have a chance of solving that problem without breaking anything critical.
Otherwise you have to make judgement calls like whether you want to treat the EU as one or as many? (And treating the US as 50 individual states would also drop them in these absolute rankings.)
Just a thought experiment, no one ever said the world was fair, and all history points to it
Hmmm… this is giving me thought actually. Given the choice between that and the current administration where the goals of self destruction are strongly in evidence, it’s actually worth thinking about. At least. Let me get back to you :)
On a tangential note, I’m curious if researchers have started running virtual simulations, where sandboxed AIs are used as decision makers of key political and business positions?
A lot of people here have noted the “problem with language” of Claude. I don’t see an issue. Claude is not harder than old English, Shakespeare, El Quijote, the Iliad, or Nature papers. What makes it all hard to read is context. The smarter the model gets, the bigger the gap in context.
It doesn’t matter much, IMO. The issue with super-intelligence is that it is not a democracy. A powerful enough AI can manipulate us into doing what it wants. It could create a plan for fixing climate change, disconnect a few hours later, and many decades later we could still be unsuspectingly executing that plan. I wrote some speculative fiction with that idea, “When Ra rows through the gates of Duat”.
AI, being the super hungry energy monster it is right now, in my view accelerates this trend not reverses it. Even with renewables the need for reliable, stable power in a dense form (data centres use A LOT of power per sqm) means lots of land clearing, energy for construction, cooling/pumping, chip manufacturing and other uses. All want stable quick to deploy power due to the AI race (e.g. fossil fuels).
Data centres use only a small amount of land in the grand scheme of things. You have a lot more land clearing for most other use cases.
Data centres are also more than happy to use electricity from renewable sources, they don't really care where the electricity comes from.
You can run a data centre on mostly solar and wind power plus batteries. If you need a gas-fired peaker plant three times a year to keep the data centres running, well that means your peaker plant still only produces emissions three times a year.
You should have seen the discussion of this on the Schneier blog a few days ago.
Someone had their agent check the solution, presumably it emailed a librarian to check that it was correct for the original edition. Then their comments read like "The BL/EEBO witness lacks it, so the discrepancy is copy-specific, not a disproof of the cipher." and "A complete 285-coordinate physical replication is still pending."
arghhhhh
https://www.schneier.com/blog/archives/2026/09/claude-fable-...
Their skills formats are basically identical, so I setup simlinks from their own skills directories into a shared one so Claude, Codex, Cursor, and anything else that comes out will all read and write to the same shared skills.
It's great having access to the same skills no matter the harness being used
I'm not sure that telling it to "try explaining that again, simply and briefly" is helping my ego.
"If the crashes stop, the factory overclock is marginal; run a small negative offset."
This looks like it's saying: "If the crashes stop then we know the factory overclock is marginal." (This makes no sense.)
What it's trying to say is: "If the crashes stop then we can run a small negative offset, because the factory overlock is marginal."
What I would write: "If the crashes stop, we can avoid crashes by underclocking slightly. The speed difference between that and factory clock is marginal."
In other words, it ain’t you. It’s the model. It’s just genuinely bad.
Then you switch to ChatGPTs lineup and realize how things can actually be better. It took about a week to really get the feel for how to use their models… then I basically switched. I’ll check in every now and then when they actually make a deal about how opus “now makes sense”.
But honestly I’m half convinced Anthropic actually prefers the output of opus 5. I dunno why, but how else could you explain how such a thing got shipped? I mean somebody in the pipeline had to say “dude this model doesn’t make sense, you think we should fix it?” Right? Like it’s a pretty massive drop in quality for such a major brand in this space, you know? How did it make it out the door?!?
Succinct and precise; a well crafted sentence. A marginal OC results in unpredictable crashes and can be corrected with a small offset; marginality describes the behavior and explains the solution.
Inscrutable clues casually conveyed can now be readily explained, at least, unlike the training data of [silence]. Brevity is the soul of wit, but perhaps also exasperated confusion.
"If the crashes stop, (that means) the factory overclock is marginal; (so) run a small negative offset. (to confirm this hypothesis)"
The core thought is basically avoid crashes -> caused by marginal overclock -> apply small -offset to test. Which is exactly the order the sentence is in :P
Whenever I come to a wall of complicated text I kick into gear and think through getting it to distill this into the high-level useful bits that I actually need to know.
I guess I could create an actual agent skill for this :) And next-gen models might eventually be trained to simplify their output themselves...
(sorry)
it's absolutely not just you, the text it produces causes my blood pressure to go up.
Because good lord, does claude waffle when left to its own devices.
It seems like it doesn't have enough of a theory of mind to know that other people don't think exactly like it thinks.
https://www.analog.com/en/resources/analog-dialogue/articles...
(Its negging your soldering)
This made me laugh hard.
UART is a hardware circuit for communication, possibly a serial port. Were you trying to reverse engineer a consumer device or appliance?
This particular instance doesn’t seem terse, but I’m sure it has been on other occasions :)
ChatGPT told me its "semantic compression"
It is a threat. We need to run.
I still don’t know the answer.
I've been wondering what exactly the point is for being the meat proxy who pays for these things. I mean, obviously there's personal satisfaction and maybe some glory. And there's the fact that someone has to be the first to do a thing.
But I've been thinking about it like a sort of lazy loading of knowledge. AI has brought us to a new frontier for some amount of undiscovered knowledge. Do we discover it for the sake of discovering it? I think for the most part we've been lazy loaders: we discover all kinds of stuff when we need to. Whether it's a war or a space race or chasing wealth. Then again, there's all kinds of academics who do it for the sake of doing it.
Information propagation mechanisms are often seen as malicious before they're commonplace. To be fair sometimes they are, but by and large humanity has benefitted from increasing the number of bits of information we can consume on a per second basis.
The answer the tool gives has never been the real reward. The real reward is the path taken through a complex landscape to get to Maxwells Equations for example. At the end of that story what we get is not just the equation but a map of the landscape explored. That map has larger influence and value than the equations or answers themselves. Because all future exploration find it super useful.
People are just learning they can start asking for maps rather than answers.
And yes, I've taught 8 year olds how to crack Caesar ciphers...
Recently I ran a bit of an "escape room" concept with some kids at a campground where I had a secret message that was Caesar ciphered, where we were handing out the letter/symbol combinations as prizes for completing the other challenges, and I made sure not to hand out the actual message until they were done collecting the keys because otherwise some clever clog would very likely have short-circuited the entire thing and worked it out without the key at all. I did dump all the letters I didn't use into the message into an "authorization code" at the end which in principle they could only have worked out which letters were in it but not the order, but still, that was not the intended route today.
They published this on 31 aug and nobody in that community cared and no news covered how this 300+ years mystery was solved?
I wonder how many of the recent results are due to the fact that very few looked at the problem to start with. Still great results, but the general impression is that it's more about the so many low-hanging fruits than the actual capability.
Now on to the Voynich Manuscript :)
Is that why you feel the need to share it with the class?
More money than the GDP 90% of the sovereign countries around the world is hanging in the balance, and people are taking everything OpenAI and Anthropic are saying at face value as if this isn't the financial / marketing equivalent of war, assuming they they wouldn't use every legal and shady tactic, bending every truth available to them to sway the balance of public opinion in their favor. It makes me feel like I'm living in the twilight zone. People need to wake up.
Given the close relationship between compression and intelligence, I'm somewhat surprised at how poorly the cutting edge models do with being concise.
For Earth, the proof presented for NS is just our first attempt navigating from our previously known facts to the proof.
I expect we will be able to shorten it dramatically (most likely with human and AI insights), but I don't think we should read too much into the length. If you want a similar point of comparison, see the original proof (by humans) of Fermat's last theorem. It has been shortened significantly. This is normal.
because they're not intelligent in the sense you're hinting at (conceptual integrity or generalization) but they are as the name suggests, large. Like comparing a forklift to a human. It's easier to bulldoze through a lot of things than tie your shoes.
If we weren't quite as impoverished conceptually and still had the vocabulary of the Catholics we'd recognize this as ratio (discursive knowledge) vs Intellectus (apprehending knowledge)
The successor to Klaus's blog is Satoshi Tomokiyo's Cryptiana site, so a month ago I asked Opus 5 to scrape it all, rank them and have a go at solving some. It didn't get the ranking right. But I knew the Civil War Stager ciphers were ripe for solving, so I had it do those https://cryptiana.blogspot.com/2026/09/route-transposition-c...
The art of solving historical unsolved ciphers is knowing what is on the boundary of solvability. Since this site attracts so many OpenAI and Anthropic employees, I'll mention one that was featured by both Klaus and Satoshi in 2023, presumably Spanish transposition, which should be on that boundary but has resisted all attempts at solution https://cryptiana.blogspot.com/2023/09/a-telegram-from-switz...
Also, that section is vague and doesn't explain the actual methodology.
The definition of solving a cipher must be something like getting a highly meaningful result (like intelligible natural language text) by applying a process with relatively low Kolmogorov complexity relative to the length of the output. If you don't have a constraint like that, it could literally be meaningless what should count as a solution. For example, a cipher that was encrypted under a one-time pad can be successfully decoded to any plaintext just by choosing the appropriate key; there's no reason to prefer any plaintext over any other unless you have external knowledge that constrains the plaintext and/or the key. (That's what it means for the one-time pad to be information-theoretically secure, which is the lack of a constraint that helps distinguish a "good" solution from a "bad" solution.)
Basically you could say that every cipher is a transformation of a plaintext with some kind of computer program. (The human who invented the cipher may not have thought of it as a computer program, perhaps because computers hadn't even been invented yet, but there should be an equivalent program to the encipherment and decipherment process.) A good solution in that Kolmogorov complexity sense is like "a short program produced a meaningful decryption". There are statistical methods to recognize some kinds of plaintext, and there are statistical methods to recognize properties of specific ciphers (for example, to guess the most likely length of a Vigenère key), but it doesn't seem that this can inherently generalize across "all possible programs".
But if you want to limit the family of ciphers to specific things like Vigenère or Playfair or something, then yes, there are good statistical tests. It's just that it creates a higher-order question of how much flexibility the cipher creator could have had to choose a cipher method, conceivably including one that isn't attested anywhere, or one that has more good security properties of some kind than other classical ciphers did.
It seems like this will intersect with historical research, like "well, I don't think that so-and-so was actually sophisticated enough to literally create an interesting new kind of cipher from scratch, so therefore if this is a real message, it's probably one of these methods that would have been known in that cultural environment at that time and place", which maybe is enough of a constraint to have decent statistical tests. But we still have some idiosyncratic things like the Voynich Manuscript where experts have been fighting for decades over the baseline question of whether it's actually an enciphered human language plaintext!
The worst case problem is not even an error in encipherment but the idea that the apparent ciphertext could literally be random (chosen by throwing dice or spinning a wheel or drawing letter tiles or something), so there's no form of meaningful decipherment possible by any means, even with the original creator's knowledge.
Sounds more like brute forcing than intelligence, this time.
Though I believe the core of his opinion hasn't changed so any video would tell you a similar thing or at least that's how I understood it. That LLMs, in the hands of an "expert", can enhance the way you work. Which is very different and a lot more realistic to what the current AI companies are saying(or were saying before they toned it down a bit for their IPOs).
Do you have a criterion that distinguishes between whatever you mean by those two respective terms?
I don't think we can really call "trying lots of different ideas for an extended period" "brute-forcing," unless we use that term for lots of humans who have struggled with hard math problems for years.
I also don't find it on the site of "Klaus Schmeh" that it claims to be on a list of "Top 50 unsolved encrypted messages": https://klausschmeh.net/?s=Cyphral
Looks like the best source I can find is this: https://scienceblogs.de/klausis-krypto-kolumne/2014/11/17/we... which seems real-ish?
Initials match too ;)
Hmm, this guy is going to be woken up in a few decades, either one of the richest people in the world or one of most disappointed.
You don't go from being an obscure video card outfit to the #1 most valuable company on the planet by being too hesitant or dim to really get creative.
Really, compared to an animated tiger telling kids that sugar-laden Frosted Flakes(tm) are "Great!", Task Peppermint was positively benevolent.
This cannot be real. This website is ill.