There is no problem with YAML, just like there was no problem with XML at all. The problem is with the software devs who use it for the wrong thing. I just wonder why JSON doesn't get all the hate, while it is a terrible format also: it is not streamable (YAML is), it doesn't have comments... and non-standard workarounds are used for these features. Norway better than YAML, in my opinion.
There are plenty of problems with YAML as the article shows. The biggest problem with YAML is, that serializer and deserializer need to agree on rules outside of the spec, because otherwise they can be both compliant and still don't understand each other.
I also wonder why it’s the 2012 version of the language who gets all the hate, instead of the hate being directed at Google, who doesn’t care about putting resources towards updating the main offender, Kubernetes’ YAML 1.1 parser. https://github.com/kubernetes/kubernetes/issues/34146#issuec...
Would totally love to do that bro, is not a resource or finance problem- not even talent - AI could probably do it with oversight .. but No Promo.. cant touch it once another guy declared it complete.. that s them rules.
I wrote a C++ library at work that, among other things, serializes/deserializes to JSON, CBOR and BSON. I don't get that point.
The first two formats I've implemented as fully streaming, the serializers emit into an output iterator one character/byte at a time and the deserializers only need to hold one value at a time (mostly for strings, and not even for CBOR's byte arrays, which are streamed). BSON requires full document instantiation because of byte offsets and is genuinely anti-streaming in its design.
I happen to have implemented a serializer for YAML, fully streaming, mostly as a compact-ish human-readable debug data representation. I can scarcely think of fates worse than having to write a spec-compliant YAML parser.
> YAML is (or at least, is declared by its maintainers to be) a superset of JSON, making any JSON document also a valid YAML document.
> So surely if YAML is streamable, JSON must be too?
Given YAML is not a subset, but a superset of JSON, as you correctly noted, I'm not sure how you reached that conclusion.
Offtopic note: I wanted to dig the wiki superset link up the lazy way, so searched for superset, and all I got was endless list of marketing trash links. Even on DDG. RIP, we are on the Dead Internet.
The "streamed JSON" needs --- line sparators, which is not a JSON spec feature, thus that data stream is valid YAML from that point, bot not a valid JSON stream.
How was my comment snarky? The YAML maintainers may have intended YAML to be a superset of JSON, but failed by not incorporating some obscure feature of it that bore on whether it was streamable or not.
There numerous problems with YAML. It is underdefined, ambiguous, fragile, and is just impossible to write a portable parser for.
JSON is naive, lacks almost everything. But is supereasy to parse and read.
As funny as it sounds, XML is overspecified. Parsing is possible and predictable but supporting the full standard (all of them) is so much work that nobody does it anyway.
So, given the choice above, i'd go for naive simplicity.
XML being overspec'd is good, IMO, because we get things like xpath, XLSX, schemas, etc very easily. It would be fine if people just use libraries for these. I mean, I certainly wouldn't roll a JSON parser by hand either, even if it is easier.
XML just looks nasty, but IMO not worse than JSON. JSON is not easy to read, because brackets suck for deep nesting. XML nesting is 100x easier to read, because the closing tags tell you exactly what is closing.
> I just wonder why JSON doesn't get all the hate, while it is a terrible format also: […] it doesn't have comments […]
The lack of comments in JSON is not a bug; comments in JSON were an anti-feature. Douglas Crockford intentionally removed them to prevent developers from using comments to store parsing directives (which he saw people doing in earlier revisions).
I think they meant "no way," but maybe it's a joke about the Norway problem? YAML famously confuses Norway's country code ("NO" or "no" without quotes) with unquoted "no" meaning Boolean false.
Next time the author would like to write an article like that, it’s better that they offer a new standard that is as flexible while also doing the work YAML is doing today
I've stumbled upon this page few years back, and was very confused with the language. It took me some time to parse it, and realise that all comments are actually sarcastic. Overall, the general style feels like I'm reading comments on TikTok. It's not that I'm against this style, it's just you don't expect it in a technical article. Am I the only one?
You really should not blame it on YAML that if you put a random hash value in a field that it may be a number. Like, duh. If you want a string, use quotes.
- "something broke" is way better than stack traces with line numbers
That sums it up. Its the same since the days of XML. Someone decides to make something configurable externally and then everything gets put in there for convenience. Now we are programming in XML/YAML without even a stack trace (or a debugger, or call hierarchy, or type checking or ... )
Surprised TOML didn't get mentioned in the alternatives section, it's probably my favourite configuration format.
I'm working on a project now that uses YAML for configuration (it makes more sense in this context than TOML) and am parsing it using Rust's `serde` crate (via `serde_norway`). It's not so bad if you use `#[serde(deny_unknown_fields)]` everywhere.
Modern editors make this a non-issue. I use both Python and YAML fairly regularly, both have significant whitespace, and I don't have any issues with either.
Also, because YAML doesn't involve things like curly braces all over the place, it's actually notably less friction to type rapidly by hand than something like JSON. I don't think YAML is amazing or anything like that but I do find myself nowadays preferring it, for config at any rate, over JSON.
JSON definitely wins for machine readable data sent over a pipe though because the lack of significant whitespace means it'll compress better.
It’s because ppl are not using it correctly, it was designed for configs and now it’s being used as a programming language… a language without types and debugger. It’s frustrating devs because they find the issue at the deployment time, not the compile time.
Personally I did never pick the YAML as a first format for the configs, only my ruby friends did that.
I think KYAML[1] is very relevant here. KYAML is a strict subset of yaml[2], supported natively by kubernetes tooling in most recent versions.
In my opinion it's closer to JSON with comments than to YAML, but overall it looks like a pretty nice format, that solves some of pain points of both technologies. This is not the first nor the only JSON-with-comments format, so I wonder if it manages to break out of kubernetes world and become popular.
My usage of YAML has been limited. For the times I have used it (kubenetes, docker compose, other things, etc) -- I don't get what it so special about it.
It just seems reinventing the wheel again and again.
I am not suggesting JSON is perfect, either. However I do prefer it to YAML (and XML) by a mile.
It's a shame s-expressions never taken off.
(select (id name)
(:from users)
(:where (= name "Peter")))
I would wager that just in the time it takes me to write this comment, millions of yaml documents have been successfully parsed.
Works for me. Until it doesn't. Just like anything else. No silver bullets.
58 comments
[ 3.0 ms ] story [ 5.2 ms ] threadEither you have an infinite spec, or you always have things outside it.
[0] YAML would allows you to tag the type, but nobody does it anyway
Google put forward KYAML recently (https://www.kubernetes.dev/resources/keps/5295/), it looks like a more sane way forward.
I agree that Norway is better than YAML. Let's keep Norway and dump YAML.
> There is no problem with YAML
I am reminded of: "There are no American infidels in Baghdad. Never!"
So surely if YAML is streamable, JSON must be too?
Or the YAML maintainers incorrect?
The json spec just doesnt do anything like that but jsonl does.
The first two formats I've implemented as fully streaming, the serializers emit into an output iterator one character/byte at a time and the deserializers only need to hold one value at a time (mostly for strings, and not even for CBOR's byte arrays, which are streamed). BSON requires full document instantiation because of byte offsets and is genuinely anti-streaming in its design.
I happen to have implemented a serializer for YAML, fully streaming, mostly as a compact-ish human-readable debug data representation. I can scarcely think of fates worse than having to write a spec-compliant YAML parser.
Given YAML is not a subset, but a superset of JSON, as you correctly noted, I'm not sure how you reached that conclusion.
> Or are the YAML maintainers incorrect?
I think you should check this page before making snarky comments: https://en.wikipedia.org/wiki/Superset
See also the comment below: https://news.ycombinator.com/item?id=49476489
Offtopic note: I wanted to dig the wiki superset link up the lazy way, so searched for superset, and all I got was endless list of marketing trash links. Even on DDG. RIP, we are on the Dead Internet.
Given: - YAML support streaming - any JSON is valid YAML
Conclusion - any JSON is streaming
The "streamed JSON" needs --- line sparators, which is not a JSON spec feature, thus that data stream is valid YAML from that point, bot not a valid JSON stream.
NDJSON is a different beast.
See Examples section.
How was my comment snarky? The YAML maintainers may have intended YAML to be a superset of JSON, but failed by not incorporating some obscure feature of it that bore on whether it was streamable or not.
JSON is naive, lacks almost everything. But is supereasy to parse and read.
As funny as it sounds, XML is overspecified. Parsing is possible and predictable but supporting the full standard (all of them) is so much work that nobody does it anyway.
So, given the choice above, i'd go for naive simplicity.
XML just looks nasty, but IMO not worse than JSON. JSON is not easy to read, because brackets suck for deep nesting. XML nesting is 100x easier to read, because the closing tags tell you exactly what is closing.
In xml's they clearly overdid it. all the formats/substandards you listed... there is just not a single library that supports all of the standard.
And because implementations are so different, most users just stick to a single C library. That library, btw, is also incomplete and undermaintained.
So that's the problem with complex standards: they are hard to implement and support. Nobody does this, unless it is a business-critical matter.
The lack of comments in JSON is not a bug; comments in JSON were an anti-feature. Douglas Crockford intentionally removed them to prevent developers from using comments to store parsing directives (which he saw people doing in earlier revisions).
* https://web.archive.org/web/20120507093915/https://plus.goog...
* 2012: https://news.ycombinator.com/item?id=3912149
There are things like https://www.baeldung.com/jackson-streaming-api but maybe there's more in you definition of "streamable" that JSON doesn't satisfy.
You really should not blame it on YAML that if you put a random hash value in a field that it may be a number. Like, duh. If you want a string, use quotes.
You mean other than when we won't be? :-(
That sums it up. Its the same since the days of XML. Someone decides to make something configurable externally and then everything gets put in there for convenience. Now we are programming in XML/YAML without even a stack trace (or a debugger, or call hierarchy, or type checking or ... )
I'm working on a project now that uses YAML for configuration (it makes more sense in this context than TOML) and am parsing it using Rust's `serde` crate (via `serde_norway`). It's not so bad if you use `#[serde(deny_unknown_fields)]` everywhere.
Well most of that is due to 1.1 which has been deprecated god knows for how long.
1.2 does not have the 'Norway' problem any more.
Tooting my own horn, this is how a modern YAML library looks like nowadays:
https://github.com/pantoniou/libfyaml
Oh, what sad times are these when passing ruffians can say Nicaragua at will to old ladies! There is a pestilence upon this land, nothing is sacred!
Modern editors make this a non-issue. I use both Python and YAML fairly regularly, both have significant whitespace, and I don't have any issues with either.
Also, because YAML doesn't involve things like curly braces all over the place, it's actually notably less friction to type rapidly by hand than something like JSON. I don't think YAML is amazing or anything like that but I do find myself nowadays preferring it, for config at any rate, over JSON.
JSON definitely wins for machine readable data sent over a pipe though because the lack of significant whitespace means it'll compress better.
Nobody anywhere needed more INI files.
Care to elaborate? We are using TOML for config files and just do fine.
Personally I did never pick the YAML as a first format for the configs, only my ruby friends did that.
I won't even look at it, and it will be just fine.
I think KYAML[1] is very relevant here. KYAML is a strict subset of yaml[2], supported natively by kubernetes tooling in most recent versions.
In my opinion it's closer to JSON with comments than to YAML, but overall it looks like a pretty nice format, that solves some of pain points of both technologies. This is not the first nor the only JSON-with-comments format, so I wonder if it manages to break out of kubernetes world and become popular.
[1] https://kubernetes.io/blog/2026/08/11/how-to-pretty-print-ku...
[2] https://www.kubernetes.dev/resources/keps/5295/
It just seems reinventing the wheel again and again.
I am not suggesting JSON is perfect, either. However I do prefer it to YAML (and XML) by a mile.
It's a shame s-expressions never taken off.
Or I know - its a sample, but you get it. :-)Is this supposed to read as Norway's YAML?