5 comments

[ 2.9 ms ] story [ 14.3 ms ] thread
This is already happening and it absolutely does let llms understand the page better. Search for JSON-LD
JSON-LD already covers a chunk of that list inside the HTML, so a separate AI-friendly JSON file helps mostly when the page itself is too noisy to parse (don't tell me )...

What still sits underneath is earlier than extraction:: can the crawler fetch the main HTML at all, and is there readable text once the JS shell is gone? "A twin behind a link tag" does not help if the UA never gets a readable document.

BUT there's IsReady.AI that help you creating AI readable content of your website. Plus consider firecrawl.dev!