Show HN: 500 TB internet index in ClickHouse, with congestion pricing (scry.io)
There is a long history of people trying to do very fancy things that end up being done in relational databases and a little SQL. There is a gravity to them, a bitter lesson, just like scaling of generalized ml training methods. I mean many, many information products can be built off essentially giant real-time OLAP databases and frontier LLMs writing brilliant SQL+Datalog+vector+Jev etc. queries.
Google Search, Tavily, Exa essentially have the problem of mapping your agents' context you are willing to provide, to a tiny subset of their index. You pay a fixed cost to an extremely hard problem that has a distribution of hardness, which means YOU eat the downsides when they are running out of budgeted compute to help you out.
Their algorithms are opaque to the caller, there's really not much user control, and there's not a serious opportunity to communally improve search recipes, like the lexical+Jev recipes you trust to select bleeding edge AI builders.
Furthermore, search companies aren't even pursuing text-to-SQL anymore (several have talked to me)... they made up their minds during the traumatic 2024 text-to-sql days. They were just too early.
Meet Scry, where I have a 500 TB NVMe internet index (I'm doing my best indexing and normalizing all the intelligence explosion alpha) that you can run ~arbitrary readonly SQL and some of Datalog over, and I handle the problem of resource-contention with congestion-based micro-auction pricing. When there's capacity, the service is free for non-commercial use.
I hope you enjoy. I'm intent on scaling this paradigm on differentiated hardware over much more data, so any compelling use cases or queries I could show off, would be much appreciated!
29 comments
[ 0.21 ms ] story [ 44.5 ms ] threadAlso HN: This site is doesn't look like other sites, awful to read.
Please, just explain in your own words, how the pricing works, without using made up terms invented by an LLM.
As load increases, the price goes up. There's also exponential egress pricing because Scry is a place where computation happens, not a metaphorical torrent. I'm still tweaking things, balancing between multi-user server load, giving normal active human users lots of priority and no fear in using it hard and deliberately, and fair pricing for agents and automated heavy workflows people set up in the background.
I suppose the scraping you're doing is a huge part of your value proposition, but I would like to gently nudge you in the direction of making the datasets available via p2p (e.g. a torrent) like how Wikipedia distributes its snapshots in the spirit of democratizing access to data that is becoming increasingly walled off. Also, I think another potential benefit that kind of bulk sharing would have is relieving the congestion from those doing the equivalent operation to extract data via the querying interface.
Can you expand on this? Is text-to-sql a deadend? (I personally think it is, having worked on a project at $work. But curious to hear about others experience.)
Will it be added? Are you forced to use a traditional search engine to make it in the leaderboard?
https://scry.io/deepsearchqa
https://www.kaggle.com/benchmarks/google/dsqa