27 comments

[ 2.7 ms ] story [ 36.5 ms ] thread
Always enjoyed his lecture series from CMU, hopefully those continue in a sponsored format from Clickhouse.
Clickhouse Labs is a research lab for ClickHouse’s engineers and their db product?

Whatever floats your boat. Sounds like you just work as an engineer at a db company

> The goal of ClickHouse Labs is to establish a best-in-class industry research organization focused on databases. It will not operate as an isolated research organization that throws ideas over the wall to engineering. Instead, we will work closely with ClickHouse engineers, customers, collaborators, and industry partners to develop and disseminate new ideas that keep ClickHouse at the bleeding edge.

This is cool, though bittersweet that the public research infrastructure (universities) is not really configured to support this kind of high-impact research any more.

That's great news for Clickhouse and for Andy I guess, they are both obsessed with databases and are at the cutting edge.
Congrats, that's incredible news. Watched apavlo@'s lectures while studying in the university and finished my bachelor thesis implementing features and doing research at ClickHouse. Surprised to see these worlds being together now!
Hi Andy, that’s very cool news. Will you be at the HPTS workshop in October?
What has been your experience with HPTS? It will be my first time presenting.
(comment deleted)
(comment deleted)
> ClickHouse had features that at the time were only found in a handful of closed-source, commercial analytical DBMSs. For example, ClickHouse was written in C++ and supported vectorized query execution using SIMD in 2016. Most prominent open-source analytical DBMSs in 2016 were JVM-based and did not support SIMD optimizations until years later.

Performance is a feature. "Written in C++" is a strange idea of a feature.

Clickhouse just became the hottest talent-attraction on the market.

Congrats Andy, hope you enjoy the ride =)

I'm very curious about the convergence of the best in class fast OLAP products (StarRocks, ClickHouse) with Trino. It sounds like everybody is going for decoupled compute/storage, using S3 or similar as the storage layer, and thus forgoing colocated joins (ok I know ClickHouse joins suck)...

So what does this mean for ingestion (and indexing)? Iceberg V3? Paimon? Bespoke ingestion through the DB engine to do the indexing?

Opinionated take: ingestion should be treated as a streaming reorganization workload, separate from whatever a "database" is.

You do not even need to change the Iceberg spec, although you could. Just rejig things a bit to produce better statistics (think liquid clustering but on the write path). You still write the same dataset queried the same way, but automatically get better query performance. Workload owners should be able to decide the tier at which this streaming reorganization happens: they know their data staging and bandwidth constraints the best.

Hope it goes well for you Andy. Have loved your lectures over the years.

Best of luck!

Looks like Andy's here, so if you see this - please also try to convince ClickHouse to consider funding DB research in academia. With all the money being poured into AI and the chaos in government funding, there is almost nothing for DB research anymore.
Remy - Good to hear from you. Yes things are mess right now. Let me get settled with setting up the lab and we'll figure out how we want to engage with academia.
Yooo is this the CMU funny guy?
Andy taught me a lot about trolling. I also hear he's banned from every post office within 50 miles of Baltimore.
Its refreshing to see corporate research labs in a non AI area. Clickhouse has been a huge beneficiary of the AI wave and its good to see some of the value going into advancing fundamental research in infrastructure
Andy, go ahead and correct the https://dbdb.io/db/bigquery How did you come to the conclusion that BigQuery's data model is "Document / XML" :joy:

This entry is not only wrong, it's way outdated, therefore misinforming. I'm wondering about the overall quality of this dbdb.io database at large.

There are several out of date entries in there. We used to have students help with writing entries. I just did a major refresh of the site this summer. I need to go through all the entries and figure out how to update them.
Andy, you had me at KillerMike + El-P super group. Best of luck running the jewels on the dbase world with CH