Show HN: EterDB, a Postgres fork that makes it easy to recover from incidents (eterdb.com)
Hey everyone, a few months ago, after recovering from a Clade-generated bug, I was thinking to myself “wouldn’t it be nice if prod DB writes were easy to roll back”. I’ve put something together: it’s a two-container deploy plus a CLI tool. It’s far from being production ready, so please be gentle. All feedback welcome!
The nitty gritty: https://eterdb.com/tech
GitHub: https://github.com/eterdb/eterdb
21 comments
[ 0.34 ms ] story [ 40.2 ms ] threadHmmmmmm......
1. "Claude-generated bug". No it was PBCAK (Problem Between Chair And Keyboard) a.k.a "foolish person ran Claude against the production database without testing it elsewhere". There, fixed it for you.
2. This "product" is solving a problem that is already solved. You can for example use a SaaS provider such as Aiven[1] who will provide you with PITR (Point-In-Time Recovery) point and click solutions. Alternatively there is more than one piece of Postgres backup software that lets you do the same on a DIY basis.
3. "Out of the box" you have pg_dump. You could have just done a simple pg_dump before letting Claude loose on your database.
[1] https://aiven.io/
The point remains that if they ran Claude against a test database they would have found the bug without killing their production database, and therefore also not need to come up with an over-engineered "solution".
Sometimes also the less over-engineered the better. Stuff like PITR and pg_dump is battle-tested and easy to reason about.
1. This is such an insane example, I don't get why would you give agents write access to your production db in the first place. Are people really doing this? I don't even have production connection urls on my laptop. Any manual statements executed against the db must be treated as a war-room situation with at least another engineer reviewing your SQL before you execute it.
2. Dropping the table could have easily caused writes to fail. Most likely there is no way to recover these writes (especially if it's from user requests), so it could have lead to loss of data. Reversing the table drop doesn't fix this issue.
Also reversing a table drop doesn't solve extra corruption routes such as cross-table dependencies.
One extra edge case to take care of is non standard ports used for PostgreSQL. Rather than trying to make our solution work with them, it would be more prudent to segfault on any attempts of PostgreSQL processes to bind to non-5432 ports though (fail early)
It was deprecated as the performance hurdle was significant.
This seems to utilize logical WAL replication and some form of dependency generation based on proxying a transaction's reads.
Managed Neon (Acquired by Databricks) also offers time travel [2].
One of the biggest challenges with this kind of solutions is handling ACID, OCC and/or HW failures correctly, Sometimes it requires an whole new language! [3]
[1] https://www.postgresql.org/docs/6.3/c0503.htm
[2] https://neon.com/docs/postgres/backup-restore/time-travel-as...
[3] https://apple.github.io/foundationdb/flow.html
Could you write that by hand? It would make it a lot easier to understand. I’d also recommend removing the various mentions to previous versions, or at least shifting them to footnotes. It’s confusing to be reading an implementation and then finding out later that it isn’t the approach that was taken.