138 comments

[ 0.21 ms ] story [ 49.7 ms ] thread
New functional query language for PostgreSQL and SQLite by Evan Czaplicki the author of Elm
Exciting news!! Love Elm, can't wait to use it more
So this is capable of turning a one-liner of SQL into six lines of barely readable code?
It seems that is the price you pay for the power to turn a 600-line nightmare SQL query into 60 lines of barely readable code.
I would like to see this example then.

I’m all for improving on SQL, but this syntax does not even solve the dangling comma issue as far as I can tell from the example.

I'd rather take the 600 lines of SQL, provided it's not dynamically constructed. SQL is a very high level language, it's fine IMO.
SQL is a horrible language. I’d gladly program in something composable like Elm.
Unfortunately Evan removed GROUP BY in 0.19 and left to buy cigarettes.
As a programming language? Sure. As a way to work with relational data? It may be my favorite "language" across all domains because of the terse beauty. I am a self-taught, no CS coder but SQL is the one place where I feel like I get all the math I should know.

An opinionated, possibly hot take would be to call SQL "A more elegant weapon of a civilized age".

Or “the worst query language ever, except for all the alternatives”
(comment deleted)
Maybe so, but my father in law, who is a salesman and knows nothing about computers and programming still knows SQL.

SQL is a horrible language in the same way Excel is -- programmers hate it but the what makes it a horrible programming language to developers is what makes it accessible to non programmers.

It looks like the HN hug of death has found a new victim
I was hoping for an alternative to PLSQL or stored procedures. But this isn’t about „Database Programming“, it’s a SQL replacement…
It isn't that bad, at least for those of us that like Ada, and feel at home on SQL Developer.
Looks very nice. Last year I took up rust, coming from c++, and some of the modern features rust brings are just so nice to have (even something as simple as not having to forward declare a class).

This year I started working with postgres and you just can't help but notice how sql is coming from the c-Era of programming. Having better and more modern ways to express my queries would be great to improve correctness and performance.

> …can't help but notice how sql is coming from the c-Era of programming. Having … more modern ways to express my queries would be great to improve correctness …

SQL is based in pure mathematics: set theory, relational algebra.

The process of applying mathematical rigor to your database design to prove correctness is referred to as normalization.

I don’t mind criticisms like “It’s old, yuck”, but criticisms like “it’s not correct” mean you haven’t studied or applied the mathematical underpinnings of sql.

This isn’t talking about correctness of SQL. It’s talking about correctness of queries.
Syntax aside, programmers and mathematicians have a very different view on how things should be done.

Programmers look at data and see opportunities for running a pipeline of transformations (map/filter/...). And they tend to write their SQL like this as well. Or use something like Linq or one of the various pipe syntax SQL extensions.

I would say that this is a major reason why there is this sentiment of "SQL is yucky" by developers. The mental models just don't match.

Data storage and retrieval is a different domain than data processing. SQL is very good at the former, not so great for the latter.

SQL is closer to array programming than the usual imperative implementation of looping (and stream programming like the one in Java and Javascript). A better implementation is functional programming like haskell and clojure (lazy and composition of functions).

I think developers should be able to switch their mental model on the fly according to the current domain instead of getting stuck in the first paradigm they have learned.

Or they did a proper Software Engineer degree that teached on how to use SQL properly, including implementing their own toy SQL engine backed by B-Tree indexes, with raw i-node blocks for storage.
It is older than C. It is based on COBOL era idea of structured English as a computer language. There are better alternatives, e.g. Datalog.
COBOL is indeed the spiritual predecessor of SQL. We have learned a lot since then about PL design, to put it mildly.
Yes, we now programm in straight English, and hope the machine gets it right.
So this really is a NL to query translation problem. That exactly is why the target language matters. Simpler target languages makes AI’s work easier, as it saves tokens and context, so it is less likely for AI to make mistakes.
These kind of comments don't age well in the days of AI programming using English.
AI programming using English makes the database query language choice more important than before. Different languages require different context sizes. A better language is one requires less tokens and context.
I'm curious if you've personally used datalog in any projects. I've written some prolog, but haven't ever worked with datalog.

Minigraph looks promising for some introductory goofing around.

I am the author of Datalevin, a Datalog database. In addition to using it in production personally, I am aware of other people using it. So, the answer is yes.
Interesting! I ordered a copy of your book about Datalevin. I'll give it a try - even though I don't know Clojure.

Thank you for your contributions to open source.

Thank you for the interest and support.
pretty sure clang is older than sql. hal agrees.
Having reusable functions and pipelines compiling to SQL sounds amazing. (EDIT: and sum types!) Will want to try this out on some side project later.

Although for my Elm + backend needs I feel like I still prefer Lamdera: https://dashboard.lamdera.app/ - WebSocket communication and being able to push new data to clients immediately instead of juggling HTTP endpoints and the client having to pull/refresh. `sendToBackend`, `sendToFrontend`, `broadcast` are a great primitive.

I'm wary of languages that seek to own the database. In particular, the claim "Coexist with SQL" seems a bit suspect given that e.g. sum types have a custom binary encoding, which likely makes them difficult to interop with from other languages. This makes the claimed interop with other languages really more of a temporary stopping point towards full Acadia adoption rather than a viable long-term equilibrium, unless you e.g. eschew using sum types.

This makes the database closer to something that Acadia compiles to, rather than something Acadia sits on top of. From my own developer experience this feels off, because I generally expect the data layer to be king and application code to revolve around that, rather than having data representation created in code and the database created off that (this is why I also dislike things like ORMs).

In general I view databases as usually having more longevity than application code, especially as you accumulate more data over time. For serious production applications, the database often outlives multiple rewrites of the production application.

I suspect though my concerns are overall rather minor. The ergonomics of the language itself seem enjoyable. Acadia seems like it would be great as an embedded DSL. It's a bit unfortunate that it currently seems coupled to creating an HTTP server. I think that Acadia has greater ambitions beyond just the database, as evidenced by creating a binary web connection with frontend Elm code to presumably obviate the need for encode-decode layers. It seems like Acadia is meant to be a stepping stone towards a closer frontend-backend fusion. But I agree with mjaniczek that something like Lamdera seems a better fit for that.

But given how early Acadia is, I'm still very excited for where it goes. What I've listed is surmountable and I also feel that often a closer frontend-backend fusion might be worthwhile.

I think, the reality is SQL being simply to old to coexist with a web app use case. All the nice things that article talks about are not possible to nicely integrate with SQL. Current development is done by either writing SQL by hand or by letting ORMs to autogenerate it. Both feel bad because of how bad SQL is. But there is no other option. I hope https://substrait.io/ will gain traction and will be supported natively by databases
I'm curious what you meant by the web app use case and why you find SQL bad?
Is this at all similar to LINQ in C#? I never used it, but I'm vaguely aware of it being a functional approach to querying an RDBMS.
From what I seen (not an expert). It’s mostly sql with a c# flavor and auto translation to native type.
LINQ is based on FP ideas on data manipulation.
This looks reasonably interesting, and Evan is extremely thoughtful about design; I know he’s put a huge amount of work into this.

Personally, I’d be very cautious about adopting closed-source software with such a restrictive license as part of an application, especially given the context of Elm’s trajectory. When Elm went through breaking changes or regressions, or was not worked on publicly for years, users had access to the source and the right to modify it. With Acadia’s licensing, you’d be stranded.

On the other hand, with Elm there was no correlation between adoption and funding for development. With Acadia, he's trying a different funding model, so that might mean better support for both Acadia and Elm.
The Elm project forked into a bunch of different Elms because Evan basically abandoned / killed it. Then he got more interested with this project. What’s to say that won’t happen again?
I would make the larger point that I do not like my software to depend on any software with a bus factor of one that I can't control. Elm had this problem and Acadia has it too.
The bus factor for Elm is currently 2, since Tereza (his Wife) works on both Elm and Acadia.
Unless they make a point to always travel separately, this seems more like a variable bus factor between 1 and 2, just as the Presidency of the United States has a much lower bus factor during the State of the Union address.
(comment deleted)
None of the forks really have the same design goals as Elm though, so if you're compiling Elm code into JS you're probably using the Elm 0.19.2 compiler, not a fork.

- There's an outside chance you're doing Zokka to allow for custom package repos (I think that's the only difference).

- You may be using the Lamdera compiler to use Set and Dict natively with your custom types.

If you're doing something other than compile Elm to JS for UIs, you may in fact be using one of the actual forks.

I agree with the premises, but the result proposed here doesn't look like anything I would like to use unfortunately. Even just looking at a glance you cannot see what it's doing and what each part means.
Oh man. If this lobste.rs comment is correct about the subscription terms then this feels like a really hard pill to swallow: https://lobste.rs/s/ykq7ym/rethinking_database_programming#c...

Still might be viable, but would be tricky to sell.

> SUBSCRIPTION TERMS

> This license is subscription-based and will remain valid only for the duration of your active subscription. Upon expiration or termination of your subscription:

> a) Your rights to use the Software will cease; b) You must uninstall and stop using the Software; and c) You may lose access to any data or content created with or stored in the Software.

On the other hand, norms in software right now are that suckers build and maintain software for free, and the love of the game should be enough for anyone, so it's shocking when people break the norm.
That's not the norm that's being broken here. Most DB technologies provide a "pay for updates, if you stop paying you keep the last version you paid for" model. This is how Oracle prices its DB tech, this is how jOOQ is priced (which is probably the closest thing to Acadia), this is how MS prices its DB tech etc.
The responses to the pricing aspect of this announcement around the web disagree.

And it's not a database, and Java ecosystem an outlier for proprietary software.

> The responses to the pricing aspect of this announcement around the web disagree.

Which responses are you thinking of?

I mean, it's the same for - say - Photoshop?
No, not at least for Photoshop. If you have the subscription version and fail to pay it downgrades you to the free version which has more limited editing capacity but still has read capacities.

More broadly I think the only subscription products most software developers are used to where access to data is revoked is cloud infra. Most software stuff follows models like Jetbrains (where e.g. you pay for updates but keep the oldest version). E.g. this is how things like SQL Server or other paid DB technologies work, where you effectively are subscribing to yearly updates, but get to keep the current version if you stop paying the subscription fee.

Needs proper docs

stuff like "The endpoint keyword" just gets a mention on the front page/readme with no further detail

I don't see anything special here. Haskell has had stuff like this for more than a decade, Selda is probably the one closest to Acadia: https://valderman.github.io/selda/

Despite their claims, this is not substantially different from ORM platforms in many languages.

The issue with defining schemas in a non-SQL programming language is they always lag behind what the underlying database can do. Sure, your ORM-like framework can define basics like primary keys and maybe uniqueness constraints, but can it define partitioning schemes, compression methods or more advanced constraints?

Look at all the features supported here:

https://www.postgresql.org/docs/current/sql-createtable.html

And then consider that other databases have even more. If you manage your schemas in code then you lose access to all of those, and will eventually need to write SQL anyway.

For queries it isn't such a problem, especially if you have a nice compiler. However, I recently lost faith in SQL wrappers/abstractions. The usual justification was that a lot of developers don't know SQL well, but LLMs are great at it. It's easier for the LLM to write SQL than some less familiar DSL. And SQL was written to be relatively easy to understand, especially if you do things like use CTEs and views correctly it should be possible to factor logic out to make even complex queries understandable.

The question for frameworks like Acadia is really: assuming I am fluent in SQL and know every feature of my database, what does the framework buy me? Because that's the perspective an LLM comes to it with.

The point is end to end type safety. Whether that is worth the tradeoff of losing direct developer access to the db primitives is another question.
I agree with end to end type safety but that needs more details to sell what problem its solving. Folks dont buy it for itself
SQL is end to end type safe.
Isn’t sql weakly typed? Or does this depend on the engine?
SQLite is the only one I know of that doesn’t enforce types by default, but I don’t know what the SQL spec requires.
No it is strongly typed, there is no accident that all PL extensions to the base query language have such a Ada/Pascal similarity.

Additional DML has plenty of options to enforce rules that keep data consistency.

While they make the life harder to delete/update/insert items in specific sequences, they can save the day on bad queries.

What happens if a query compares a string to a number?
You get a type error from the database.
As far as I can tell, some engines will implicitly coerce types so “7” = 7
But when do you get that type error?

This is the important bit.

You get it after the app is deployed, the query is ran and a result is expected.

When do I get a type error from my language if it's statically typed? That's right, before I even deploy.

Which end? This moves one end to reach frontend code
Only backend to database. This is talking about typesafe from database - backend - frontend.
You can write raw sql and use the "describe" clause in script, and then generate code with the result. This gives full db-backend-frontend type safety with raw sql queries
They might mean static typing.
There's a lot of benefit in these systems, though there's rough edges and I agree about the basics like PK's and uniqueness.

I've been using Ormin [1] in Nim which works by parsing the SQL tables and uses it to compile time check queries:

    # Multiple joins with pagination
    let page = query:
      select Post(title)
      join Person(name) on author == id
      join Category(title) on category == id
      orderby desc(post.creation)
      limit 5 offset 10
I think that's better since defining SQL should be the source-of-truth for the DB and the code. ORM's always ended up causing trouble in my experience.

Things like indexes, defaults, partitions, etc generally aren't expressible in code without a lot of kludges. Then each DB engine have pretty different rules, syntax, etc for tables.

However having the queries compile time checked, type conversions handled, and the nuances between SQL query syntax handled is rather nice. As you mention it's a much easier subset.

1: https://github.com/Araq/ormin

Just learn SQL. I believe all these SQL replacement layers are just because people don't like SQL and don't learn it, so they learn a training wheels version of it that will cripple their ability to grow because it's simplifications remove expressiveness that caused SQL to be more complex to begin with.

Just learn SQL, it's not that hard. A lot of very very smart people put a lot of effort into it. It's very good. The things that are annoy you about it are often there because of something you don't yet even realize is something you need to be aware of, or because your fundamental understanding of things is just wrong or incomplete.

I already know SQL which is why I like the above. It's SQL with some tweaks to match Nim syntax and to have less ambiguous table/column identification.

Meanwhile embedding SQL in a string with `?` everywhere, manually converting the results, and remembering some of the SQL syntax is annoying.

I think the issue is that while ORMs etc, stuff like ecto…whilst they’re never going to be database native like actual SQL, the value in the abstraction isn’t making querying easier, but making more robust and useful the integration into the host language. It brings it out of database domain and into application domain so that doesn’t have to to constantly reinvented.

You can always be more expressive and portable in raw SQL, that’s obvious, but the things you’re doing have to be used somewhere, so at some point the things you are doing have to cross a barrier. For the 90% use case, ORMs are a pragmatic choice because the good abstractions aren’t about the syntax, they’re about allowing you to talk about and mutate data within the language paradigms that everything else is written in.

No. SQL is just bad. It's an old way of doing things. It's not hard but it's not good.

Take this for example. Why do we have static type checking for typescript? Why do we have a build step for this?

Why DON'T we have it for SQL? Why is it runtime strings? So no static checking and the only way to test if a query works is to run it?

The purpose of these replacement layers is to get it all under one language. Once it's all under one language you get full safety and fusion across the two concepts. Query builders and ORMs are shooting for an ideal, and the ideal makes sense. It's just a nightmare to implement and thus fundamentally there are compatibility issues and that's why a lot of people in general don't like orms.

There's also a sync step where the model in the language has to be aligned with the model in the database which is just an extra mutating state layer which further compounds the bugs.

Only true when avoiding stored procedures.
Yeah. Most systems avoid stored procedures imo. They way to go imo is to use stored procedures for everything, but the standard pattern has stored procedures as some sort of secondary thing.

Either way the types of the stored procedures do not statically mesh well with the types of the application server. So there's a lot of syncing issues here that can only be caught at runtime.

Runtime strings? No static checking? I think you have used mysql and think that mysql is somehow what you should expect, because you are just saying wildly incorrect things. (You don't know what you're talking about)
i do. default way of doing things is sending a string from server to database.

You don't know what you're talking about.

Learning SQL doesn’t absolve you from the fact that, from the perspective of your PL, you’re smashing arbitrary strings together like a Neanderthal, and you can be offered all the support otherwise given to your string smashing problems (exactly none)

It also doesn’t absolve the fact that SQL is not a particularly well-designed language for smashing strings together like a Neanderthal. In fact, you might even say it’s absolutely horrid at it, with random keywords, extraneous syntax, and general lack of compositional capabilities.

The relational model is fantastic. The engines are a work of art. The SQL language is a shitshow. The programmatic interface to a database is an utter mess. None of this is contentious, or should be.

Only because some people are very opinated in avoiding stored procedures, and think smashing strings together is a much better solution.
PL/SQL is cursed and the unstandardized library system is cursed such that every DB’s ecosystem is anemic. Instead of smashing strings, you can code with all the affordances of C90 and still get the chance to smash strings together if you need to do anything beyond utilizing simple variables (EXECUTE) — now with an even worse string manipulation stdlib. And you also get the privilege of working with the some of the most worthless parser errors known to modern man

You can reuse code through extensions/external instead, and have access to real programming languages with actual libraries… but now you’re kicked out of managed environments because it’s not whitelisted, and even if you manage to run it, you’re back to smashing strings together like a Neanderthal trying to communicate to your DB.

PL/SQL is great and using SQL Developer definitely better than smashing strings together.

If only C90 was half as good.

Stored procedures have the wrong versioning model. If they were version-locked to the application code, instead of to the database schema, they'd be less of a pain and people might be more willing to use them.
There are CI/CD processes for deployment, versioning problem is solved at least for 30 years.

Also it is hardly any different from handling version differences in distributed systems, or split between frontend and backend on Web applications.

Version differences in distributed systems (including Web apps) are a real pain! In many circumstances they're unavoidable, and we've developed various techniques to make them marginally easier, but if you can avoid the issue entirely by just not having the thing be distributed, that's the more maintainable choice.
Oracle has a feature called 'editions' that does this. Different DB sessions can have different versions of redefinable objects like stored procs and packages.
> Just learn SQL....

I agree. In my experience, ORMs are more complex and harder to learn to an expert level than SQL. Knowing Java (but not SQL) doesn't help much with learning Java ORMs (Again, to an expert level). Besides not supporting all the SQL features of some DB, ORMs also covers other things such as caching.

Learning ORMs is likely just as difficult as learning SQL. It is likely harder to learn how to optimize performance with ORMs.

SQL as opposed to code has the advantage that it can be kept in a separate file, and thus modified by experts in databases without changing the code. The article claims the author found migrations harder with SQL than with his framework. I would think it would depend a great deal on the database one is migrating.

I'm not convinced that LLMs make things easier, you still need an expert to verify the generated code, and to tune it, as often the database is business critical with serious consequences if wrong, slow, or turs out to be infringement of someone's copyright.

Just learn SQL!

Agreed with you here. In my experience the best solutions go the opposite way, and parse the SQL in ways that can be used from the application.
The problem here is that you still have to do some pointless and tedious conversion between generated data structures and your domain data structures. Maybe with LLMs, some of that tedium goes away, but you still have to test, maintain and understand that part of the code.
Agreed, that's why I chose to implement a simple ORM for my language's multi-platform database library. It has a mandatory 'id' column, for simple updating and deleting, but table creation and complex queries are done in plain SQL.
A core idea of the relational model is to seperate the logical model from the physical layer including optimizations, indexes etc.

So it makes sense to only expose the logical model at the ORM layer.

The problem comes if you want to define the database schema through the ORM layer, rather than just represet it.

> Look at all the features supported here:

> https://www.postgresql.org/docs/current/sql-createtable.html

Unironcally, yesterday i was vibe-coding a small app for personal use using Django and was quite shocked to discover that Django's orm does not support something as simple as specifying a database schema other than the default "public" one out of the box.

You either have to add options specific from libpq:

    DATABASES = {
        "default": {
            "ENGINE": "django.db.backends.postgresql",
            "NAME": "mydatabase",
            "USER": "myuser",
            "PASSWORD": "mypassword",
            "HOST": "localhost",
            "PORT": "5432",
            "OPTIONS": {
                "options": "-c search_path=myapp,public",
            },
        }
    }
Or you have to do it from the postgresql side:

    ALTER ROLE myuser
    IN DATABASE mydatabase
    SET search_path = myapp, public;

It's not ergonomic at all.
Reading this, I mistook it for a slightly different idea: using these functional languages directly inside the database process, avoiding SQL altogether.

I've wanted to try that out with e.g. Roc and a reimplementation of SQLite's on-disk format. O(f course, that's a non-starter for production use but it could be an interesting experiment to see what that programming model was like.) The database would become kind of like a library you use to build your tables and queries with.

Also, thank you for calling it a 1+n query, not an n+1 query ;)

It seems like this is a few things:

1. An Elm-like programming language that lives in .db files

2. A compiler from this language to strongly-typed database procedures in a target backend language

This has more in common with a semantic layer than an ORM.

What you gain is a shared language that connects the table definitions (say a SQL migrations folder) and your API language (often handwritten SQL). This can be type checked and optimized for you.

But for me the big question is what functionality do you lose? Can I express everything that PostgreSQL can?

Hmm. I've skimmed the article. It looks to be another ORM/FRM type thing. There are many issues with such things, but for me the most troubling is this: in most systems (obviously...it depends) you don't want to wind the database around the axle of any one software component or language. Having the data separate from the code, and defined/managed with a language that suits data management is a feature not something to be designed out. My hunch is that people who come up with these "solutions" fail to realize this. They then condemn everyone using their layer to endless hair pulling trying to figure out "what SQL did it make from that?" and "how do I make it do this SQL?".