43 comments

[ 1.8 ms ] story [ 38.7 ms ] thread
[deleted]
Literally the first backend they mention is Vega-Lite.
When does it get too abstract? What’s wrong with plotly? Or plotly express? How is yet another JSON spec era-anything?
>When does it get too abstract?

When the abstraction detracts more value than it provides.

I think this is an interesting direction to explore. Plotly is fine. So is D3 and plenty of other solutions, but neither represent the ultimate pinnacle of charting.

Why is this "for the AI era"? Isn't it kind of not for the AI era since an LLM can create a gnarly python or js chart of whatever you want in seconds?
(comment deleted)
It's cool...but is it needed? I'm thinking of Apache echarts and any of many other mature charting libraries. Kinda seems like reinventing the wheel; thing gets released, thing gets more and more new features bolted onto it, eventually someone offers another thing that's basically the original thing with some slightly different design and syntax opinions...

Switch backends to use their native strengths: ECharts for hierarchical sunbursts, Plotly for statistical and analytical traces, or Excel for editable charts embedded in a workbook.

Or just find a charting library that you like and actually get to know what it can do, vs mixing and matching presets from different libraries but never tweaking them.

Even in the Era AI, GGPlot's API is still the best charting API. The name "Grammar of Graphics" isn't just marketing, they literally sought to write a god damn grammar to was capable of expressing all possible qualitative graphics.

They even wrote a book about how they went about it (not that it speaks to the quality of the API) https://link.springer.com/book/10.1007/0-387-28695-0

I actually stumbled upon this book when I was trying to look up how draftsmen (with pens and pencils on paper) did qualitative graphics as I found they had a lot of charm as opposed to modern charting libraries. It's something I noticed when looking through a bunch of historical RBA (Reserve bank of Australia) annual reports, the 1960-1980 charts had a lot of character, but then you go into the early 2000s and its a stale chart from excel.

Anyways ggplot doesn't really recapture the magic of those older charts, but it seems use quite a few of those as a baseline for how to communicate information. Like in figure 20.1 they talk about efforts to replicate older inforgraphics that showed Napoleon’s March on Russia, this graphic here (I think the example in the book is a bit nicer than the one in this blogpost IMO)

https://www.andrewheiss.com/blog/2017/08/10/exploring-minard...

On top of the charts just look nicer than anything you could produce with pyplot (and any API built on top of it) as pyplot seems to be have some really limited raster based rendering or something and the text handling is incredibly limited, I've never had this issue in ggplot.

I feel like most software engineers aren't exposed to because it exists in the R ecosystem which is more so data scientist, econometricians, statisticians and other quantitative data professions, but it definitely one of the nicer APIs and I wish more people in the node and python ecosystem copied their homework. I see vega's full name is something to do with grammars, but idk it's for the same reason.

Yeah this is not needed and every single day it is needed even less.
So this is one interface that can render to multiple charting backends?

If AI is writing the "Flint", why not just have it write the backend code instead? I'm not sure why I would want pluggable charting backends.

I can see an argument for providing simpler APIs for LLMs, though, so that it can be more token efficient for example.

This is just json right? One issue with llms right now it's that if you give them a json specification, they're not always amazing at following it. I think it makes sense to have an agent tool that takes llm json, a file name and a spec path and only writes out the file if it conforms.
unfortunately JSON is doomed to fail in "the AI era"

LLMs are surprisingly bad at generating JSON.

Does it do flowcharts and arch diagrams? Doesn’t seem like it.
Can it respond to my several open github tickets?
Maybe they’re not selling this well. Maybe they’re so close to the AI research that this seems like an obviously good idea.

But there’s not one word of why this is good for LLMs, or how they tested/measured that.

My gut would tell me that a new solution put up against all the vega lite specs it’s already be trained on would be a hard thing to win.

Every time I needed to do charts at work I had a conundrum. Libraries are aplenty but basically only plotly does everything, but it’s not the prettiest. I’d much rather developers contributed new features into existing libraries.
(comment deleted)
So a stringly typed JSON-based DSL without a linter or LSP?
If it's made for LLMs, the spec should be yaml not json. Way more token efficient
What is wrong with "please make plot XYZ in plotly?"
(comment deleted)
(comment deleted)
I have tried using Flint vs asking the AI to generate a Vega lite spec directly, and in my opinion Flint was not as nice of a solution.

Flint is fine for doing predetermined chat types, with very low customization. But I found using an agent or sub agent to create the Vega spec directly allowed for a lot more flexibility, and ultimately that means higher quality visualizations (stuff like adding points for min and max on a timeseries, or adding a callout marker for a date where some event happened).

That being said, with Vega lite you have to validate your chart specs, provide specific guidance, and play whack a mole with Vega bugs/idiosyncrasies. So Flint is more reliable if you don’t want to dedicate a whole skill to making charts and want to get running quickly.

idk ~ it's either a bit thin or i'd expect something more substantial "for the AI era"

... as the kids would say: "weak sauce"

... and it's just a tool to create charts not a "visualization language"

you could probably vibeslop a pipeline from some ad-hoc DSL to excel or pandas or R and get a better integration with your business context

This needs a side-by-side config comparison with something like echarts config schema format. I really don't see the point.

I'm 99% sure the verbosity required in the system prompt to teach non-M$ models this new ever-so-slightly-different-but-not-obviously-necessary chart def abstraction format, and the iterations required to get it right, will outweigh any supposed efficiency gains resulting from using it.

Just stating the obvious.