I’m not convinced GPU acceleration is a meaningful advantage for most charting use cases. Most dashboards don’t render enough data for it to matter. Once a chart is dense enough for rendering to become the bottleneck, it normally is already be too crowded to be meaningful.
Zooming can justify supporting larger datasets, but sampling/viewport culling and level of detail often avoid drawing unnecessary points...
Why settle for sampling when you can have the whole dataset?
The spiral pattern is an excellent example. "Sure it looks like this when you zoom out, but when you zoom in, you can see the finer structure of the points..."
I can give you a prime example where you want to render all the data and quickly, oscilloscopes. It's quite common that you fetch traces with millions of samples, but then want to zoom into specific regions. There are lots of similar applications in experimental signal processing, where you want to have large data sets, but sampling easily will hide details that you want to see (unless you already know what exactly your data looks like).
it's possible to render data out-of-core with XY, allowing it to render the entirety of OpenStreetMaps (that's 10,742,674,832 nodes!) with sub-second pan/zooms. it's a bit difficult to host online but you can try it out locally: https://github.com/reflex-dev/xy/tree/main/examples/osm
Interesting; how do the examples compare to datashader?
Edit: for my use cases, I use napari (~1e7-8 points) if I need true interactivity; otherwise, datashader/holoviz, or even just fast-histogram's 2D histograms work.
For extremely large point clouds, these caveats[0] still apply. It irks me when people make dense scatterplots without any indication of just how dense some portions are.
Still, if it can indeed handle 1e10 points, that's pretty impressive.
Check out mosaic from uwdata which works on top of Observable plot
Or plotly-resampler which works on top of plotly and uses the rust package tsdownsample to aggregate on the 4pixels per pixel shown level (to make antialias work)
the grammar of graphics approach really is a great abstraction, and I'd love to see xy work in that direction
Interesting approach to large scale visualisation. Moving reduction into Rust and sending screen bounded data to WebGL seems much more sensible than pushing millions of raw points into the browser. How does it perform with real time updates? I am assuming it is much more performant? Any plans for a prod deployment?
I can imagine this useful to 'compress' gigabytes of data onto a 2d canvas quickly. For that, I appreciate the effort.
One thing that would be useful is to read up on Ed Tufte's principles of data visualization. Many graph libraries don't implement basic visualization principles to make they key point clear, easy to see while still keeping the full depth and complexity of data visible.
23 comments
[ 0.28 ms ] story [ 34.0 ms ] threadThe spiral pattern is an excellent example. "Sure it looks like this when you zoom out, but when you zoom in, you can see the finer structure of the points..."
it seems a lot of people don't know about histograms...
> XY is an extremely fast, interactive, customizable Python charting library
which is it?
Love rust as the impl.
Edit: for my use cases, I use napari (~1e7-8 points) if I need true interactivity; otherwise, datashader/holoviz, or even just fast-histogram's 2D histograms work.
For extremely large point clouds, these caveats[0] still apply. It irks me when people make dense scatterplots without any indication of just how dense some portions are.
Still, if it can indeed handle 1e10 points, that's pretty impressive.
[0]: https://datashader.org/user_guide/Plotting_Pitfalls.html
Or plotly-resampler which works on top of plotly and uses the rust package tsdownsample to aggregate on the 4pixels per pixel shown level (to make antialias work)
the grammar of graphics approach really is a great abstraction, and I'd love to see xy work in that direction
One thing that would be useful is to read up on Ed Tufte's principles of data visualization. Many graph libraries don't implement basic visualization principles to make they key point clear, easy to see while still keeping the full depth and complexity of data visible.