36 comments

[ 0.33 ms ] story [ 33.8 ms ] thread
I co-authored this paper. It's a new technique to generate 3D graphics as source code instead of a point cloud.

Under the hood, it generates 3D objects with separate, sophisticated internal assembly, producing an editable "kit of parts" (instead of monolithic blobs).

E.g. imagine you generated a 3D washing machine via this approach. It's not merely going to be just "geometry" that looks like a washing machine. We actually know that there is a `Door`, `Drum`, `Control_panel` etc. Which things belong to which assemblies. What moves and where its pivot is. And eventually what those components are supposed to do.

Most current 3D GenAI cannot do this since it generates "monolithic blobs" that look good, but are unusable in downstream workflows (e.g. game engines). I.e. if you generate a 3D bicycle using traditional approaches, it's basically a blob. When you need the wheels to turn, a human (or another AI) must spend time cutting the blob into parts, naming them, placing pivots and rigging joints. I.e. you need post-generation segmentation workflows of some sort.

The paper breaks down the whole technique, and comes with a github repo too if you're interested.

The paper breaks down the whole technique, and comes with a github repo too if you're interested in viewing it.

How far are we from speaking a GameCube-era game into existence as a pastime?
GameCube-era game (with full synchronous multiplayer play): 1-2 quarters. Early "self-generating" Metaverse: ~ 3-6 quarters. The Matrix: ~5-7 years
I am doubtful about the timeline.

Not because AI can't do it; it totally can. LLMs have been able to run the full artist + code pipeline at least since the beginning of the year. I've built several physics-synced network simulation stacks without reading a single line of code. Agents playtest my games overnight and I wake up to a list of technical issues fixed, and FPS boosted. If you know how to ask the shaders will look great.

The problem is that making a game actually worth playing (something that Nintendo would allow to be released) isn't something that was ever possible to do as a pasttime, AI or not. You have to be in front of the computer all day guiding it. Worse, AI does not have any notion of experiencing or evaluating fun, so you can't automate this. That would be a killer research problem to tackle, though!

If we're talking about making something that passes a sniff test, you could make a metaverse right now. It just wouldn't beat the bottom of the Steam barrel in terms of what players prefer.

I agree. AI does not solve "product market fit". It mostly solves engineering. Currently AI is great as a tool, but not really a co-creator with taste.
> Agents playtest my games overnight

I'm curious if you want to elaborate. What kind of games? Turn based? Do you just feed it repeated screenshots?

> Agents playtest my games overnight I'm curious if you want to elaborate. What kind of games? Turn based? Do you just feed it repeated screenshots?

3D ARPG with dozens of systems, think Genshin or Fortnite. But it's all typescript running in the browser, so native browser introspection/debuggability came for free.

Fable has a cromulent time building its own tests, tools and pipelines. But there's no magic, it literally uses the gamepad and plays the game itself, taking screenshots, profiling, and debugging as it goes. The game ticks are fully controllable so it can frame-advance at its own pace.

This looks super interesting. I'm trying out the hosted app using "bring your own key", I've added an OpenAI key but it doesn't seem to let me generate a 3d model. It's still saying I need credits. Is this expected?
Have you explored optimizing the assets to be game-ready? This kind of decomposition works if you have a single object on screen, and it's super artist + programmer friendly. But the generated assets have ~50 mesh parts, which means importing just a couple of these into a scene and you've blown your entire draw call budget for a shippable game. It's the brick wall every gamedev realizes after trying to make a scene out of easy-to-work-with primitives. You just can't hit a playable frame rate like this unless your entire game consists of just a few objects.

Have you experimented with atlasing, mesh fusion, baking animations, standardizing PSO's to a scene budget, etc? Because if this can't be automated, I've found it really limits the utility of such freeform generation techniques in practice.

[flagged]
My ideal would be the model generates the sources (so it's programmatically tweakable after the fact, which is the entire appeal), but there is an open source compilation tool or something that can optimize it for the different use cases after you've made your tweaks.

Devs could probably make their own version of this; every engine/consumer of assets is different and needs tweaking. But there is no engine that won't choke on the raw version, so someone needs to make an optimization baseline or show how it's possible.

Thanks for considering!

[flagged]
Isn't today's graphics stack moving in the direction of dynamically-optimized meshes (ie. Unreal's Nanite)?
Yup. But that's 10x more complicated, and needs an even more specialized baking phase that's even further from the raw representation. It only multiplies the issue.

And Nanite is not really designed for the kinds of lower fidelity fully articulated objects we're talking about here.

I will say though: I think things are going to move to neural rendering faster than people expect. So maybe the future is low fidelity highly articulated objects rendered with img2img. But nobody is seriously doing that yet.

neural rendering isn't rendering
If all of these meshes use the same shader (which they do, its just PBR) they can be all drawn with a single multidraw indirect call.
Cool. Do the individual parts still use point clouds? Or are they meshes or CSG?
The parts are originally defined in code, not stored as point clouds. That code builds the geometry using primitives, curves, custom mesh operations and sometimes CSG/booleans. When executed in Blender, the final exported GLB contains meshes.
How does this compare to parametric modeling tools like Fusion/Solidworks/ProE ? Is it more about the integration with game specific tools?
> there's a showcase (+ github repo) you can play around with: https://nova3d.xyz/

Before anyone else bothers giving them your Google account, there's zero free generations, something they conveniently don't disclose until after funneling you to sign up.

[delayed]
The current output is polygon meshes + the source code. So technically "surfaces", rather than CAD/B-rep solids. Many parts are closed volumes, but we don't claim manufacturing-grade solid geometry. This paper is focused on the Blender/mesh path. A CAD-solid backend would be a different target.
@baigy Huh. Interesting. I was just doing the final clean-up for something convergent to this research that I had been working on for the past few months. I think I arrived at your thesis (code first semantics from a different direction in CAD, so I think it'd be interesting for us to compare notes.

Have you formalized this into a compiler infrastructure yet? I think Python on its own would be too slow to build complex parts, especially since for triangle mesh, accuracy inversely correlates to performance.

Vision is generally not the most reliable form of checks for LLMs, even on GPT 5.6 Sol, so a recommendation I would have is to instead emit JSON or CSV of the color/topology data for the LLM to inspect directly, and this is the instance where ray query for topology checking will greatly improve accuracy in general. SDFs are a bit more complicated right now, I have a full implementation designed for 3D analysis

My own experimental compiler generated mesh suffers from the spiderweb effect: it's very polygon efficient but not very friendly towards UV unwrapping in general, and I'm struggling to find the correct approach for that. If you have any suggestions, I'd love if you can point me towards the correct approach.

Definitely very interesting though.

[dead]
https://github.com/yuechen-li-dev/Aetheris

So, yeah, mine is full geometry kernel that was made for CAD. I'll do a full write-up on Show HN later as there is just way too much stuff to cover for the project. The gist of it is that it is a compiler for 3D models that takes the high-level language, which I named Firmament, and lowers it to STEP AP242 mapped BRep in C#.

So, for SDF, it was originally conceived as a method to solve the general case 3D BRep boolean problem through methodology similar to libfive/Fidget in FRep directly. The problem is that we immediately ran into the same wall that everybody else did attempting to recover mesh/BRep structure from the SDF blob, spent like a week doing BRep patches for it, before we ultimately concluded that it was not really possible to do as FRep is a lower representation than either BRep or mesh. It's probably more useful for continuum physics/fluid dynamics in the future than it is for CAD/solid mechanics/3D modeling, but currently the SDF pipeline is just kinda sitting there as dead code and not being used much.

https://github.com/yuechen-li-dev/Aetheris/tree/master/Aethe...

This is super interesting. Especially the high-level language > STEP/BRep lowering. Tt feels very complementary to the mesh path we're exploring and potentially the right backend when exact solids matter. I feel your SDF conclusion is useful too. It's tempting to treat it as a universal intermediate, then discover you've lost the structure you need later.

I'll dig into the repos and would compare notes afterwards.

Very cool.

"Asset as a service" is something I've been kicking around for a while. I did some work on something similar at Apple.

The idea was to create "Assets as a service" where the generation system can decide how much configurability remains live at run time, and how much is "compiled away" at asset generation time.

https://patentimages.storage.googleapis.com/43/19/69/c4c2dce... https://patentimages.storage.googleapis.com/d3/4a/df/d329bb5... https://patentimages.storage.googleapis.com/d2/31/6d/123b055...

Whoa super cool, you guys got the patent too. I hope there's no cease and desist in my future lol (kidding..... not kidding).
I think this approach (separate parts) is the right call. This is how human artists build models, and once the model is built and segmented you can decide which parts go together, and then then group, remesh, UV-map and bake those parts. All of which are hard problems too but are getting closer to being automated.
I was thinking about this approach, and this paper basically validates it.

The inference cost must be extremely high compared to diffusion based approaches. I wonder if this would ever be useful for more organic sculpting type workflows. E.g. for organic non-hard surface models, use a diffusion model for generation, and leverage this LLM codegen and tool call approach for retopology and cleanup.

My belief is it'll get better over time with organic shapes as LLMs improve their ability to synthesize higher order differentials. I also believe the future may not be coded 3D everywhere, but a mix of code and dumb 3D. But I'm not a huge fan of inverse code recovery; I found that to be extremely lossy.

Do you want to connect over discord or linkedin or something, to cross-pollinate ideas?