I’m a bit skeptical about LLMs actually being able to perceive 3D context properly. I think a well-designed structured scene graph could be promising.
I’m a bit skeptical about LLMs actually being able to perceive 3D context properly. I think a well-designed structured scene graph could be promising.