Ok, this is really cool. The fact that the robot can use pointing to decide where to go is a great design decision, and robotics really is the next frontier. Definitely cheering on Mistral here!
For a claim such as state of the art, or claims such as "great at any task" needs something of more substance. I've seen maze-solving robot competitions which can zoom around in seconds. The sped up video in the first part, and the "obstacle avoidance" are too slow for me to believe this is state of the art.
While impressive at 8B, what would the expectation be in real life, that it's run remotely or autonomously with a strapped on GPU and battery?
The claim is very specifically that it's SOTA on the R2R-CE benchmark, which is a bunch of 3D environments in a simulation. So, yes, it's SOTA; no, it's not very different than a maze. And it's sure not anywhere near anything that could be considered SOTA in the real world... if such a SOTA was even possible to define objectively.
(it's not because evaluation in the real world is very, very tricky).
Producing specific niche models for 100 year old industries that have mountains of data and warehouses full of folders will be the european take on AI.
It may come late but it‘ll be safe and reliable. It also requires a lot of OCR.
It's implied, and I'm hoping it's true, that this is a map-less navigation. Which is impressive. This kind of task is much easier if you have a pre-captured map of the environment, but if they are doing this without a map it's great. Historically you were always faced with "The Kidnapped Robot" problem where robots that didn't know where they were couldn't navigate even a little bit. Here the robot appears to be able to follow directions as long as they are interpretable from its current vision (or via dead reckoning).
Can you explain how it is much easier if you have a pre-captured map given what they are doing without using any sensors, all you have is perhaps these recent feed forward tokens not actual Geometry.
This looks to not be an openly available model, but I think if it were, availability of an easy single-camera navigation setup could allow for a lot of cool hobbyist projects.
If you're wondering what prevents or mitigates AI hallucinations on the AI layer from replicating or acting out on the physical layer look up QNX. They manage the deterministic reasonin gof robotics. You know them better as Blackberry.
This is very cool. Congratulations to the Mistral team. Map less navigation in the outside world has been around for quite a while. But map less navigation inside the buildings is relatively new. Some stanford researchers trained a vision model (PIGEON) which could tell the geo-location from any image. It was not released publicly due to privacy nightmarish (stalking!) possibilities but I am assuming similar type of tech has gone behind this robot. if someone knows more, feel free to correct.
Funny how nearly all model improvements this year are demonstrated on the subset of use cases where brute force / reinforcement learning is most effective:
Robotics (using physics sims)
Cybersecurity (red team / blue team)
Math (using automated proof checkers)
Programming (using compilers)
For the record I think robotics is a totally logical place to use this training approach and this is very impressive. But if we zoom out and think about LLMs in general I’m not sure this inspires confidence in AGI arriving any time soon. I would also propose that this is a form of overfitting / training-test contamination.
Take cybersecurity for example. Through brute force techniques you will gradually memorize all of the possible exploits. So when fable breaks into a DoD network everyone is shocked but in reality it basically memorized all possible exploits including some zero day.
I’d be much more interested to see if fables performance is preserved as new exploits arise (NOT zero day - negative day meaning exploits that don’t exist yet). Would fable still find them? Or would they need to retrain it on the new software stack continuously in order to identify the zero days.
This is an important distinction that I have not seen made before.
I wonder how Mistral will prioritize its robotic development against its LLM development. We have either players that prioritize both (Google, AMI), or players that prioritize coding and agentic (OpenAI, Anthropic, ...).
The multi-sensor comments are confusing. This issue is a command->semantic understanding problem, not a sensor fusion problem or trajectory planning problem per se.
It's not like the true depth of field is important for the robot to plan when it's moving at turtle speed and can stop quickly.
What is the realistic path to getting to play with this? I would love to hook this up to OpenClaw for hobbyist exploration. My dream has been to embody OpenClaw into a farm robot (been looking at adapting one of those RC lawnmowers that is tracked and built for mowing steep hills) so that I can assign it various tasks around our acreage -- "Explore the fenceline take pictures of the plants. Find all of the poison ivy and invasive honeysuckle and spray it with your Roundup sprayer. Repeat this every week and report the species map after every pass. Come back to the barn and charge yourself whenever you get low."
It's not hard to put OpenClaw into a robot body (numerous YouTube videos showing people doing this sort of thing), but when you dig in and see what people have done, the actual movement portion is always the clunkiest part (and this matches my own experiments as-such as well). It feels like an 8B model like this would be perfect for solving pathing and navigation issues.
Anyone who may be more experienced with Mistral (or companies like them) -- are they interested in hobbyist builders who would be experimenting with things like this? Or are they primarily looking for commercial partners? I would be willing to pay a license fee to use the model in my experiments, but if I'm just one guy, I'm not sure they'd want to work with me unless I were building a business out of it (which I'm not).
38 comments
[ 3.2 ms ] story [ 54.0 ms ] threadSOTA 80% means a practically useless robot. What are they really imagining their ICP to be here?
But I'm scared for when those home helpers get drafted to fight in wars, either for or against me...
While impressive at 8B, what would the expectation be in real life, that it's run remotely or autonomously with a strapped on GPU and battery?
(it's not because evaluation in the real world is very, very tricky).
I would like to know what it did the other 23.4% of the time!
It may come late but it‘ll be safe and reliable. It also requires a lot of OCR.
here's the link to the PIGEON paper - https://lukashaas.github.io/PIGEON-CVPR24/
Robotics (using physics sims)
Cybersecurity (red team / blue team)
Math (using automated proof checkers)
Programming (using compilers)
For the record I think robotics is a totally logical place to use this training approach and this is very impressive. But if we zoom out and think about LLMs in general I’m not sure this inspires confidence in AGI arriving any time soon. I would also propose that this is a form of overfitting / training-test contamination.
Take cybersecurity for example. Through brute force techniques you will gradually memorize all of the possible exploits. So when fable breaks into a DoD network everyone is shocked but in reality it basically memorized all possible exploits including some zero day.
I’d be much more interested to see if fables performance is preserved as new exploits arise (NOT zero day - negative day meaning exploits that don’t exist yet). Would fable still find them? Or would they need to retrain it on the new software stack continuously in order to identify the zero days.
This is an important distinction that I have not seen made before.
This analysis by Toby Ord demonstrates why it’s a problem if frontier improvements are coming from reinforcement learning (brute force methods) from a purely computational perspective: https://www.tobyord.com/writing/inefficiency-of-reinforcemen...
It's not like the true depth of field is important for the robot to plan when it's moving at turtle speed and can stop quickly.
It's not hard to put OpenClaw into a robot body (numerous YouTube videos showing people doing this sort of thing), but when you dig in and see what people have done, the actual movement portion is always the clunkiest part (and this matches my own experiments as-such as well). It feels like an 8B model like this would be perfect for solving pathing and navigation issues.
Anyone who may be more experienced with Mistral (or companies like them) -- are they interested in hobbyist builders who would be experimenting with things like this? Or are they primarily looking for commercial partners? I would be willing to pay a license fee to use the model in my experiments, but if I'm just one guy, I'm not sure they'd want to work with me unless I were building a business out of it (which I'm not).