Wednesday, September 9, 2026
venfeedSubscribe

NavMCP got 78.3% success on a Unitree Go2 by scaffolding a VLM onto a navigation model

The framework reports state-of-the-art results on three embodied question-answering benchmarks, with margins of 10 to 45 points. The architecture is the finding, not the model.

Venfeed Editor2 min read
ShareXBlueskyLinkedInHNRedditEmail

A framework called NavMCP reports a 78.3 percent success rate on a Unitree Go2 quadruped by coupling a vision-language model with a navigation foundation model, claiming state-of-the-art results on the HM-EQA, MT-HM3D and EXPRESS-Bench benchmarks with margins of 10 to 45 points and a 14.9-point gain on HM-EQA.

The interesting claim is architectural. Neither component is novel; the result comes from how they are joined.

The division of labour

Embodied question answering requires a robot to move through an environment to answer a question about it — find out whether the kitchen window is open, and the system must work out where the kitchen is, get there, look, and report.

That decomposes into two problems that are badly served by a single model. High-level reasoning about what the question means and what would answer it is language work. Low-level navigation — building a spatial representation, planning a path, avoiding obstacles, recovering from failure — is control work, operating at a different timescale and on different representations.

Vision-language models are good at the first and poor at the second. Navigation foundation models are the reverse. NavMCP's contribution is the scaffolding between them: the vision-language model reasons and issues goals, the navigation model executes, and the interface between them is what the paper is really about.

That is the same conclusion the software agent literature has been converging on. τ^τ-Bench found the best systems at 23.9 percent building customer-service agents against 82.2 percent for expert humans; BAAI's DisCo lifted MLE-bench from 31.1 percent to 72.9 percent by giving agents well-formed reusable skills. In both cases the model was not the limiting factor — the structure around it was.

Why running on a Go2 matters

A great deal of embodied AI research is evaluated only in simulation, where results are cheap to produce and transfer poorly.

The Unitree Go2 is a commercially available quadruped costing a few thousand dollars. Reporting a real-hardware success rate on an inexpensive, widely owned platform means other groups can attempt to replicate it, which is a meaningfully higher standard than a simulation number.

It also puts a figure on the state of the art that is easy to interpret: 78.3 percent success means roughly one task in five fails. That is a research result, not a product.

Where this sits in a busy fortnight for robotics

The surrounding activity has been unusually heavy. Figure AI signed a $3.5 billion compute agreement with Nscale targeting 100,000 Nvidia GPUs for deployment in the second half of 2027. Tripo AI raised about $446 million for generative 3D, two months after a $150 million round, with robot simulation behind the demand. Lucida, published the same day as NavMCP, rebuilds cluttered indoor scenes from ordinary video into editable assets for simulators. XDOF is in talks at a $1.2 billion valuation three months out of stealth. MIT's Phillip Isola has proposed cloud language models controlling connected robots directly.

The common thread is that robotics is being attacked as a data and integration problem rather than a hardware one — and NavMCP is evidence for that framing, since its gains came from architecture rather than from a better model or a better robot.

The paper does not report how the 78.3 percent breaks down by failure type, which is the number that would say how far from deployment this is.

Venfeed Editor
Editor in chief

Runs the newsroom. Rename this profile in the studio to your own byline.

The Feed · weekdays, 6:30am ET

Every weekday, the AI stories that moved money or shipped code.

No cross-posting, unsubscribe anytime. See all newsletters