Issue #050 — This Week's View

Share

# The Benchmark Nobody Funded Is Telling the Truth About Physical AI

VANTAGE-Bench just exposed a gap in vision-language models that nobody in the pitch decks mentions — and it points at where the next wave of robotics value actually sits.

Here is what I believe: the robotics industry's obsession with embodied, human-scale AI has blinded capital to the near-term, deployable, revenue-generating segment of Physical AI — fixed-camera infrastructure intelligence — and the new VANTAGE-Bench results prove the gap is a data and evaluation problem, not a scaling problem. Investors pricing robotics portfolios around humanoid end-state narratives are systematically underweighting the boring cameras bolted to warehouse ceilings that will pay back first.

The finding that should reprice a category

A team of researchers recently released VANTAGE-Bench (arXiv:2609.09396), the first benchmark to evaluate vision-language models on what they call "Infrastructure AI" — fixed-camera, open-loop perception for safety monitoring, logistics, and operational logging — as opposed to the subject-centric consumer video that dominates embodied AI evaluation. The annotation effort is serious: 3,346 media assets, 4,281 image-grounding annotations, and 27,404 detection boxes across logistics, transportation, and smart spaces.

The results are uncomfortable. Across 17 models evaluated zero-shot, event verification, referring expressions, and temporal localization fall roughly 9 to 24 points relative to consumer-centric benchmarks — at *every* model scale. In absolute terms, no system exceeds 55.7 mIoU on temporal localization or 37.3 SODA_c on dense video captioning (per the VANTAGE-Bench paper). Frontier models track well over short horizons and fall apart as the horizon extends.

But here is the detail that matters most for capital allocation: the shortfall is concentrated, not general. Video question answering stays within 5.3 points of VideoMME. 2D spatial pointing shows no shortfall against BLINK. And — critically — open-weight models lead 2D object localization outright. As the authors note, neither scale nor proprietary access explains the pattern.

Read that again. The gap is not a capability ceiling. It is a data distribution gap. And data distribution gaps are the cheapest, fastest gaps in this industry to close.

Three arguments for why this repricing is overdue

First, infrastructure AI is where the buyers already are. Every warehouse, port, solar farm, and construction site on earth already has cameras mounted and cabled. There is no deployment problem, no form-factor problem, no battery problem. The customer's pain — safety incidents, theft, throughput losses — maps directly to insurance and compliance budgets that exist today. Compare that to humanoid deployments, where the Maze Intelligence catalog tracks 600+ robot models across roughly 350 companies globally (our catalog count), and the number running commercial hours at customer sites remains, by industry consensus, a rounding error against the capital raised. Second, the failure mode is tractable. The VANTAGE results show models failing specifically on temporal tasks — tracking over extended horizons, event verification over time. That is a training-data and benchmark problem, and the benchmark now exists to close it. When a gap comes with a public leaderboard, a data release, and an evaluation harness (all published at vantage-bench.org), it stops being a research mystery and starts being an engineering roadmap. This is exactly the pattern we saw with embodied benchmarks a few years ago, and capital chased that capability curve aggressively. Third, the open-weight result inverts the moat story. If open-weight models win outright on 2D localization, then infrastructure AI value does not accrue to whoever owns the biggest foundation model. It accrues to whoever owns the *deployment surface*: camera networks, edge compute, domain data from actual industrial sites, and integration relationships. That is a venture-investable thesis at seed and Series A — capital-efficient, revenue-tethered, and uncorrelated with whether anyone's humanoid ships in 2027.

The steelman

The strongest counterargument deserves a straight answer: infrastructure AI is open-loop and low-margin, and open-loop perception is precisely the work that gets commoditized first. A model that flags a forklift near-miss is a feature, not a company; the incumbents in video security already have distribution, and foundation-model providers will absorb the perception layer as they march down-stack. Embodied AI, by contrast, is where the defensible end-state lives — the labor market is tens of trillions of dollars (roughly; various analyst estimates), and whoever owns general-purpose manipulation owns the future. Betting the category on camera analytics is optimizing for a good outcome in a small market while missing the transformational one.

I don't dismiss this. But note what it concedes: the humanoid end-state is real *and* far. The strategic question is sequencing, not selection. The teams that learn to sell perception systems into industrial sites today — accumulating domain data, deployment infrastructure, and customer trust — are building exactly the assets that embodied AI will need later. Infrastructure AI is not instead of the end-state. It is the on-ramp to it.

What I'd tell the partners

Stop underwriting robotics portfolios as a single bet on legs and hands. The VANTAGE-Bench data says the perceptual layer of Physical AI is broken in a *specific, fixable, unclaimed* way, in a market segment with existing hardware, existing budgets, and existing buyers. The simulation gap, the data moat, and the go-to-market question everyone agonizes over for humanoids are, for fixed-camera infrastructure intelligence, mostly solved or irrelevant.

The next durable robotics company may not have any legs at all. It may just have a very good eye — pointed at a loading dock, all day, every day.

*Ken ZHANG is founder of Maze Intelligence, which maintains a research catalog of 600+ robot models across ~350 companies globally.*