An interview with Hantz Févry

Spatial intelligence promises to move enterprise AI beyond interpreting documents and images toward modeling the physical world as it changes over time. In this Serious Insights interview, Geolava CEO Hantz Févry explains why the real value lies not in photorealistic 3D scenes or object detection, but in systems that can make falsifiable predictions about assets, risks, and operational consequences that range from deteriorating infrastructure to logistics performance and insurance exposure.
Top 3 Takeaways
- Spatial intelligence is predictive, not merely descriptive. A GIS layer, digital twin, or vision system can identify what and where something is; a true world model estimates how physical conditions will evolve in response to time, events, and decisions.
- The enterprise opportunity is detecting “drift” between records and reality. Physical assets change while inspections, drawings, reports, and other enterprise records age, creating hidden maintenance, compliance, and risk exposures that spatial systems could surface.
- Buyers should evaluate outcomes, not demos. Févry recommends historical backtesting, calibrated confidence measures, resilience to sparse data, temporal consistency, and pricing tied to verified decisions rather than seats or visual fidelity.
The Hantz Févry Interview
Spatial intelligence is being framed as one of AI’s next frontiers. What does that phrase mean in practical terms, and where does it risk becoming another overused AI label?
In practical terms, it means AI that models where things are, how they relate in space, and how they change over time. Not describing the world, computing it. The test I apply: can the system predict a physical outcome you can later verify? If yes, it is spatial intelligence. If it just displays locations or labels objects in imagery, it is a map with better marketing. The label is already drifting. Anything that touches coordinates now calls itself spatial AI. Geocoding a spreadsheet is not intelligence.
Large language models have become very good at manipulating symbols. What breaks when symbolic reasoning gets applied to environments governed by geometry, physics, time, and uncertainty?
Language has no penalty for being physically impossible. A sentence about water flowing uphill is grammatically perfect. Three things break specifically. Geometry: symbols do not encode adjacency, occlusion, or load paths, so the model cannot tell you what a wall does. Time: physical state evolves continuously, and text is a snapshot of whenever someone bothered to write. Uncertainty: in the physical world errors compound through causal chains, and a model with no causal structure cannot bound its own error. LLMs have read every physics textbook and never watched anything fall.
Many enterprise AI use cases still live inside documents, chats, dashboards, and workflows. What changes when AI starts from the built environment rather than from text?
The direction of trust inverts. Today enterprises treat documents as ground truth and the world as unknown. But documents lag reality: the inspection report is two years old, the rent roll is aspirational, the as-built drawings are neither. When AI starts from the physical asset, documents become claims to verify against observed state rather than the other way around. That is a different product category. You stop asking “what does the file say” and start asking “what is actually true, and what will be true next year.”
What makes a world model different from a richer data model, digital twin, GIS layer, or computer vision system?
Those are all descriptions. GIS tells you where things are. A digital twin tells you what things are. Computer vision tells you what things look like. A world model predicts: given this state and this action, what is the next state? Defer the roof repair, water finds the sheathing. It is the difference between a photograph of a chessboard and knowing the legal moves. Dynamics, not inventory. A digital twin without dynamics is an expensive screenshot.
Where are today’s spatial AI systems most brittle? Is the hard problem perception, prediction, grounding, reasoning, data quality, or something else?
Perception is largely solved for common cases. The brittleness is in temporal grounding: tying observations captured at different times, resolutions, and angles to one persistent object and its trajectory. Knowing this facade in 2021 imagery and that facade in a 2025 drone pass are the same wall, and estimating what happened in between. Sparse, irregular observation is the norm outside of robotics labs. The systems that win will be the ones that predict well from incomplete data and know how confident to be, not the ones that need perfect capture.
In real estate, construction, infrastructure, and logistics, decisions often depend on messy, incomplete, or aging physical-world data. What does spatial intelligence make visible that current enterprise systems miss?
Drift. Every enterprise system records last known state. Spatial intelligence estimates current state and projects future state, which surfaces the gap between the two. Deferred maintenance that never made a report. Unpermitted modifications. Deterioration between inspection cycles. Risk exposure that changed because the environment changed, not the asset. In my industry, that drift is where losses live. The record says one thing, the building says another, and until now nobody was systematically measuring the difference.
What should business leaders understand about the difference between recognizing an object in an image and reasoning about how that object functions in a physical environment?
Recognition assigns a label. Reasoning assigns consequences. A model that detects “crack” has done recognition. Knowing that this crack sits near a load path, that it has widened since the last observation, and that freeze-thaw cycles will accelerate it: that is function. Same pixels, entirely different decision value. Leaders should be skeptical of accuracy numbers on detection benchmarks. Detecting objects at 99 percent means little if the system cannot say what any of them imply.
How should organizations evaluate spatial intelligence systems? What are the useful measures of performance beyond demo quality or visual impressiveness?
Demand falsifiable predictions and backtests. Have the vendor predict outcomes on your portfolio for a historical period where you know what happened, then score them. Measure calibration, not just accuracy: a system that says 70 percent should be right 70 percent of the time. Test degradation on sparse inputs, because your data is worse than the demo data. Check temporal consistency: does the model give coherent answers about the same asset across time? And price it per verified decision, not per seat. Visual fidelity is the easiest thing to fake and the least correlated with any of this.
Fei-Fei Li and World Labs have put a large spotlight on 3D understanding and world models. What does that funding signal validate, and where do you think the market narrative is getting ahead of the technology?
It validates the core thesis: the next frontier is models of the world, not models of text, and serious researchers are staking careers on it. Where the narrative runs ahead is conflating generative quality with predictive utility. Synthesizing a beautiful explorable 3D scene and accurately predicting what happens to a real asset are different problems, and the second one is what enterprises pay for. Photorealism is progress on rendering. Enterprises need progress on consequences. The market is currently pricing the demo, not the decision.
Looking three to five years out, which industries will likely lead in spatial intelligence, and any thoughts on industries that should be leaders but will get in their own way?
Leaders: insurance, because pricing physical risk is their entire business and the feedback loop is brutal and fast. Logistics, because prediction converts directly to margin. Defense and critical infrastructure, because the budget and urgency exist. Real estate finance will follow once one lender demonstrates loss avoidance.
Getting in their own way: construction, which has the most to gain and the most fragmented incentives, with every party hoarding data and nobody owning the outcome. And parts of commercial real estate, where some participants benefit from opacity. When not knowing is a feature of your business model, a technology that makes the physical world legible arrives as a threat before it arrives as a tool.
About Hantz Févry
CEO, Geolava

Hantz Févry is a former Google Technical Product Manager who led AI products and research with his team, now integrated into DeepMind. He specializes in building deeply technical systems that solve real-world problems at scale. At Geolava, he is building AI systems that can perceive, reason, and act in physical environments. To do so, he is building what will be the next frontier in the AI industry: a world model.
For more serious insights on AI, click here.
Did you find this interview with Hantz Févry useful? If so, please like, share or comment. Thank you!

Leave a Reply