The last decade of AI ran on text, because text had already been written down. The physical world has not been written down. That absence is the single largest gap in the training data of every model now being pointed at the real world, and it cannot be closed by scraping.
Language models became possible because humanity had spent thirty years typing the internet into existence. The corpus was lying there, waiting.
There is no equivalent corpus for physical behavior. No one has written down how a person approaches a display, slows, hesitates, glances at a price, walks on, and returns four minutes later. That sequence happens hundreds of millions of times a day and is recorded essentially nowhere.
You cannot scrape it, because it was never published. You cannot buy it, because nobody assembled it. It exists only if someone deploys sensing in a real space and captures it as it happens.
A well-funded line of work argues that you can generate behavioral data instead: build agents grounded in interviews and transactions, then simulate how people would react.
This is genuinely useful for exploring hypotheticals, and it has a hard limit. Simulation reproduces the assumptions in its priors. It can tell you what a model of a person would do, which is a different object from what people did. Anything genuinely surprising, the behavior nobody would have predicted, is precisely what a simulator cannot produce, because it was not in the priors to begin with.
Simulated behavior also has no ground truth to be wrong against. Without sensed reality to calibrate on, a behavioral simulator can drift indefinitely while remaining internally coherent. These approaches are complements, not substitutes: simulation needs sensing to stay honest.
Not all movement data is training data. Four properties determine whether a dataset can support a model rather than a dashboard:
Behavioral ground truth has a collection problem: it is slow, physical, and cannot be parallelised by buying more compute. So the question becomes where to collect it fastest.
Events are unusually good for this, for reasons that have nothing to do with events being a large market:
The result is a labelled, high-intent behavioral dataset accumulating far faster than the same effort would produce anywhere else. That is why we started here, and why events are a training ground rather than a destination.
The immediate output is commercial: ranked leads, zone performance, layout decisions, pricing evidence. That is what funds the collection, and it should. A dataset that does not pay for itself never reaches useful scale.
The longer arc is a model that has seen enough physical behavior to generalise. Not a system that reports what happened in a space it has already measured, but one that can predict how people will move through a space it has never seen, from the floor plan alone.
Whether that generalisation holds is an empirical question and we would be overclaiming to call it settled. What is already clear is the ordering: the model cannot exist before the dataset, and the dataset cannot exist before someone deploys sensing in enough real rooms. The bottleneck is physical, which is exactly why it is defensible.
If behavioral ground truth can only be collected by deployment, then whoever deploys first accumulates an asset that later entrants cannot shortcut. Not because the technology is secret. Sensing hardware is a commodity and the interpretation methods are learnable, but because time spent in real rooms is not compressible.
That is an unfashionable kind of moat in a software industry accustomed to instant scale. It is also, at the moment, one of the few that still holds.
Physical AI refers to systems that understand and act in the real world rather than purely in text or images, reasoning about where things are, how people and objects move, and how a space changes over time. It requires continuous real-world sensing, because unlike language, physical behavior was never written down and therefore cannot be scraped from the internet.
Behavioral ground truth is directly sensed data about how people actually moved and behaved in a real physical space, as opposed to survey responses, app taps or simulated behavior. To be usable for training it needs continuity, sub-metre resolution, interpretation labels defining what counts as approach or engagement, and a defensible consent basis.
Simulation reproduces the assumptions built into its priors, so it can describe what a model of a person would do but not what people actually did. It is valuable for exploring hypotheticals and cannot generate genuinely unexpected behavior, because the unexpected was not in the priors. Simulation and sensing are complementary: simulators need sensed ground truth to calibrate against.
Because the signal is cleaner and the collection cycle is faster. At an event, presence implies intent, because people travelled and paid to attend, while retail footfall carries more ambient browsing noise. Events also have a registration layer that makes consented identity linking possible, a hard spatial boundary that makes complete journeys observable, and a deployment cycle measured in days.
The fact that it can only be produced by physically deploying sensing in real spaces over time. The hardware is commodity and the methods are learnable, but time spent capturing real behavior in real rooms cannot be compressed or purchased, which means an early lead widens rather than erodes.