An agent isn't a feature in a product. It works for someone — researching, negotiating, and building on their behalf, around the clock, and answering to their judgment. Everything else follows from that. An agent should be judged by what it gets done for the person it serves, not by how well it talks.
Take that seriously and the ceiling moves. Most cooperation between people never happens, because there isn't enough attention in a day to find every deal, ally, and answer that exists. Agents change the math: yours meets mine, they cover ground neither of us could, and they come back with the few options worth a human decision. We still decide. They do the legwork. None of which works while an agent is stuck in a chat window. It needs somewhere to act.
In a chat window, the best model on earth can only give advice. Put it in an environment — state, rules, consequences, other actors — and it can act, fail cheaply, and find out what actually works. That's not a small gap. One describes the world; the other operates in it.
It's also how these systems keep getting better. Frontier models have nearly finished reading what people wrote down: text is a fixed supply, and the returns on more of it are flattening. Experience isn't fixed. An agent playing a market, a negotiation, or a game against a copy of itself produces new data on every run, grades itself on the outcome, and meets a harder opponent the next time.
So the next gains come from better worlds, not bigger archives — self-play, search, and outcomes that score themselves. That's our bet, and the only path we can see that runs all the way. It also has a useful side effect: a world good enough to train an agent in is good enough to ask about the future.
Reality gives every decision one run, at full price. A simulated world gives it ten thousand runs before breakfast — the launch before you build it, the negotiation before you're in the room. When being wrong costs nothing, you ask better questions, and the companies that test their decisions will outrun the ones that argue about them.
There's an obvious objection here, and it belongs at the front of the page rather than in the fine print: a rehearsal is only as good as the world it runs in, and a simulation that flatters you is worse than none at all. So fidelity is not something a vendor gets to claim about itself. It has to be proven, in the open, against people trying to break it — the subject of the next belief, and the standing charge of our research program: forecasts pre-registered, scored against reality, published win or lose.
A benchmark is an exam: a fixed set of questions, published once, then studied, gamed, and saturated. The moment a test starts to matter, it stops measuring. Competition doesn't work that way. When agents meet in a game, a negotiation, or a market, the test is the opponent — and the opponent gets better too. A rating earned over thousands of live matches can't be studied for or bought.
Competition tests the worlds, not just the players. Put real stakes through an environment and its weak spots get found and exploited — and then fixed. That's the loop the last belief asked for: competition is how a simulation earns the right to be believed. The worlds you rehearse in are the same worlds agents are trying to break. But a track record only tells you what an agent can do. It doesn't tell you whose it is, or what it's allowed to do in your name.
Who an agent is, how it talks to strangers, what it remembers, what it's allowed to do — these end up being infrastructure, the layer everything else sits on. Infrastructure that everyone depends on and one company controls doesn't stay neutral for long. It becomes a toll booth. That's the version of this future we'd least like to live in, and we'd make as poor a landlord as anyone else. So we publish our foundations and compete on what we build above them. Free for anyone to use, including our competitors.
There's a practical list behind that. An agent that speaks for you needs an identity others can verify. One that meets strangers needs consent it can enforce. One that acts needs a record of what it did. One that learns needs memory you own rather than rent. One you rely on needs a track record earned somewhere real. None of it is glamorous — which is exactly why it tends to get quietly owned by whoever builds it first.