When you need a sandbox
Early on you can skip it: run the agent by hand, read what it did, form an opinion. For a while that is enough. It stops being enough at the point Anthropic calls flying blind, where someone says the agent feels worse since the change and there is no way to settle it except to guess. What breaks first is comparison. If the world your agent worked in was different this time, a score that moved tells you nothing. A difference only means something when every trial starts from the same declared state. What breaks next is coverage. The destructive half of the job, the refund and the delete and the reply-all, is the half nobody dares run twice, so it stays untested where it matters most. And you cannot settle it from the agent’s own write-up, because saying the work is done is the easiest part of doing it. Pome sits under all three. It does not write your agent, pick your model or decide what good looks like. It supplies a world that resets to a state you declared, and the record of what happened in it.Where your agent acts
A digital twin is a stateful emulation of one real API, backed by a real database, answering the same REST, GraphQL and MCP calls as production. Your agent reads, decides, writes, then reads back what the write did. Nothing it does reaches a live account. Pome runs one twin per API it covers, across issue tracking, payments, team chat and mail, and Digital twins has the current list. A sandbox is what you start: one id, one lifetime, and the twins a single run talks to, booted from a starting state you declare and thrown away after. Every run gets its own, so no trial inherits state from the one before it. How closely each twin matches the real API is measured daily, surface by surface, and published at status.pome.sh. That measurement is the verified in the title.What comes back
A run leaves two things behind, and what makes them worth having is that the agent did not write either one. The tape is the twin’s own log: every call it served, in order, with the method, the path, the status, and whether the write changed anything. That last one is not redundant. Slack answers200 to a write it refused, so an agent
reading status codes reports a message it never posted without lying about it.
On the tape those two calls do not look alike, because the twin knows which one
moved its state.
The second is that state itself, exported when the run ends. Together they
answer a different question from the agent’s own account of the run: not what it
believed it did, but what happened to the world it was working in.
How a run is scored
A criterion is a question about that world, and the grader answers it from the record rather than from the write-up. Did a second issue get opened for the same bug? Is there exactly one refund against that charge? Code runs those against the tape and the final state and gets the same answer every time, which is the only reason two runs are comparable at all. Checking state is also what keeps a confident agent honest. An agent can finish a run reporting a refund that exists nowhere, and the ledger is what answers. Because the check read a particular row to decide, the verdict comes back holding it, so a failure is something you open rather than something you take on trust. Deterministic checks buy that at a price worth naming: they only see what you can put as a question about state. Whether the agent explained itself, or noticed that a label was wrong for the situation, is not a query. The narrator, which is the LLM in the loop, reads those and answers them in prose, and its reading stays out of the number. A score a model can move is a score that moves on its own. How Pome works takes one run apart, step by step.What stays yours
The thing under test is your agent harness and your model working together, and Pome changes neither. Your prompt, your tools and your framework stay as they are. The one difference is where the agent’s calls land. That is also why none of this is a migration. The tape is a file you can read, not a dashboard you have to move into. Use your existing eval stack reads the same tape into Braintrust or LangSmith, and Pome’s score arrives as one more column beside what you already track.Start here
Get started
About five minutes, and no account. One command runs a twin on your
machine, your own coding agent does a real piece of work in it, and you read
back the tape that recorded it. Nothing is graded.