About 5 minutes. You start a private Slack workspace
with two channels and an unanswered morning message, your own coding agent
answers inside that thread, and you read the twin’s own record of every call it
made — including the one column that says whether the message actually landed.
slack.com/api, boots from a declared starting state,
records every request, and never touches Slack. Reply in a thread here and the
parent message’s reply count really moves — inside this sandbox, and nowhere
else.
Before you start
- A Pome account.
pome loginopens the browser sign-in and creates one if you do not have it. No credit card. - Your own coding agent — Claude Code, Cursor, anything that can run a shell command and read JSON back.
- Node 18+ for
npx, pluscurlandjqfor the transcripts below. Your agent can read the raw JSON withoutjq; it is here to keep the blocks short.
ANTHROPIC_API_KEY, no model inference
paid for by Pome: your agent is both the operator and the actor. It drives
Pome, and it is the thing that acts on the twin.
Nothing on this page is graded. No task file, no criteria, no score — and no
agent eval is charged, because an eval is only ever burned when a run is graded.
A sandbox you start, drive and stop costs you nothing. Grading appears exactly
once in this curriculum, at the
support-triage capstone, where the agent under test
is sealed off from the criteria that judge it.
Paste this
Hand this to your coding agent as-is. It names the twin’s own Web API surface and the boundary it must not cross.The world
sandbox create boots a Slack twin from its declared starting state and hands
back the URLs that reach it. It also writes the connection secrets to
.pome-sandbox.env at mode 0600, and says so on stderr.
api_url with the one bearer the sandbox accepts.
Note the bare method paths: this twin mounts conversations.list, not
api/conversations.list, and a wrong prefix comes back 501 with every served
surface listed in the body.
#random untouched. Read those ts values,
never hardcode them. The seed fixes the content of this world, not its
identifiers: message timestamps are minted when the sandbox boots, so yours
differ from the ones above and from your last run. Capture the one you need:
replies=null before
the write and reads 1 after it, and the reply comes back nested under her ts
rather than sitting beside it in the channel. A mock has no parent to update.
Stop the sandbox and that thread is gone; create another and #general is two
unanswered greetings again.
Read the tape
Every call above was recorded by the twin as it happened. This is the part neither a mock nor a staging workspace gives you: an account of the run written by the service, not by the agent.200. That is not the recording being lazy —
it is Slack’s actual contract, and it is why the request line alone cannot tell
you what happened. Ask for the column that can:
- Six rows, in the order they happened. Read top to bottom and the exchange is legible without asking the agent what it did: it looked twice, it wrote twice, it checked.
mut=trueon exactly the two writes.state_mutationmeans the call landed. On this twin that column is the whole verdict, because Slack answers200to a refused write as readily as an accepted one.tool=nullon every row. Unlike the GitHub twin, this one stamps no action vocabulary; calls are identified by method and path, which is why the assertable checks below read final state rather than the tape.fid=semantic. These surfaces carry a full behavioural contract, not a response shape with placeholder values.
200. The tape is what says nothing was written. When you
are done, stop the sandbox to free the slot; it also expires on its own 30
minutes after it was created:
What you could assert here
Nothing on this page was graded, but a tape and a final state are exactly what a graded check reads. The Slack twin already declares these — no authoring required, andlist_checks on the Pome MCP prints the full set with what each
one actually compares:
Pointers, not a task. Turning them into a graded exam is the
capstone below, and
Write a task is where the grammar lives.
How real is this twin?
Every surface you called above is compared against a captured response from the real Slack Web API, re-run daily and published. The Slack row on status.pome.sh carries the current count of twin responses that match, how many of Slack’s own MCP tools are covered, and how old the captured baseline is. Nobody has to take our word for the fidelity, including us — a twin that drifts turns that row red. Which surfaces are covered, and which are shape-only, is on the Slack twin reference.Next: the one graded lesson
The support-triage capstone
The same twins, now an exam: a sealed agent under test, a deliberate failing
score, and one line of prompt that turns it green.
Slack twin reference
Every method and MCP tool the twin serves, by use case, with its fidelity tier.
Write a task
Turn the checks above into a graded exam for your own agent.
pome sandbox
Create, list and stop sandboxes — including multi-twin ones.