Skip to main content

Capstone run failed?

The rest of this page assumes the CLI. If you came through the graded capstone, you likely don’t have it installed. Debug from the conversation instead:
  • Ask the coach for the report again: get_report(run_id) names each criterion, its verdict, and the evidence behind it.
  • finalize_run errored. Read the error text first. If it says the sandbox is still live, the failure is retryable (a capture or transient server error) and the sandbox is preserved: call finalize_run again with the same session_id instead of re-running anything. Only an expired sandbox is unrecoverable — it cannot be scored, so re-run steps 5–8 of the quickstart prompt and finalize the moment the examinee exits.
  • Pome tools missing. The MCP is not connected, or its OAuth grant has lapsed. Connect to the MCP carries the wiring for each client, how to verify a live connection, and what a revoked grant looks like.
  • The dashboard link in every report shows the full trace and the run’s handoff.

Debugging with the CLI

Start with pome inspect latest. It usually points at the layer that failed: task, agent, twin, or grading.

Agent crashed

Most common causes:
  • Agent did not no-op on POME_PREFLIGHT=1.
  • Agent ignored POME_GITHUB_REST_URL and hit the real internet.
  • Task prompt referenced state the agent could not parse.

Score below 100

If a call has fidelity: "unsupported", the agent reached for an endpoint the twin does not implement yet. File an issue at github.com/pome-sh/digital-twins.

Twin will not start

A port is stuck. pome twin start names the port it could not bind, so pick another one:
To check whether the twin is alive, ask the CLI:
It probes the twin the last pome twin start recorded and requires the health response to name that twin, so another dev server on the port reads as not running instead of as a green. The unauthenticated root health endpoint is the lower-level fallback — reach for it when you want the raw payload, or when the twin was not started by pome twin start:
The session-scoped /s/<sid>/healthz requires the JWT and returns 401 without it.

Hosted run hangs or 401s

Re-mint your key:
Or set POME_API_KEY directly in CI.

The hosted upload rejects my events.jsonl schema

The hosted upload requires the discriminated-union RecorderEvent schema. Upgrade the CLI:
Then re-run.

pome fix-prompt says “failed outright” for a failed hosted run

pome fix-prompt reads the verdicts a hosted run persisted as verdict.json. A run set with no verdict.json — a --local capture, or a run that never finalized — has nothing for it to group.
  • For a --local capture, run pome eval <run-dir> to get a verdict, then ask again.
  • For a run set where nothing failed outright but a trial is INCOMPLETE, pome fix-prompt says so and names the trial: the grader did not finish, so a prompt built from it would claim more than was checked. Re-run the task to grade the rest.
  • Either way, the run URL printed by pome run carries the narrator’s handoff.

pome login says “already logged in” but I do not see ~/.pome/credentials.json

On macOS, pome stores the key in the Keychain by default:
On Linux/Windows, or when the Keychain is unavailable, the key falls back to ~/.pome/credentials.json. Both paths are valid; the CLI tries the env var, then the Keychain, then the JSON file.

pome run exited 1 and I want to know which kind of 1

Exit 1 covers two outcomes: a task scored below its threshold, and a task the grader could not finish (INCOMPLETE). Read the printed verdict word to tell them apart, or open the run URL, where the score names its denominator. CLI reference has the full table.

pome sandbox list returns more rows than the dashboard

The CLI defaults to --state running, but if you have ever passed --state all, that selection persists in your shell history. The dashboard’s Twins page defaults to running sandboxes. To match it:

My agent’s MCP transport fails against a hosted sandbox

pome sandbox create prints two lines per twin, and they differ by the /mcp suffix:
Do not append /mcp yourself. The CLI already emits the suffix on the MCP line, so appending a second one produces .../mcp/mcp, which is not a route the twin serves. Older copies of this page told you to append it; that advice is retired. Two things to check instead:
  • You mounted the API line. They differ by four characters and are one line apart. An agent handed the REST root as an MCP transport gets REST responses and never negotiates the handshake. Use the MCP line, or POME_<TWIN>_MCP_URL from --secrets-file — that variable already carries the suffix.
  • The bearer is missing. Every call to a hosted twin needs Authorization: Bearer $POME_AUTH_TOKEN. Without it the handshake fails on auth rather than on transport, which reads the same way from most clients.
On a multi-twin sandbox each twin gets its own pair, prefixed by twin id — the sandbox has one id and one bearer, but no shared MCP endpoint. See pome sandbox.

pome doctor is red and I want to know what to change

pome doctor runs four checks in order and stops at the first red. pome run runs the same four itself and refuses to start when one fails — there is no --force, because a run whose agent quietly reached the real API produces a trace that looks complete and grades nothing. Exit 0 means every applicable check passed; exit 1 means one failed and the printed cause names which.