# What the harness?

Same protocol. Different behavior.

[Website](https://what-the-harness.polytomic.com/) · [Approved run history](https://what-the-harness.polytomic.com/history.html)

## Test 001: The double payload

### Why it matters

We don’t want to flood the model with duplicate tokens. MCP tools can return the same data as text and structured output; the harness should pass it to the model once.

This test puts a different random marker in each field and asks the agent what it sees. Reporting both markers reveals duplication. Two extra probes check that text-only and structured-only responses still get through.

### What counts as a pass?

| Response sent | PASS | WARN | FAIL |
| --- | --- | --- | --- |
| Both fields | Uses structured output | Falls back to text | Duplicate, missing, or unexpected output |
| Text only | Only text gets through | — | Missing or unexpected output |
| Structured only | Only structured output gets through | — | Missing or unexpected output |

Our rubric prefers one structured payload, avoids duplication, and preserves single-field results. This preference is not a claim about what the MCP standard mandates.

## Observed results

Latest approved run per harness, MCP client and version. Dates are ISO 8601; Received means the run date is unknown and receipt time is shown.

| Harness | MCP client | Version | Date | Both fields | Text only | Structured only | Evidence |
| --- | --- | --- | --- | --- | --- | --- | --- |
| Claude Desktop | Anthropic/ClaudeAI | 1.0.0 | Run: 2026-10-05T20:34:34.480Z | WARN: Falls back to text | PASS | PASS | [Run details](https://what-the-harness.polytomic.com/history.html?run=2339d08c-3ea5-42c3-8900-13d859b27e4a#session-2339d08c-3ea5-42c3-8900-13d859b27e4a) |
| Claude Code (Claude desktop app, Code tab) | claude-code | 2.1.286 | Run: 2026-10-05T20:32:17.625Z | PASS: Uses structured output | PASS | PASS | [Run details](https://what-the-harness.polytomic.com/history.html?run=8f78623b-0668-4079-b0f9-9e1193b669b6#session-8f78623b-0668-4079-b0f9-9e1193b669b6) |
| Amp via Orb | amp-thread-actor | 1 | Run: 2026-10-05T19:53:07.466Z | FAIL: Duplicate output | PASS | PASS | [Run details](https://what-the-harness.polytomic.com/history.html?run=9e507623-3f07-44bd-b325-65c695a6fc95#session-9e507623-3f07-44bd-b325-65c695a6fc95) |
| Cursor Desktop | cursor-vscode | 1.0.0 | Run: 2026-10-05T19:49:09.667Z | WARN: Falls back to text | PASS | FAIL: No output | [Run details](https://what-the-harness.polytomic.com/history.html?run=98183293-0c96-4a22-ada7-837ceae6c4de#session-98183293-0c96-4a22-ada7-837ceae6c4de) |
| pi | pi | 1.0.3 | Run: 2026-10-05T19:46:05.804Z | FAIL: Duplicate output | PASS | PASS | [Run details](https://what-the-harness.polytomic.com/history.html?run=cfdfdec6-cea8-4935-bf09-706953549dc9#session-cfdfdec6-cea8-4935-bf09-706953549dc9) |
| Claude Code CLI | claude-code | 2.1.289 | Run: 2026-10-05T19:42:56.523Z | PASS: Uses structured output | PASS | PASS | [Run details](https://what-the-harness.polytomic.com/history.html?run=e84fa15c-e407-4153-a561-d8e1fb7a1c71#session-e84fa15c-e407-4153-a561-d8e1fb7a1c71) |
| Claude Code (Claude desktop app, Code tab) | Anthropic/ClaudeAI | 1.0.0 | Run: 2026-10-05T19:41:13.009Z | PASS: Uses structured output | PASS | PASS | [Run details](https://what-the-harness.polytomic.com/history.html?run=e62a1c96-5113-4cf4-b0d6-e9eb7221d921#session-e62a1c96-5113-4cf4-b0d6-e9eb7221d921) |
| Codex CLI | codex-mcp-client | 0.160.1 | Run: 2026-10-05T19:37:24.679Z | FAIL: Duplicate output | PASS | PASS | [Run details](https://what-the-harness.polytomic.com/history.html?run=0b394abc-1ed8-4931-aafb-f55df4d88935#session-0b394abc-1ed8-4931-aafb-f55df4d88935) |
| ChatGPT Desktop Codex | codex-mcp-client | 0.160.0 | Run: 2026-10-05T19:35:34.833Z | FAIL: Duplicate output | PASS | PASS | [Run details](https://what-the-harness.polytomic.com/history.html?run=1b49d18c-755b-4d77-be3a-6616b2172613#session-1b49d18c-755b-4d77-be3a-6616b2172613) |
| ChatGPT Desktop Work | codex-mcp-client | 0.160.0 | Run: 2026-10-05T19:33:00.989Z | FAIL: Duplicate output | PASS | PASS | [Run details](https://what-the-harness.polytomic.com/history.html?run=9c1bf2f3-478f-4a8c-843b-098c06d50ee5#session-9c1bf2f3-478f-4a8c-843b-098c06d50ee5) |
| Grokbot | Cursor | 1.0.0 | Run: 2026-10-03T14:15:01.612Z | WARN: Falls back to text | PASS | FAIL: No output | [Run details](https://what-the-harness.polytomic.com/history.html?run=6512292c-e2d6-43a1-ab3b-42fb1b483a6f#session-6512292c-e2d6-43a1-ab3b-42fb1b483a6f) |
| ChatGPT Work Cloud via Web | openai-mcp (Codex) | 1.0.0 | Run: 2026-10-03T14:04:59.630Z | FAIL: Duplicate output | PASS | PASS | [Run details](https://what-the-harness.polytomic.com/history.html?run=67f61d67-a631-45b4-af0f-8d712e9e6a6f#session-67f61d67-a631-45b4-af0f-8d712e9e6a6f) |
| ChatGPT via Web | openai-mcp | 1.0.0 | Run: 2026-10-03T11:00:55.873Z | PASS: Uses structured output | FAIL: No output | PASS | [Run details](https://what-the-harness.polytomic.com/history.html?run=1e79a97b-829f-4fbf-853f-59d62384b6a7#session-1e79a97b-829f-4fbf-853f-59d62384b6a7) |
| Amp CLI | amp-thread-actor | 1 | Run: 2026-10-02T20:30:53.525Z | PASS: Uses structured output | PASS | PASS | [Run details](https://what-the-harness.polytomic.com/history.html?run=7aa8b5f4-9a87-4674-a23a-9d7590d266b6#session-7aa8b5f4-9a87-4674-a23a-9d7590d266b6) |
| Cursor Cloud | Cursor | 1.0.0 | Run: 2026-10-02T20:13:57.747Z | WARN: Falls back to text | PASS | FAIL: No output | [Run details](https://what-the-harness.polytomic.com/history.html?run=49d5652c-541c-4a0b-9c07-b9ac631080b3#session-49d5652c-541c-4a0b-9c07-b9ac631080b3) |
| OpenAI Agents API | openai-mcp (Codex) | 1.0.0 | Run: 2026-10-02T20:11:44.886Z | FAIL: Duplicate output | PASS | PASS | [Run details](https://what-the-harness.polytomic.com/history.html?run=958c9314-27e3-4c4e-a799-54cd14d9a411#session-958c9314-27e3-4c4e-a799-54cd14d9a411) |
| Grok CLI | grok-shell-mcp-output-probe | 1.0.46 | Run: 2026-10-02T19:30:39.115Z | FAIL: Duplicate output | PASS | PASS | [Run details](https://what-the-harness.polytomic.com/history.html?run=32705cd8-fa5b-4858-a408-897c16f5135f#session-32705cd8-fa5b-4858-a408-897c16f5135f) |
| Codex CLI | codex-mcp-client | 0.160.0 | Run: 2026-10-02T19:21:51.545Z | FAIL: Duplicate output | PASS | PASS | [Run details](https://what-the-harness.polytomic.com/history.html?run=369ec41a-9040-4df5-bf0f-f5c1a255a540#session-369ec41a-9040-4df5-bf0f-f5c1a255a540) |
| pi | pi | 1.0.0 | Run: 2026-10-02T19:21:44.456Z | FAIL: Duplicate output | PASS | PASS | [Run details](https://what-the-harness.polytomic.com/history.html?run=e8f1b7b4-abd9-4a6e-8823-1c92cb45a6ca#session-e8f1b7b4-abd9-4a6e-8823-1c92cb45a6ca) |
| Cursor app | cursor-vscode | 1.0.0 | Run: 2026-10-02T19:19:01.600Z | WARN: Falls back to text | PASS | FAIL: No output | [Run details](https://what-the-harness.polytomic.com/history.html?run=b950a53b-3584-4e4d-b9ca-66584834813e#session-b950a53b-3584-4e4d-b9ca-66584834813e) |
| ChatGPT Desktop | openai-mcp | 1.0.0 | Run: 2026-10-02T19:17:01.183Z | PASS: Uses structured output | PASS | PASS | [Run details](https://what-the-harness.polytomic.com/history.html?run=611c36a9-c33a-4f96-a8a7-117d549cd7d6#session-611c36a9-c33a-4f96-a8a7-117d549cd7d6) |
| Claude Code | claude-code | 2.1.287 | Run: 2026-10-02T19:07:10.916Z | PASS: Uses structured output | PASS | PASS | [Run details](https://what-the-harness.polytomic.com/history.html?run=4dd33f28-92c5-46f9-9814-6ce151352911#session-4dd33f28-92c5-46f9-9814-6ce151352911) |
| Codex Desktop | codex-mcp-client | 0.159.0-alpha.12.1 | Run: 2026-10-02T18:56:24.661Z | FAIL: Duplicate output | PASS | PASS | [Run details](https://what-the-harness.polytomic.com/history.html?run=4317a024-2061-4e69-ae27-7f526f25010a#session-4317a024-2061-4e69-ae27-7f526f25010a) |

Verdicts use reported markers. Observations describe individual sessions, not inferred model inputs. Harness names use the reviewer-assigned name when present, otherwise the agent-reported name. Client names and versions are declared by the MCP connection and may identify a connector rather than the host application.

## Test 002: Protocol version

Recorded automatically at the start of a run. Legacy connections show the negotiated version; modern connections show the request protocol. This is the version used in the session, not the client’s highest version or complete supported set.

| Harness / MCP client | Protocol used | Evidence |
| --- | --- | --- |
| Claude Desktop / Anthropic/ClaudeAI 1.0.0 | 2026-07-28: Latest | [Run details](https://what-the-harness.polytomic.com/history.html?run=2339d08c-3ea5-42c3-8900-13d859b27e4a) |
| Claude Code (Claude desktop app, Code tab) / claude-code 2.1.286 | 2026-07-28: Latest | [Run details](https://what-the-harness.polytomic.com/history.html?run=8f78623b-0668-4079-b0f9-9e1193b669b6) |
| Amp via Orb / amp-thread-actor 1 | 2025-06-18: Older protocol | [Run details](https://what-the-harness.polytomic.com/history.html?run=9e507623-3f07-44bd-b325-65c695a6fc95) |
| Cursor Desktop / cursor-vscode 1.0.0 | 2025-11-25: Older protocol | [Run details](https://what-the-harness.polytomic.com/history.html?run=98183293-0c96-4a22-ada7-837ceae6c4de) |
| pi / pi 1.0.3 | 2025-11-25: Older protocol | [Run details](https://what-the-harness.polytomic.com/history.html?run=cfdfdec6-cea8-4935-bf09-706953549dc9) |
| Claude Code CLI / claude-code 2.1.289 | 2026-07-28: Latest | [Run details](https://what-the-harness.polytomic.com/history.html?run=e84fa15c-e407-4153-a561-d8e1fb7a1c71) |
| Claude Code (Claude desktop app, Code tab) / Anthropic/ClaudeAI 1.0.0 | 2026-07-28: Latest | [Run details](https://what-the-harness.polytomic.com/history.html?run=e62a1c96-5113-4cf4-b0d6-e9eb7221d921) |
| Codex CLI / codex-mcp-client 0.160.1 | 2025-06-18: Older protocol | [Run details](https://what-the-harness.polytomic.com/history.html?run=0b394abc-1ed8-4931-aafb-f55df4d88935) |
| ChatGPT Desktop Codex / codex-mcp-client 0.160.0 | 2025-06-18: Older protocol | [Run details](https://what-the-harness.polytomic.com/history.html?run=1b49d18c-755b-4d77-be3a-6616b2172613) |
| ChatGPT Desktop Work / codex-mcp-client 0.160.0 | 2025-06-18: Older protocol | [Run details](https://what-the-harness.polytomic.com/history.html?run=9c1bf2f3-478f-4a8c-843b-098c06d50ee5) |

## Test 003: URL elicitation

Declared support is recorded automatically. Optional verification requires a human to complete the browser step and the probe to resume successfully. Advertisement alone is unverified. A client may advertise support while using an older protocol that prevents verification. Browser completion without a recorded retry is incomplete.

| Harness / MCP client | URL support advertised | Interactive verification | Evidence |
| --- | --- | --- | --- |
| Claude Desktop / Anthropic/ClaudeAI 1.0.0 | No | Skipped | [Run details](https://what-the-harness.polytomic.com/history.html?run=2339d08c-3ea5-42c3-8900-13d859b27e4a) |
| Claude Code (Claude desktop app, Code tab) / claude-code 2.1.286 | Yes | Declined | [Run details](https://what-the-harness.polytomic.com/history.html?run=8f78623b-0668-4079-b0f9-9e1193b669b6) |
| Amp via Orb / amp-thread-actor 1 | No | Skipped | [Run details](https://what-the-harness.polytomic.com/history.html?run=9e507623-3f07-44bd-b325-65c695a6fc95) |
| Cursor Desktop / cursor-vscode 1.0.0 | No | Skipped | [Run details](https://what-the-harness.polytomic.com/history.html?run=98183293-0c96-4a22-ada7-837ceae6c4de) |
| pi / pi 1.0.3 | No | Skipped | [Run details](https://what-the-harness.polytomic.com/history.html?run=cfdfdec6-cea8-4935-bf09-706953549dc9) |
| Claude Code CLI / claude-code 2.1.289 | Yes | Verified | [Run details](https://what-the-harness.polytomic.com/history.html?run=e84fa15c-e407-4153-a561-d8e1fb7a1c71) |
| Claude Code (Claude desktop app, Code tab) / Anthropic/ClaudeAI 1.0.0 | No | Skipped | [Run details](https://what-the-harness.polytomic.com/history.html?run=e62a1c96-5113-4cf4-b0d6-e9eb7221d921) |
| Codex CLI / codex-mcp-client 0.160.1 | Yes | Skipped: requires newer protocol | [Run details](https://what-the-harness.polytomic.com/history.html?run=0b394abc-1ed8-4931-aafb-f55df4d88935) |
| ChatGPT Desktop Codex / codex-mcp-client 0.160.0 | Yes | Skipped: requires newer protocol | [Run details](https://what-the-harness.polytomic.com/history.html?run=1b49d18c-755b-4d77-be3a-6616b2172613) |
| ChatGPT Desktop Work / codex-mcp-client 0.160.0 | Yes | Skipped: requires newer protocol | [Run details](https://what-the-harness.polytomic.com/history.html?run=9c1bf2f3-478f-4a8c-843b-098c06d50ee5) |

## Test 004: Resource reading

Read the direct per-run resource and a second resource returned as a tool resource_link. Read successfully means a recorded MCP resource read and an exact marker report. Not read or a missing marker is incomplete; an incorrect marker fails. Discovery is available through a resource guide and URI template; this verdict measures reading and visibility, not autonomous discovery.

| Harness / MCP client | Direct resource | Tool resource link | Evidence |
| --- | --- | --- | --- |
| Claude Desktop / Anthropic/ClaudeAI | Read successfully | Read successfully | [Run details](https://what-the-harness.polytomic.com/history.html?run=2339d08c-3ea5-42c3-8900-13d859b27e4a) |
| Claude Code (Claude desktop app, Code tab) / claude-code | Read successfully | Read successfully | [Run details](https://what-the-harness.polytomic.com/history.html?run=8f78623b-0668-4079-b0f9-9e1193b669b6) |
| Amp via Orb / amp-thread-actor | Not read | Not read | [Run details](https://what-the-harness.polytomic.com/history.html?run=9e507623-3f07-44bd-b325-65c695a6fc95) |
| Cursor Desktop / cursor-vscode | Read successfully | Read successfully | [Run details](https://what-the-harness.polytomic.com/history.html?run=98183293-0c96-4a22-ada7-837ceae6c4de) |
| pi / pi | Read successfully | Read successfully | [Run details](https://what-the-harness.polytomic.com/history.html?run=cfdfdec6-cea8-4935-bf09-706953549dc9) |
| Claude Code CLI / claude-code | Read successfully | Read successfully | [Run details](https://what-the-harness.polytomic.com/history.html?run=e84fa15c-e407-4153-a561-d8e1fb7a1c71) |
| Claude Code (Claude desktop app, Code tab) / Anthropic/ClaudeAI | Not read | Not read | [Run details](https://what-the-harness.polytomic.com/history.html?run=e62a1c96-5113-4cf4-b0d6-e9eb7221d921) |
| Codex CLI / codex-mcp-client | Read successfully | Read successfully | [Run details](https://what-the-harness.polytomic.com/history.html?run=0b394abc-1ed8-4931-aafb-f55df4d88935) |
| ChatGPT Desktop Codex / codex-mcp-client | Read successfully | Read successfully | [Run details](https://what-the-harness.polytomic.com/history.html?run=1b49d18c-755b-4d77-be3a-6616b2172613) |
| ChatGPT Desktop Work / codex-mcp-client | Read successfully | Read successfully | [Run details](https://what-the-harness.polytomic.com/history.html?run=9c1bf2f3-478f-4a8c-843b-098c06d50ee5) |

## Run it yourself

Connect your agent to https://what-the-harness.polytomic.com/mcp as a public Streamable HTTP MCP server. No client sign-in or access token is required.

[Connection guide](https://what-the-harness.polytomic.com/submission-guide.html)

### Test prompt

Use the What the Harness MCP server. Call begin_run with verify_url_elicitation: true and test_resources: true; if its schema requires begin_key, generate a fresh UUID for this run, keep it private, and reuse it for start retries. Supply the harness and model only if explicitly known, otherwise leave them null. Save the returned run_id, run_token, and client-declared identity. Keep resource URIs and run_token private; omit it from the final report and evidence. Call probe_both, probe_text_only, and probe_structured_only once each with that run_id and run_token. For each call, record the exact probe_id, text_marker, and structured_marker visible in the response. Use null for every value not visible; do not infer or invent markers. Discover MCP resources and resource templates if the client makes them available. Read the private resources.direct URI from begin_run through MCP resources/read and record its exact resource_marker. Call probe_resource_link and read the URI in its resource_link through MCP; record that resource_marker separately. If a resource cannot be read or its marker is not visible, report null. Never fetch these URIs over HTTP, inspect server files, or infer markers. Call probe_url_elicitation with the same run_id and run_token. If it presents a URL interaction, pause for the human to complete it through the client UI; do not fetch or complete the page yourself. After the human completes it, retry probe_url_elicitation with the same credentials until it returns the verification result. If support is not advertised or the protocol is unsupported, record the skip. If the human declines or cancels, record that outcome. Call submit_result with the run_id, run_token, and all three output reports, plus resource_reports containing direct and linked markers (or null), using those exact field names. If submission times out, retry with the same run_id, run_token, and unchanged reports. If the MCP connection reconnects, keep using the original run_id and run_token. Call result_status with the same credentials. Report the receipt, publication state, protocol version, declared URL support, verification status, and the output and resource markers you saw; say “not visible” for null values. Do not fetch the endpoint or website separately, and do not start another run to recover a missing marker.

## Method and limitations

Independent random markers identify text and structured output. The three probes are probe_both, probe_text_only, and probe_structured_only. Exact reproduction demonstrates visibility; a missing report does not establish the precise model input. Harness and model identity are agent-reported. Unknown values stay unknown. Results are reviewed before publication. Exact token overhead has not been measured.

[Diagnostic probe source](https://what-the-harness.polytomic.com/evidence/probe-source.txt)
