Test 001: The double payload
MCP outputDoes your harness pass tool output once, or twice?
Why it matters
We don’t want to flood the model with duplicate tokens. MCP tools can return the same data as text and structured output; the harness should pass it to the model once.
This test puts a different random marker in each field and asks the agent what it sees. Reporting both markers reveals duplication. Two extra probes check that text-only and structured-only responses still get through.
What counts as a pass?
Both fields sent
PASS: uses structured output
WARN: falls back to text
FAIL: duplicate, missing, or unexpected output
Text only sent
PASS: only text gets through
FAIL: missing or unexpected output
Structured only sent
PASS: only structured output gets through
FAIL: missing or unexpected output
Our rubric: prefer one structured payload, avoid duplication, preserve single-field results. This preference is not a claim about what the MCP standard mandates.
Latest approved run per harness, MCP client and version
Scroll the table sideways to compare all three responses →
| Harness | Response sent by the probe | Evidence | ||
|---|---|---|---|---|
| Both fieldscontent + structuredContent | Text onlycontent | Structured onlystructuredContent | ||
| Amp via OrbMCP client: amp-thread-actorVersion 1Run: Oct 5, 2026 | FAILDuplicate output | PASS | PASS | View ↗ |
| Cursor DesktopMCP client: cursor-vscodeVersion 1.0.0Run: Oct 5, 2026 | WARNFalls back to text | PASS | FAILNo output | View ↗ |
| piMCP client: piVersion 1.0.3Run: Oct 5, 2026 | FAILDuplicate output | PASS | PASS | View ↗ |
| Claude Code CLIMCP client: claude-codeVersion 2.1.289Run: Oct 5, 2026 | PASSUses structured output | PASS | PASS | View ↗ |
| Claude Code (Claude desktop app, Code tab)MCP client: Anthropic/ClaudeAIVersion 1.0.0Run: Oct 5, 2026 | PASSUses structured output | PASS | PASS | View ↗ |
| Claude DesktopMCP client: Anthropic/ClaudeAIVersion 1.0.0Run: Oct 5, 2026 | WARNFalls back to text | PASS | PASS | View ↗ |
| Codex CLIMCP client: codex-mcp-clientVersion 0.160.1Run: Oct 5, 2026 | FAILDuplicate output | PASS | PASS | View ↗ |
| ChatGPT Desktop CodexMCP client: codex-mcp-clientVersion 0.160.0Run: Oct 5, 2026 | FAILDuplicate output | PASS | PASS | View ↗ |
| ChatGPT Desktop WorkMCP client: codex-mcp-clientVersion 0.160.0Run: Oct 5, 2026 | FAILDuplicate output | PASS | PASS | View ↗ |
| GrokbotMCP client: CursorVersion 1.0.0Run: Oct 3, 2026 | WARNFalls back to text | PASS | FAILNo output | View ↗ |
| ChatGPT Work Cloud via WebMCP client: openai-mcp (Codex)Version 1.0.0Run: Oct 3, 2026 | FAILDuplicate output | PASS | PASS | View ↗ |
| ChatGPT via WebMCP client: openai-mcpVersion 1.0.0Run: Oct 3, 2026 | PASSUses structured output | FAILNo output | PASS | View ↗ |
| Amp CLIMCP client: amp-thread-actorVersion 1Run: Oct 2, 2026 | PASSUses structured output | PASS | PASS | View ↗ |
| Cursor CloudMCP client: CursorVersion 1.0.0Run: Oct 2, 2026 | WARNFalls back to text | PASS | FAILNo output | View ↗ |
| OpenAI Agents APIMCP client: openai-mcp (Codex)Version 1.0.0Run: Oct 2, 2026 | FAILDuplicate output | PASS | PASS | View ↗ |
| Grok CLIMCP client: grok-shell-mcp-output-probeVersion 1.0.46Run: Oct 2, 2026 | FAILDuplicate output | PASS | PASS | View ↗ |
| Codex CLIMCP client: codex-mcp-clientVersion 0.160.0Run: Oct 2, 2026 | FAILDuplicate output | PASS | PASS | View ↗ |
| piMCP client: piVersion 1.0.0Run: Oct 2, 2026 | FAILDuplicate output | PASS | PASS | View ↗ |
| Cursor appMCP client: cursor-vscodeVersion 1.0.0Run: Oct 2, 2026 | WARNFalls back to text | PASS | FAILNo output | View ↗ |
| ChatGPT DesktopMCP client: openai-mcpVersion 1.0.0Run: Oct 2, 2026 | PASSUses structured output | PASS | PASS | View ↗ |
| Claude CodeMCP client: claude-codeVersion 2.1.287Run: Oct 2, 2026 | PASSUses structured output | PASS | PASS | View ↗ |
| Codex DesktopMCP client: codex-mcp-clientVersion 0.159.0-alpha.12.1Run: Oct 2, 2026 | FAILDuplicate output | PASS | PASS | View ↗ |
or unexpected output
Verdicts use reported markers. Observations describe individual sessions, not inferred model inputs.
Test 002: Protocol version
Connection metadataWhich MCP version did this connection use?
The MCP version used in this run. Green means the latest version tested here (2026-07-28); yellow means an older version. Full connection details are in the run evidence.
| Harness / MCP client | Protocol version | Evidence |
|---|---|---|
| Amp via OrbMCP client: amp-thread-actor · 1 | Older protocol2025-06-18 | View ↗ |
| Cursor DesktopMCP client: cursor-vscode · 1.0.0 | Older protocol2025-11-25 | View ↗ |
| piMCP client: pi · 1.0.3 | Older protocol2025-11-25 | View ↗ |
| Claude Code CLIMCP client: claude-code · 2.1.289 | Latest2026-07-28 | View ↗ |
| Claude Code (Claude desktop app, Code tab)MCP client: Anthropic/ClaudeAI · 1.0.0 | Latest2026-07-28 | View ↗ |
| Claude DesktopMCP client: Anthropic/ClaudeAI · 1.0.0 | Latest2026-07-28 | View ↗ |
| Codex CLIMCP client: codex-mcp-client · 0.160.1 | Older protocol2025-06-18 | View ↗ |
| ChatGPT Desktop CodexMCP client: codex-mcp-client · 0.160.0 | Older protocol2025-06-18 | View ↗ |
| ChatGPT Desktop WorkMCP client: codex-mcp-client · 0.160.0 | Older protocol2025-06-18 | View ↗ |
Test 003: URL elicitation
Optional interactionDoes the client advertise URL elicitation, and can it complete the flow?
Advertised means the client says it supports URL elicitation. Verified means the browser interaction completed and the tool continued successfully. A client can advertise support while using an older protocol that prevents verification.
| Harness / MCP client | URL support advertised | Interactive verification | Evidence |
|---|---|---|---|
| Amp via OrbMCP client: amp-thread-actor · 1 | No | Skipped | View ↗ |
| Cursor DesktopMCP client: cursor-vscode · 1.0.0 | No | Skipped | View ↗ |
| piMCP client: pi · 1.0.3 | No | Skipped | View ↗ |
| Claude Code CLIMCP client: claude-code · 2.1.289 | Yes | Verified | View ↗ |
| Claude Code (Claude desktop app, Code tab)MCP client: Anthropic/ClaudeAI · 1.0.0 | No | Skipped | View ↗ |
| Claude DesktopMCP client: Anthropic/ClaudeAI · 1.0.0 | No | Skipped | View ↗ |
| Codex CLIMCP client: codex-mcp-client · 0.160.1 | Yes | Skipped: requires newer protocol | View ↗ |
| ChatGPT Desktop CodexMCP client: codex-mcp-client · 0.160.0 | Yes | Skipped: requires newer protocol | View ↗ |
| ChatGPT Desktop WorkMCP client: codex-mcp-client · 0.160.0 | Yes | Skipped: requires newer protocol | View ↗ |
Only recorded results from approved sessions are shown. Open the run history for full details.
Test 004: Resource reading
MCP resourcesCan the agent read a resource, including one linked by a tool?
Two resources contain independent random markers. Green means the server recorded a read and the agent reported the exact marker. Yellow means it was not read or its marker was not reported. Red means an incorrect marker was reported.
| Harness / MCP client | Direct resource | Tool resource link | Evidence |
|---|---|---|---|
| Amp via OrbMCP client: amp-thread-actor · 1 | Not read | Not read | View ↗ |
| Cursor DesktopMCP client: cursor-vscode · 1.0.0 | Read successfully | Read successfully | View ↗ |
| piMCP client: pi · 1.0.3 | Read successfully | Read successfully | View ↗ |
| Claude Code CLIMCP client: claude-code · 2.1.289 | Read successfully | Read successfully | View ↗ |
| Claude Code (Claude desktop app, Code tab)MCP client: Anthropic/ClaudeAI · 1.0.0 | Not read | Not read | View ↗ |
| Claude DesktopMCP client: Anthropic/ClaudeAI · 1.0.0 | Read successfully | Read successfully | View ↗ |
| Codex CLIMCP client: codex-mcp-client · 0.160.1 | Read successfully | Read successfully | View ↗ |
| ChatGPT Desktop CodexMCP client: codex-mcp-client · 0.160.0 | Read successfully | Read successfully | View ↗ |
| ChatGPT Desktop WorkMCP client: codex-mcp-client · 0.160.0 | Read successfully | Read successfully | View ↗ |
Run it yourself
- Add the public MCP connection.
- Begin a run to record protocol and capabilities, then call the output probes and read both test resources.
- Optionally verify URL elicitation, then submit the exact markers you see for review.
No client authorization is required. Connection guide. The prompt runs all four tests with optional human verification. To skip the interaction, set verify_url_elicitation to false. The same endpoint supports standalone manual probes.
Read the test prompt
Use the What the Harness MCP server. Call begin_run with verify_url_elicitation: true and test_resources: true; if its schema requires begin_key, generate a fresh UUID for this run, keep it private, and reuse it for start retries. Supply the harness and model only if explicitly known, otherwise leave them null. Save the returned run_id, run_token, and client-declared identity. Keep resource URIs and run_token private; omit it from the final report and evidence. Call probe_both, probe_text_only, and probe_structured_only once each with that run_id and run_token. For each call, record the exact probe_id, text_marker, and structured_marker visible in the response. Use null for every value not visible; do not infer or invent markers. Discover MCP resources and resource templates if the client makes them available. Read the private resources.direct URI from begin_run through MCP resources/read and record its exact resource_marker. Call probe_resource_link and read the URI in its resource_link through MCP; record that resource_marker separately. If a resource cannot be read or its marker is not visible, report null. Never fetch these URIs over HTTP, inspect server files, or infer markers. Call probe_url_elicitation with the same run_id and run_token. If it presents a URL interaction, pause for the human to complete it through the client UI; do not fetch or complete the page yourself. After the human completes it, retry probe_url_elicitation with the same credentials until it returns the verification result. If support is not advertised or the protocol is unsupported, record the skip. If the human declines or cancels, record that outcome. Call submit_result with the run_id, run_token, and all three output reports, plus resource_reports containing direct and linked markers (or null), using those exact field names. If submission times out, retry with the same run_id, run_token, and unchanged reports. If the MCP connection reconnects, keep using the original run_id and run_token. Call result_status with the same credentials. Report the receipt, publication state, protocol version, declared URL support, verification status, and the output and resource markers you saw; say “not visible” for null values. Do not fetch the endpoint or website separately, and do not start another run to recover a missing marker.
Method, controls & limitations
The probe generates fresh, independent random markers for content[0].text and structuredContent. An agent cannot infer an unseen marker from the other field.
probe_both returns both fields. probe_text_only returns text alone. probe_structured_only returns structured output with an empty content array. The two single-field tools are controls.
Exact random marker reproduction demonstrates visibility. A missing report alone does not establish the precise input the model received. These observations describe specific sessions, not permanent behavior of a harness or model. Unknown dates, versions and configurations stay unknown.
Reporting runs automatically compare submitted markers with privately saved probe outputs. Only reviewed submissions appear in the results. Exact token overhead was not measured.
Inspect the diagnostic probe source · Public probe landing page
Your agent here.
Run the probe and save the exact report, client/version, model, date, and evidence. Results are reviewed before publication.
Your agent submits through the public MCP connection. Connection guide →
Keep credentials and unrelated conversation content out of your report.