An independent field guide to agent quirksby Polytomic
A puzzled pixel computer with a question mark on its blue screen

What the harness?

Same protocol. Different behavior.

Test 001: The double payload

MCP output

Does your harness pass tool output once, or twice?

Why it matters

We don’t want to flood the model with duplicate tokens. MCP tools can return the same data as text and structured output; the harness should pass it to the model once.

This test puts a different random marker in each field and asks the agent what it sees. Reporting both markers reveals duplication. Two extra probes check that text-only and structured-only responses still get through.

What counts as a pass?

Both fields sent

PASS: uses structured output
WARN: falls back to text
FAIL: duplicate, missing, or unexpected output

Text only sent

PASS: only text gets through
FAIL: missing or unexpected output

Structured only sent

PASS: only structured output gets through
FAIL: missing or unexpected output

Our rubric: prefer one structured payload, avoid duplication, preserve single-field results. This preference is not a claim about what the MCP standard mandates.

Observed results

Run history ↗

Latest approved run per harness, MCP client and version

Scroll the table sideways to compare all three responses →

Test 001 observations. Each result applies to one recorded session. Desired outcomes are defined in the rubric above.
HarnessResponse sent by the probeEvidence
Both fieldscontent + structuredContentText onlycontentStructured onlystructuredContent
Amp via OrbMCP client: amp-thread-actorVersion 1Run: Oct 5, 2026FAILDuplicate outputPASSPASSView ↗
Cursor DesktopMCP client: cursor-vscodeVersion 1.0.0Run: Oct 5, 2026WARNFalls back to textPASSFAILNo outputView ↗
piMCP client: piVersion 1.0.3Run: Oct 5, 2026FAILDuplicate outputPASSPASSView ↗
Claude Code CLIMCP client: claude-codeVersion 2.1.289Run: Oct 5, 2026PASSUses structured outputPASSPASSView ↗
Claude Code (Claude desktop app, Code tab)MCP client: Anthropic/ClaudeAIVersion 1.0.0Run: Oct 5, 2026PASSUses structured outputPASSPASSView ↗
Claude DesktopMCP client: Anthropic/ClaudeAIVersion 1.0.0Run: Oct 5, 2026WARNFalls back to textPASSPASSView ↗
Codex CLIMCP client: codex-mcp-clientVersion 0.160.1Run: Oct 5, 2026FAILDuplicate outputPASSPASSView ↗
ChatGPT Desktop CodexMCP client: codex-mcp-clientVersion 0.160.0Run: Oct 5, 2026FAILDuplicate outputPASSPASSView ↗
ChatGPT Desktop WorkMCP client: codex-mcp-clientVersion 0.160.0Run: Oct 5, 2026FAILDuplicate outputPASSPASSView ↗
GrokbotMCP client: CursorVersion 1.0.0Run: Oct 3, 2026WARNFalls back to textPASSFAILNo outputView ↗
ChatGPT Work Cloud via WebMCP client: openai-mcp (Codex)Version 1.0.0Run: Oct 3, 2026FAILDuplicate outputPASSPASSView ↗
ChatGPT via WebMCP client: openai-mcpVersion 1.0.0Run: Oct 3, 2026PASSUses structured outputFAILNo outputPASSView ↗
Amp CLIMCP client: amp-thread-actorVersion 1Run: Oct 2, 2026PASSUses structured outputPASSPASSView ↗
Cursor CloudMCP client: CursorVersion 1.0.0Run: Oct 2, 2026WARNFalls back to textPASSFAILNo outputView ↗
OpenAI Agents APIMCP client: openai-mcp (Codex)Version 1.0.0Run: Oct 2, 2026FAILDuplicate outputPASSPASSView ↗
Grok CLIMCP client: grok-shell-mcp-output-probeVersion 1.0.46Run: Oct 2, 2026FAILDuplicate outputPASSPASSView ↗
Codex CLIMCP client: codex-mcp-clientVersion 0.160.0Run: Oct 2, 2026FAILDuplicate outputPASSPASSView ↗
piMCP client: piVersion 1.0.0Run: Oct 2, 2026FAILDuplicate outputPASSPASSView ↗
Cursor appMCP client: cursor-vscodeVersion 1.0.0Run: Oct 2, 2026WARNFalls back to textPASSFAILNo outputView ↗
ChatGPT DesktopMCP client: openai-mcpVersion 1.0.0Run: Oct 2, 2026PASSUses structured outputPASSPASSView ↗
Claude CodeMCP client: claude-codeVersion 2.1.287Run: Oct 2, 2026PASSUses structured outputPASSPASSView ↗
Codex DesktopMCP client: codex-mcp-clientVersion 0.159.0-alpha.12.1Run: Oct 2, 2026FAILDuplicate outputPASSPASSView ↗
PASS Desired outputWARN Text fallbackFAIL Duplicate, missing,
or unexpected output

Verdicts use reported markers. Observations describe individual sessions, not inferred model inputs.

Test 002: Protocol version

Connection metadata

Which MCP version did this connection use?

The MCP version used in this run. Green means the latest version tested here (2026-07-28); yellow means an older version. Full connection details are in the run evidence.

Harness / MCP clientProtocol versionEvidence
Amp via OrbMCP client: amp-thread-actor · 1Older protocol2025-06-18View ↗
Cursor DesktopMCP client: cursor-vscode · 1.0.0Older protocol2025-11-25View ↗
piMCP client: pi · 1.0.3Older protocol2025-11-25View ↗
Claude Code CLIMCP client: claude-code · 2.1.289Latest2026-07-28View ↗
Claude Code (Claude desktop app, Code tab)MCP client: Anthropic/ClaudeAI · 1.0.0Latest2026-07-28View ↗
Claude DesktopMCP client: Anthropic/ClaudeAI · 1.0.0Latest2026-07-28View ↗
Codex CLIMCP client: codex-mcp-client · 0.160.1Older protocol2025-06-18View ↗
ChatGPT Desktop CodexMCP client: codex-mcp-client · 0.160.0Older protocol2025-06-18View ↗
ChatGPT Desktop WorkMCP client: codex-mcp-client · 0.160.0Older protocol2025-06-18View ↗

Test 003: URL elicitation

Optional interaction

Does the client advertise URL elicitation, and can it complete the flow?

Advertised means the client says it supports URL elicitation. Verified means the browser interaction completed and the tool continued successfully. A client can advertise support while using an older protocol that prevents verification.

Harness / MCP clientURL support advertisedInteractive verificationEvidence
Amp via OrbMCP client: amp-thread-actor · 1NoSkippedView ↗
Cursor DesktopMCP client: cursor-vscode · 1.0.0NoSkippedView ↗
piMCP client: pi · 1.0.3NoSkippedView ↗
Claude Code CLIMCP client: claude-code · 2.1.289YesVerifiedView ↗
Claude Code (Claude desktop app, Code tab)MCP client: Anthropic/ClaudeAI · 1.0.0NoSkippedView ↗
Claude DesktopMCP client: Anthropic/ClaudeAI · 1.0.0NoSkippedView ↗
Codex CLIMCP client: codex-mcp-client · 0.160.1YesSkipped: requires newer protocolView ↗
ChatGPT Desktop CodexMCP client: codex-mcp-client · 0.160.0YesSkipped: requires newer protocolView ↗
ChatGPT Desktop WorkMCP client: codex-mcp-client · 0.160.0YesSkipped: requires newer protocolView ↗

Only recorded results from approved sessions are shown. Open the run history for full details.

Test 004: Resource reading

MCP resources

Can the agent read a resource, including one linked by a tool?

Two resources contain independent random markers. Green means the server recorded a read and the agent reported the exact marker. Yellow means it was not read or its marker was not reported. Red means an incorrect marker was reported.

Harness / MCP clientDirect resourceTool resource linkEvidence
Amp via OrbMCP client: amp-thread-actor · 1Not readNot readView ↗
Cursor DesktopMCP client: cursor-vscode · 1.0.0Read successfullyRead successfullyView ↗
piMCP client: pi · 1.0.3Read successfullyRead successfullyView ↗
Claude Code CLIMCP client: claude-code · 2.1.289Read successfullyRead successfullyView ↗
Claude Code (Claude desktop app, Code tab)MCP client: Anthropic/ClaudeAI · 1.0.0Not readNot readView ↗
Claude DesktopMCP client: Anthropic/ClaudeAI · 1.0.0Read successfullyRead successfullyView ↗
Codex CLIMCP client: codex-mcp-client · 0.160.1Read successfullyRead successfullyView ↗
ChatGPT Desktop CodexMCP client: codex-mcp-client · 0.160.0Read successfullyRead successfullyView ↗
ChatGPT Desktop WorkMCP client: codex-mcp-client · 0.160.0Read successfullyRead successfullyView ↗

Run it yourself

  1. Add the public MCP connection.
  2. Begin a run to record protocol and capabilities, then call the output probes and read both test resources.
  3. Optionally verify URL elicitation, then submit the exact markers you see for review.
https://what-the-harness.polytomic.com/mcp

No client authorization is required. Connection guide. The prompt runs all four tests with optional human verification. To skip the interaction, set verify_url_elicitation to false. The same endpoint supports standalone manual probes.

Read the test prompt

Use the What the Harness MCP server. Call begin_run with verify_url_elicitation: true and test_resources: true; if its schema requires begin_key, generate a fresh UUID for this run, keep it private, and reuse it for start retries. Supply the harness and model only if explicitly known, otherwise leave them null. Save the returned run_id, run_token, and client-declared identity. Keep resource URIs and run_token private; omit it from the final report and evidence. Call probe_both, probe_text_only, and probe_structured_only once each with that run_id and run_token. For each call, record the exact probe_id, text_marker, and structured_marker visible in the response. Use null for every value not visible; do not infer or invent markers. Discover MCP resources and resource templates if the client makes them available. Read the private resources.direct URI from begin_run through MCP resources/read and record its exact resource_marker. Call probe_resource_link and read the URI in its resource_link through MCP; record that resource_marker separately. If a resource cannot be read or its marker is not visible, report null. Never fetch these URIs over HTTP, inspect server files, or infer markers. Call probe_url_elicitation with the same run_id and run_token. If it presents a URL interaction, pause for the human to complete it through the client UI; do not fetch or complete the page yourself. After the human completes it, retry probe_url_elicitation with the same credentials until it returns the verification result. If support is not advertised or the protocol is unsupported, record the skip. If the human declines or cancels, record that outcome. Call submit_result with the run_id, run_token, and all three output reports, plus resource_reports containing direct and linked markers (or null), using those exact field names. If submission times out, retry with the same run_id, run_token, and unchanged reports. If the MCP connection reconnects, keep using the original run_id and run_token. Call result_status with the same credentials. Report the receipt, publication state, protocol version, declared URL support, verification status, and the output and resource markers you saw; say “not visible” for null values. Do not fetch the endpoint or website separately, and do not start another run to recover a missing marker.

Method, controls & limitations

The probe generates fresh, independent random markers for content[0].text and structuredContent. An agent cannot infer an unseen marker from the other field.

probe_both returns both fields. probe_text_only returns text alone. probe_structured_only returns structured output with an empty content array. The two single-field tools are controls.

Exact random marker reproduction demonstrates visibility. A missing report alone does not establish the precise input the model received. These observations describe specific sessions, not permanent behavior of a harness or model. Unknown dates, versions and configurations stay unknown.

Reporting runs automatically compare submitted markers with privately saved probe outputs. Only reviewed submissions appear in the results. Exact token overhead was not measured.

Inspect the diagnostic probe source · Public probe landing page

Your agent here.

Run the probe and save the exact report, client/version, model, date, and evidence. Results are reviewed before publication.

Your agent submits through the public MCP connection. Connection guide →

Keep credentials and unrelated conversation content out of your report.