Odin MCP · root-caused & fixed

The intermittent MCP failure: found, and fixed.

Reproduced on isolated GKE previews of real Odin code, isolated to a single mechanism, and fixed. The headline: tools/list was silently dropping ~30% of the time — a stale read-after-write on the Redis session store — now 0%.

30% → 0%
tools/list empty-202 drops, after the read-your-writes fix
42 KB
tools/list response written into the session blob, then read back stale
no HTML
/_mcp now answers application/json on every path
≈
worker mode = classic at equal CPU (didn't help)

The verdict

What was actually wrong, in order of impact.

1 · The real bug: a stale read-after-write.
The MCP SDK writes the 42 KB tools/list response into the session, saves it, then immediately reloads the session to hand it back. That reload returns the stale, pre-write 142 B blob ~35% of the time → the response is lost → empty 202. Only tools/list (the one big payload) trips it; every tools/call succeeds.
2 · The fix: read-your-writes.
A request-local write-through buffer on the session store serves reads from what was just written this request. tools/list drops 30% → 0% at every concurrency, latency unchanged.
3 · text/html → JSON.
/_mcp used to answer errors (and the empty 202) as text/html — even on prod — so strict clients rejected before reading the body. A response listener now forces application/json on every path. Benjamin's #1.
4 · Not the things we first suspected.
Not request volume (a heavy request path × few workers, since fixed by scan_dirs). Not cache eviction (a persistent pool didn't help). Not Cloudflare Access (prod reaches Odin directly). Not the SDK version (drop code identical in latest). Not worker mode (≈ classic at equal CPU).

The investigation, step by step

Eight milestones from report to fix. Click any test to open the side-by-side comparison of every run.

The root cause, caught red-handed

A transparent decorator on the session store recorded every read/write and surfaced it in an X-Mcp-Diag header. The trace is unambiguous — same session, sequential, no concurrency:

DROP (empty 202)
read (id)=142B          ← session loaded
write(id)=42533B        ← 42KB tools/list response written
read (id)=142B          ← STALE reload! response missing
→ empty 202
OK (200)
read (id)=142B
write(id)=42533B
read (id)=42533B        ← reads back the response
→ 200 + tools list
After the fix, the reload is served from the request-local buffer:
write(id)=42533B
read (id)=42533B[buf]   ← read-your-writes: always the fresh value
→ 200, every time
This is why moving cache pools never helped (both stores share the flaky read-after-write) and why only the 42 KB tools/list trips it, never the tiny tool-call responses.

The fix — before & after

tools/list drop rate by concurrency. The read-your-writes fix flatlines it at zero; the earlier persistent-pool attempt did not.

Recommended (shippable)
read-your-writes on the MCP session store
Make the store read-your-writes consistent within a request (a request-local write-through, or a ChainAdapter[array,redis] pool). Kills the drop at the source. Pair with the JSON-response listener so no path ever returns text/html.
Also worth doing • Report the drop upstream — it reproduces on the latest SDK (v0.7).
• The SDK/bundle upgrade (protocol 2025-11-25, kills the per-request file scan) is a good modernization on its own ticket — it needs a discovery/schema refactor and does not fix this drop.
• Skip ARO-4256 worker mode as a perf play — ≈ classic at equal CPU.

Versions & protocol

InstalledLatestNote
mcp/sdkv0.5.0v0.7.0transport drop code byte-identical → won't fix the drop
symfony/mcp-bundlev0.9.0v0.12.0replaces file-scan discovery w/ DI autoconfig (BC)
MCP protocol (default)2025-06-182025-11-25latest bumps default; enum already knew it
Controlled experiment on isolated GKE previews of real Odin code · classic vs worker vs fixed · Cloudflare Access read-only + scoped bypass · root cause pinned via X-Mcp-Diag instrumentation · generated for ARO-4202.
← Back to timeline

Comparison of every run

All four experiments on the same isolated preview + harness (Benjamin's exact sequence, ramped 1→48 concurrent).

tools/list empty-202 drop rate (%) by concurrency — lower is better.
median latency (ms) of successful calls — the fix does not change latency; worker mode ≈ classic.