The intermittent MCP failure: found, and fixed.
Reproduced on isolated GKE previews of real Odin code, isolated to a single mechanism, and fixed. The headline: tools/list was silently dropping ~30% of the time — a stale read-after-write on the Redis session store — now 0%.
The verdict
What was actually wrong, in order of impact.
The MCP SDK writes the 42 KB
tools/list response into the session, saves it, then immediately reloads the session to hand it back. That reload returns the stale, pre-write 142 B blob ~35% of the time → the response is lost → empty 202. Only tools/list (the one big payload) trips it; every tools/call succeeds.A request-local write-through buffer on the session store serves reads from what was just written this request. tools/list drops 30% → 0% at every concurrency, latency unchanged.
/_mcp used to answer errors (and the empty 202) as text/html — even on prod — so strict clients rejected before reading the body. A response listener now forces application/json on every path. Benjamin's #1.Not request volume (a heavy request path × few workers, since fixed by scan_dirs). Not cache eviction (a persistent pool didn't help). Not Cloudflare Access (prod reaches Odin directly). Not the SDK version (drop code identical in latest). Not worker mode (≈ classic at equal CPU).
The investigation, step by step
Eight milestones from report to fix. Click any test to open the side-by-side comparison of every run.
The root cause, caught red-handed
A transparent decorator on the session store recorded every read/write and surfaced it in an X-Mcp-Diag header. The trace is unambiguous — same session, sequential, no concurrency:
read (id)=142B ← session loaded write(id)=42533B ← 42KB tools/list response written read (id)=142B ← STALE reload! response missing → empty 202
read (id)=142B write(id)=42533B read (id)=42533B ← reads back the response → 200 + tools list
write(id)=42533B read (id)=42533B[buf] ← read-your-writes: always the fresh value → 200, every timeThis is why moving cache pools never helped (both stores share the flaky read-after-write) and why only the 42 KB
tools/list trips it, never the tiny tool-call responses.
The fix — before & after
tools/list drop rate by concurrency. The read-your-writes fix flatlines it at zero; the earlier persistent-pool attempt did not.
• The SDK/bundle upgrade (protocol 2025-11-25, kills the per-request file scan) is a good modernization on its own ticket — it needs a discovery/schema refactor and does not fix this drop.
• Skip ARO-4256 worker mode as a perf play — ≈ classic at equal CPU.
Versions & protocol
| Installed | Latest | Note | |
|---|---|---|---|
| mcp/sdk | v0.5.0 | v0.7.0 | transport drop code byte-identical → won't fix the drop |
| symfony/mcp-bundle | v0.9.0 | v0.12.0 | replaces file-scan discovery w/ DI autoconfig (BC) |
| MCP protocol (default) | 2025-06-18 | 2025-11-25 | latest bumps default; enum already knew it |