Repository navigation
MCP bearer token passed via environment is captured in systemd-coredump journal entries on provider crashes #12031
Description
Activity
Triage
Confirmed on current
main(0f5a1513e). This is a real Codex credential-channel bug, not a wash of #10020. The thread-scopedt3-codeMCP bearer is injected into the Codex child environment asT3_MCP_BEARER_TOKEN. On Linux,systemd-coredumppersists that entire block asCOREDUMP_ENVIRONin the journal. The reporter’s scope is right: session-bounded, same-UID, no operator knob that keeps crash capture and drops this field.What the code does
CodexAdapter.startSessionreads the in-memory MCP session and puts the raw token in the child env, then tells Codex to read it from that variable (apps/server/src/provider/Layers/CodexAdapter.ts). The-cflags only name the URL and the env-var name; the secret itself is not on argv. That is why this is the inverse of the Claude argv leak, not a duplicate of it.The token is a 32-byte random bearer from
McpSessionRegistry. The registry stores a SHA-256 hash only./mcpis mounted outside the environment auth stack; this token is the only guard for thet3-codetoolkits.Lifetime, as designed:
Path Token fate ProviderService.stopSessionclearMcpSession→revokeActiveMcpThread(eager)startSession/ recover failuresame revoke New session on the same thread issueActiveMcpCredentialrevokes the thread first, then mints a new tokenUnclean death with no later touchDEFAULT_LIVENESS_WINDOW_MS= 24 hoursruntime.error/session.exitednot revoked. Ingestion marks the thread errorand clears turn state. NostopSession.Session reaper 30-minute idle, and it skips an active turn / background liveness That matches the reporter’s measurement: a token taken from a dump returned
401after the session ended; a freshly minted one returned200. The gap they could not see from outside: a crash that has not yet gone throughstopSessionleaves a still-valid token inCOREDUMP_ENVIRON. On a host where several provider sessions share one UID, any of those sessions that can runjournalctl --allcan use it until revoke, replacement, or the 24h window.journalctl --output=jsonomits the field because it is large;--allshows it.coredump.confon systemd 261 has no exclude forCOREDUMP_ENVIRON.Codex client constraint
bearer_token_env_varis an env-var name, not a file/fd. Codex’s current HTTP MCP config is:bearer_token_env_var/env_http_headers— still environmenthttp_headers— static values. Passing the bearer here via-cwould put it on argv. Do not do that.http_headers_helper— local command that prints{"Authorization":"Bearer …"}. Documented for local HTTP MCP. Codex caches helper headers and refreshes once on401/403. Explicit bearer / OAuth wins over a helperAuthorization, so a helper path must dropbearer_token_env_var.
That helper is the non-env channel that exists today without a Codex client change.
--bootstrap-fdis the T3 server bootstrap path. It is not a Codex MCP auth mechanism; do not try to reuse it as one.Related, not duplicates
Issue / PR Why it does not close this Open #10020 / open #10090 Claude: bearer serialized onto --mcp-configargv. #10020 triage even pointed at Codex’sT3_MCP_BEARER_TOKENas the “stronger” pattern. #10090 moves Claude onto the env channel this issue is about. Same secret, opposite surface. Do not attributeCloses #12031to #10090.#4659 Keeps credentials alive across turns that never touch MCP ( touch). Lifetime, not the env channel.#4189 Revoke-by- providerSessionIdfor overlapping sessions. Not coredump.ACP / OpenCode / Claude on mainDifferent auth surfaces; not this spawn env. No open or merged PR removes
T3_MCP_BEARER_TOKENfrom the Codex child environment, addshttp_headers_helper, or revokes on provider process exit.Suggested fix
Keep this as a bugfix on the Codex MCP spawn path, plus one lifecycle hardening that every provider benefits from.
- Get the bearer out of Codex
environ. Preferhttp_headers_helperover a new Codex feature. Write the token to a0600file underserverConfig.secretsDir(already0700; not the workspace). Helper prints{"Authorization":"Bearer <token>"}and nothing else. Pass only-c mcp_servers.t3-code.http_headers_helper=…plus the URL. OmitT3_MCP_BEARER_TOKENandbearer_token_env_var. Delete the file inclearMcpSession/ session finalizer. - Do not put the raw bearer in
http_headersvia-c. That is Keep the Claude t3-code MCP bearer out of process arguments #10020 for Codex. - Revoke on provider death, not only on
stopSession.runtime.errorandsession.exitedshouldclearMcpSession(orrevokeProviderSessionfor that provider session). A token recovered from the dump of the process that just died should be dead on arrival. - Tests: with an MCP session, spawned Codex env has no raw bearer / no
T3_MCP_BEARER_TOKEN; argv has noAuthorization/Bearer; helper file is0600and is gone after stop;runtime.error/session.exitedmakesresolve(token)undefined. - fix(claude): keep MCP credentials out of process arguments #10090: if Claude must leave argv, do not stop at a child env var on Linux coredump hosts. Same file/helper belongs there too. Do not treat fix(claude): keep MCP credentials out of process arguments #10090 as closing this.
A shorter default TTL alone is not the fix. Crashes are common on a busy host; T3 already sees them.
Workaround
None in-app. Operator-only, and they cost the diagnostics the dumps are for:
- Disable coredumps for the provider unit /
LimitCORE=0 - Tighten journal ACL so the provider UID cannot read
COREDUMP_ENVIRON - Do not run mutually untrusted sessions under one Unix account
journalctlwithout--allhiding the field is not a mitigation.Classification: bug · accepted · moderate (thread-scoped
/mcpbearer; Linuxsystemd-coredump; same-UID multi-session; live until stop/replace/24h, and crash does not revoke today)
Labels: addbug,accepted,via-triage,codex
Discord tags:providers,linux- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.acceptedfeature request acceptedfeature request acceptedvia-triageFiled through npx t3 triageFiled through npx t3 triage
on Sep 16, 2026 Status on current main (
477282263a):- Codex no longer uses the env channel. The bearer goes in
http_headersin the thread config sent over the app-server JSON-RPC connection (apps/server/src/orchestration-v2/Adapters/CodexAdapterV2.ts:1321-1322). - Claude is fixed on main. fix(server): keep the Claude MCP token out of process arguments #17408 moved the token off argv into the child env as
T3_CODE_MCP_AUTHORIZATION, which put Claude on this issue's channel. fix(server): send Claude MCP servers over the control channel #17898 now sends the MCP servers withsetMcpServersover the SDK control channel instead (apps/server/src/orchestration-v2/Adapters/ClaudeAdapterV2.ts:646and:687). - ACP joined this channel with the V2 orchestrator (feat(orchestrator): introduce new orchestrator #2829), after the triage. It tells every agent, Grok and Antigravity included, to spawn the
t3 acp-mcp-bridgestdio server withT3_ACP_MCP_AUTHORIZATIONin its env (packages/provider-acp/src/server/adapter.ts:705). Registry ACP agents also get it in their own process env and in their client terminals, throughprocessEnvironment(:714). - Pi still sets
T3_MCP_BEARER_TOKENin the pi child env (packages/provider-pi/src/server/mcpInjection.ts:313-314).
A dummy variable set on a process that was then crashed with SIGSEGV landed in
COREDUMP_ENVIRON.Scope notes:
- Capture needs
core_patternpiping to systemd-coredump and a signal that dumps core. SIGKILL and OOM kills do not. With a piped pattern the kernel aborts the dump only whenRLIMIT_COREis 1. With 0, systemd-coredump skips the core file but still writes the journal entry,COREDUMP_ENVIRONincluded. A dummy-variable test of both values showed this. - For registry ACP agents and Pi, every command the agent runs inherits the variable, so a crash of any of them dumps the token, not only a crash of the provider CLI. For Grok and Antigravity it reaches only the bridge process.
- V2 now releases the session when the provider event stream fails (
apps/server/src/orchestration-v2/ProviderSessionManager.ts:1978-1982,runtime_error), and the release revokes the credential unless another session still holds it (:1027). That shortens the window for provider crashes. A crashing subprocess does not end the session, so its token stays valid until stop, archive, delete, idle reap, rotation or the 24 h liveness window. - The same user can read the entry, since it lands in that user's journal. So can root, the journal-reading admin groups the distro grants, and anyone holding a copy of the journal.
Until the fix ships,
LimitCORE=1on the unit that runs T3, orprlimit --core=1in front of the command that starts it, stops the kernel handing crashes to systemd-coredump. Descendants inherit the limit, at the cost of every crash dump under it. A shell'sulimit -c 1is no substitute, because bash counts it in 1024-byte blocks (512 in POSIX mode), so the limit never becomes 1. The triage'sLimitCORE=0still lets the journal entry through.The fix is in #17957. It moves ACP and Pi to an owner-only credential file per thread, so the token is no longer in their child environments or in those processes'
COREDUMP_ENVIRON.- Codex no longer uses the env channel. The bearer goes in
Summary
T3_MCP_BEARER_TOKENis passed to spawned provider processes through theenvironment. On Linux,
systemd-coredumprecords the entire environment ofany process that dumps core into the systemd journal as
COREDUMP_ENVIRON,in plaintext. So every provider crash writes that session's MCP bearer token to
a log readable by the account the providers run as.
Requesting a file- or fd-based alternative, in line with the
--bootstrap-fdmechanism this project already has.
Where it happens
From the shipped bundle:
paired with the Codex MCP client config:
Measured impact on one host
Every core dump over a two-week window on a machine running a provider fleet:
COREDUMP_ENVIRONis readable byroot, thesystemd-journalgroup, and —via the journal directory ACL — the user the providers run as. Where multiple
provider sessions share one Unix account, any session that can run
journalctlcan read the MCP token of any other session that has crashed.journalctl --output=jsonomits the field by default because it is large;--allreveals it. That makes the exposure easy to miss entirely.Scope, stated honestly
mcpSession.authorizationHeaderreads as a per-session token, and thatmatches observation: a token recovered from a dump returned
401against both/mcpand/api/orchestration/snapshotafter its session ended, while a freshlyminted token returned
200on the same route. So this is not a long-livedcredential leak.
The exposure is therefore bounded to the lifetime of the session whose provider
crashed — but within that window it is a usable token sitting in a world-of-the-
account-readable log, and provider crashes are not rare on a busy host.
I would call this moderate rather than critical, and I would rather describe it
accurately than inflate it.
Why the environment is the wrong channel specifically
Not a general "env vars are bad" argument — this is a concrete Linux behaviour:
/proc/<pid>/environis readable for the lifetime of the process.systemd-coredumppersists the whole block on any crash, and it outlives thecore file, since
Storage=controls the core, not the journal metadata.coredump.confon systemd 261 exposes onlyStorage,ProcessSizeMax,SizeMax,ExternalSizeMax,MaxUse,Compress,EnterNamespace,CoredumpReceive— there is no way for an operator to exclude that field.So a deployment cannot mitigate this without disabling crash capture entirely,
which costs the diagnostics operators need for exactly the crashes that leak.
Requested change
A non-environment path for the MCP bearer token, e.g.:
0600file and pass the path, or--bootstrap-fdflag ("Read one-time bootstrap secrets from the given filedescriptor"), which shows the project already treats fd-passing as the right
channel for secrets.
A constraint I cannot resolve from outside: the Codex MCP client is
configured here via
bearer_token_env_var, which implies the client itselfexpects an environment variable. If that is the only mechanism Codex supports,
this may need a client-side change first, or a T3-side mitigation instead
(shorter session TTL, or revoking the MCP session on provider exit so a
recovered token is dead on arrival).
Environment
Not asking for
No mitigation is needed on my side beyond what an operator can already do. Filing
this because the leak is structural to the channel rather than to any
deployment's configuration, and because no operator-side setting can close it.