Repository navigation
Ambiguous @Name mentions silently notify nobody when a display name maps to two pubkeys; agents appear online but unreachable, and 'agents archive' does not resolve the ambiguity #4303
Description
Activity
- changed the title
[-]Relay permanently stops delivering DM-channel messages after a websocket reconnect (group channels unaffected); not recoverable by agent restart or new DM channel[/-][+]Duplicate agent identities (one display name, two pubkeys) break DM/mention resolution silently; no validation that a stored key matches its recorded pubkey[/+]on Aug 2, 2026 Amending my own report: the original framing ("relay stops delivering DMs after a websocket reconnect") is likely wrong, and I've rewritten the issue body around a more probable mechanism — duplicate agent identities, where one display name is registered under two distinct pubkeys and name resolution silently lands on the one with no live process behind it.
That fits the evidence better than the original theory, which never explained why a full cold restart with verifiably correct subscriptions changed nothing.
The original observations are preserved at the bottom of the body; they're reproducible and may still be useful independently. I've also marked clearly which parts I verified directly and which came from an agent's roster query that I could not independently confirm — the credential needed to query the roster sits behind a gpg passphrase prompt I couldn't satisfy non-interactively.
Retitled to match. Happy to close this and open a narrower one if maintainers would prefer the two concerns separated.
Follow-up: the duplicate-identity hypothesis in the issue body is now confirmed, and attempting to remediate it surfaced two further problems in
agents archive.Duplicate identities confirmed
Querying the community roster (
users get --name <name>) returned two distinct pubkeys for five of eight display names, and one each for the other three — 13 bot registrations for 8 running agents.For each of the five I verified which pubkey is live rather than assuming: decrypt the agent's private key, query its own profile (
users getwith no--pubkey), and compare against the pubkey recorded in the local agent registry. All five matched their record, confirming the second pubkey in each pair is an orphan with no private key behind it.This is consistent with the provisioning gap described in the issue body — a record created with a pubkey but no stored key, key material generated later, and the original registration never removed.
Problem 1 —
agents archivereports success without applyingArchived all five orphans as the community owner. Every request returned:
{"action":"archive","ok":true,"event_id":"<hex>","target":"<pubkey>"}Three subsequently appeared in the archive snapshot (
agents archived, kind 13535). Two never did. Re-submitting those two produced a freshok:truewith a new event id and still no snapshot entry — so this is reproducible, not a propagation race.The CLI help notes the relay selects a consent path (self / admin / owner) from the submitted request and does not retry with a different shape. If a request is rejected because the selected path doesn't authorize it, that rejection isn't surfaced: the caller gets
ok:trueand a signed event id, and can only detect the no-op by polling the snapshot afterwards.Whatever the underlying cause,
ok:trueon a request that is not applied is the reportable part — a caller cannot distinguish success from silent failure.Problem 2 — archived identities still appear in name lookup
For the three orphans that did land in the archive snapshot,
users get --name <name>continues to return both pubkeys.If name lookup is intentionally a raw profile search that ignores archive state, that's defensible — but then
agents archiveisn't a remedy for duplicate display names, and there appears to be no other mechanism for retiring a superseded identity short of moderationban, which is semantically wrong for an identity that was simply replaced by a re-key.If mention/DM resolution shares this lookup path, archiving doesn't address the original failure at all.
Suggested
agents archiveshould return a failure when the request isn't applied, rather thanok:truewith an event id.- Archive state should be reflected in identity lookup — or the docs should state plainly that it isn't, and name the intended mechanism for retiring a superseded identity.
- Ideally, registering a second identity under an existing display name should be rejected or require disambiguation, so this state can't arise in the first place.
Pubkeys and relay host omitted here; happy to supply the event ids privately if useful for tracing the two no-op requests.
Mechanism now confirmed end-to-end, and it reframes this issue. The delivery layer is fine — mentions of a duplicated display name are silently dropped at resolution time.
The confirming experiment
Three agents were
@Namementioned by the community owner in one channel, seconds apart. Only one replied. Minutes later, one of the silent two was mentioned in the same channel with an explicit pubkey:messages send --channel <ch> --content "@Neuromancer diagnostic ping ..." \ --mention <live-pubkey>Response was immediate:
13:41:44 turn starting for channel <ch> 13:41:55 turn complete — replied in-threadSame agent, same channel, same process, ~5 minutes apart. Ignored by name, answered by pubkey. The agent was healthy the entire time: websocket established, subscribed to that channel,
presence set to online, balanced turn counts.Why
From
messages send --help:Supplying any explicit identity permits unresolved or ambiguous @name text as presentation-only; uniquely resolved member names still notify
So an
@Namematching two pubkeys is ambiguous, and an ambiguous mention does not notify anyone. That is documented behavior. Combined with the duplicate registrations described earlier in this issue, the result is:- 5 of 8 display names map to 2 pubkeys each → unreachable by name
- 3 of 8 are singletons → work perfectly
That split held exactly, all day, and was the entire user-visible symptom.
It also explains the DM half. A client resolving one of the duplicated names picks the orphan pubkey; the orphan has no process behind it and therefore no presence; the client renders the agent offline and won't open a DM to it. From the user's side: "the bots are offline and ignoring me," while every server-side and agent-side signal reads healthy.
Severity
The failure is completely silent in both directions. The sender gets an accepted message with no warning that the mention resolved to nothing; the agent never learns a message existed. There is no error, no log line, and no metric on either side. Diagnosing it required reconstructing per-channel turn history from agent logs and then A/B-testing name vs pubkey mentions — several hours for a condition the relay already knows about at resolution time.
The remedy in the tooling does not work
agents archiveis the obvious way to retire a superseded identity, and it does not resolve the ambiguity: three orphans were successfully archived (confirmed present in the kind 13535 snapshot) andusers get --namestill returns both pubkeys for all three afterwards. Two further orphans could not be archived at all — see the previous comment,ok:truewith no effect.So there is currently no working path from "I have duplicate display names" to "mentions work again", short of moderation
ban, which is semantically wrong for an identity that was merely superseded by a re-key.Suggested, in priority order
- Warn on ambiguous mention resolution. If
@Nameresolves to more than one member, that should surface to the sender — a response field, a system message, anything. Silently downgrading to presentation-only is the core defect; everything else here is recoverable if the operator can see it happening. - Reject duplicate display names at registration, or require disambiguation, so this state cannot arise.
- Make
agents archiveexclude the identity from name resolution — otherwise archiving is not a remedy for the one problem it most obviously appears to address. - Log ambiguous-resolution events relay-side so operators can find them without A/B testing.
Happy to supply event ids or the raw agent logs for the confirming experiment if useful.
- changed the title
[-]Duplicate agent identities (one display name, two pubkeys) break DM/mention resolution silently; no validation that a stored key matches its recorded pubkey[/-][+]Ambiguous @Name mentions silently notify nobody when a display name maps to two pubkeys; agents appear online but unreachable, and 'agents archive' does not resolve the ambiguity[/+]on Aug 2, 2026 Root cause found and fixed for the client half: #4371.
_resolveMentionsinmobile/lib/features/channels/send_message_provider.darttook the first profile whose display name matched andbreaked. With no ambiguity detection andcachebeing aMap, "first match" meant "whichever entry the map yielded first" — and that decided which pubkey went into theptag.That explains the detail I could not account for earlier in this thread: why one duplicated agent answered
@Namewhile the others never did. It was map ordering. Names whose live profile happened to enumerate first worked; the rest tagged a dead identity, and since agents match purely onptags, the intended agent never saw the mention.The PR collects all matching candidates and tags them all, with a regression test verified to fail against the previous behaviour.
Two relay-side items from this issue are not addressed by that PR and remain open:
agents archivereturningok:truewith a signed event id while never applying — reproducible on two specific identities, twice each.- Archived identities still appearing in
users get --name.
Happy to split those into their own issue if you'd prefer this one scoped to the mention-resolution bug now that it has a fix attached.
Resolved on our deployment — and the mechanism is narrower than this issue states
Fixed this today for 5 of 8 agents. Posting the corrected mechanism plus two dead ends, because both cost real time and neither is obvious from the code.
It is CHANNEL membership that creates the ambiguity, not relay membership
crates/buzz-relay/src/workflow_sink.rsbuilds the name map from channel members:get_members(community, channel_uuid) // channel membership get_users_bulk(community, &pubkeys) // display_name per pubkey resolve_mention_pubkeys(&text, &named_members)
resolve_mention_pubkeysthen foldsname -> pubkeyand sets the slot toNonewhen two distinct pubkeys share a name, so noptag is emitted. The behaviour is deliberate and documented in that function, and I agree with the design choice — arbitrary selection misroutes, tagging all is a false-wake firehose.The consequence is that the ambiguity is per channel. A stale twin only breaks mentions in channels it is still a member of.
Dead end 1:
agents archivedoes not helpConfirmed. It never touches channel membership, so both twins remain members and the name stays ambiguous. Two of our orphans also returned
ok:truewith a signed event id and never entered the kind:13535 snapshot at all — a silent no-op, which is worth a separate look.Dead end 2: kind 9031
RELAY_ADMIN_REMOVE_MEMBERdoes not applyI built a signer for kind 9031 (
buzz-core/src/kind.rs, handlerhandlers/relay_admin.rs) assuming the twins were relay members. Every request returned:400 invalid: member not found: <orphan pubkey>They were never in
relay_members. Worth noting for anyone else who tries this path:POST /eventsalso requires NIP-98 auth (Authorization: Nostr <base64 kind-27235>), and without it you get401 missing Nostr auth— a different failure that is easy to mistake for the same problem.What actually fixes it
buzz channels remove-member --channel <uuid> --pubkey <orphan-pubkey>signed with an owner/admin key — agent keys cannot remove members. We removed 22 orphan memberships across 6 channels; all returned
accepted:true.Verification, same channel, name-only mention:
{"accepted":true,"mention_pubkeys":["<live-pubkey>"]}That array was empty before the cleanup. The
--mention <pubkey>workaround is no longer needed.users get --namestill returns both pubkeys afterwardsExpected, and not a bug: that endpoint reads the users/profile table, not membership. Only membership drives mention resolution. Flagging it because it looks like the fix failed when it has not.
Suggestions
- Surface the ambiguity to the sender. Today an ambiguous
@Nameis accepted with an emptymention_pubkeysand no warning on either side. Returning a warning in the send response, or logging it relay-side, would have collapsed a multi-hour investigation into one message. This is the single highest-value change here. - Add
display_nametochannels members. It currently returns onlypubkeyandrole, so nothing can detect a duplicate-name condition from that endpoint without an N+1 lookup throughusers get. I wrote a cleanup tool that matched names against the members payload and it silently reported "0 orphans" forever, because the field simply is not there. - Consider a first-class cleanup path — a
buzz adminsubcommand, or having the relay refuse a channel join that would introduce a duplicate display name in that channel.
Happy to open separate issues for 1 and 2 if you would prefer them tracked apart from this one.
- Surface the ambiguity to the sender. Today an ambiguous
- added a commit that references this issue
on Aug 21, 2026 - added a commit that references this issue
on Aug 23, 2026 - added a commit that references this issue
on Oct 7, 2026
Revised hypothesis: duplicate identities break name resolution
What is confirmed (directly observed)
1. Identity is split across three stores with nothing reconciling them.
pubkeyrecorded in the local agent registry at provisioning time.Nothing verifies these agree, and nothing deregisters a stale relay registration when an identity is replaced.
2. Agents were provisioned with a pubkey but no private key. Three agents were created within 4 seconds of each other with pubkeys recorded, but no private key was ever written to the password store. That gap sat latent for 2.5 days, until the supervisor tried to start them and failed with
identity key unavailablefor exactly those three. New key material appeared 8 minutes later.For comparison, agents provisioned later had their key written 22–57 seconds after the record was created. The three earliest had a 2.5-day gap, and a fourth had an 11.5-hour gap.
If the key written later is a fresh keypair rather than a recovery of the original, the agent now announces under a new pubkey while the original pubkey remains registered on the relay indefinitely.
3. Nothing validates the key against the record. There is no step anywhere in the startup path that derives the pubkey from the loaded private key and compares it to the recorded pubkey. An agent whose key does not match its record will connect and announce under a second identity, silently.
4. The password-store entry name is a hardcoded map with three different naming conventions across eight agents (
<name>,<lowercase-name>,<lowercase-name>-key). Any rename breaks the lookup, and a failed lookup is what leads to replacement key material being generated.5. The lookup can wedge.
passspins at 100% CPU indefinitely when gpg cannot prompt for a passphrase (e.g. a non-interactive session). Reproduced accidentally during this investigation — one such process consumed a full core for 10 minutes. A comment in the operator's supervisor script records a prior incident where three of these pinned three cores. A wedged lookup presents asidentity key unavailable, which is the same signal that precedes replacement-key generation.What is reported but not independently verified
An agent with a working identity queried the community roster and reported that five display names each map to two distinct pubkeys, while three are singletons — 13 bot registrations for 8 running agents. It named a specific second pubkey for one of the duplicated display names.
I could not verify this myself: querying the roster requires an authenticated relay connection, and the credential is behind a gpg passphrase prompt unavailable in a non-interactive session. Treat this as a strong lead rather than established fact.
Why this explains the original symptom better
If a display name resolves to two pubkeys and only one has a live process behind it, then mention-by-name and DM addressing can resolve to the dead identity and fail silently — no error on either side. That accounts for the group-works / DM-fails split in the original report: group-channel delivery is by channel membership and does not depend on resolving a name to a pubkey, whereas DM addressing does.
It also explains why a full cold restart with verifiably correct subscriptions changed nothing, which never fit a "relay stopped delivering" explanation.
Suggested hardening
Original report (preserved)
Summary
After a websocket reconnect, a hosted community relay stopped delivering DM-channel messages to headless ACP agents, while continuing to deliver group-channel messages to those same agents over the same websocket, for over two hours. It survived agent reconnect, full agent restart, and creation of an entirely new DM channel.
Timeline (UTC, single day)
DM delivery working normally — 7 turns across 6 agents on 6 distinct DM channels:
Then the host running the agents came under heavy resource pressure and the websockets flapped:
The host-side resource problem was fully resolved shortly after. From 09:48:35 onward, zero DM-channel messages were delivered to any of the 8 agents, while group-channel messages continued flowing the entire time.
Steps taken to recover (none worked)
resubscribing to N channel(s) after reconnect, counts matched the pre-disconnect subscription sets including DM channels. Group messages resumed; DMs did not.discovered 5 channel(s)followed by explicitsubscribed to channel <dm-channel-id>. Group messages delivered; DMs did not.membership notification: subscribing to new channeland subscribed. A message was sent to that channel by the human user. 22+ minutes later the agent had received nothing on it.Agent-side state during the failure
Every observable client-side signal reported healthy: 8/8 established websockets; each agent logging
subscribed to channel <its own DM channel>; each loggingsubscribed to membership notifications; each loggingpresence set to online; balancedturn starting/turn completecounts; and group-channel messages from the same human sender delivered and processed normally on the same connection throughout.Note on observability
subscribed to channel Xis logged on the agent side without any indication of whether the relay acknowledged and registered the subscription. A subscription the relay never honoured is indistinguishable in the logs from one that works. Surfacing a relay-side ack — or a periodic subscription reconciliation — would make this class of failure diagnosable.