Skip to content

Ambiguous @Name mentions silently notify nobody when a display name maps to two pubkeys; agents appear online but unreachable, and 'agents archive' does not resolve the ambiguity #4303

Description

@CryptoJones

Update — revised mechanism. I originally filed this as "the relay stops delivering DM messages after a websocket reconnect." Further investigation suggests that framing is probably wrong, or at least describes a symptom rather than the cause. A more likely mechanism is duplicate agent identities: one display name registered under two distinct pubkeys, which makes mention/DM resolution by name land on a dead identity. The original observations are preserved below because they are reproducible and may still be independently useful; the revised analysis is at the top.

Revised hypothesis: duplicate identities break name resolution

What is confirmed (directly observed)

1. Identity is split across three stores with nothing reconciling them.

  • The pubkey recorded in the local agent registry at provisioning time.
  • The private key, in a password store under a hardcoded entry name.
  • The registration held by the relay once the agent first announces.

Nothing verifies these agree, and nothing deregisters a stale relay registration when an identity is replaced.

2. Agents were provisioned with a pubkey but no private key. Three agents were created within 4 seconds of each other with pubkeys recorded, but no private key was ever written to the password store. That gap sat latent for 2.5 days, until the supervisor tried to start them and failed with identity key unavailable for exactly those three. New key material appeared 8 minutes later.

For comparison, agents provisioned later had their key written 22–57 seconds after the record was created. The three earliest had a 2.5-day gap, and a fourth had an 11.5-hour gap.

If the key written later is a fresh keypair rather than a recovery of the original, the agent now announces under a new pubkey while the original pubkey remains registered on the relay indefinitely.

3. Nothing validates the key against the record. There is no step anywhere in the startup path that derives the pubkey from the loaded private key and compares it to the recorded pubkey. An agent whose key does not match its record will connect and announce under a second identity, silently.

4. The password-store entry name is a hardcoded map with three different naming conventions across eight agents (<name>, <lowercase-name>, <lowercase-name>-key). Any rename breaks the lookup, and a failed lookup is what leads to replacement key material being generated.

5. The lookup can wedge. pass spins at 100% CPU indefinitely when gpg cannot prompt for a passphrase (e.g. a non-interactive session). Reproduced accidentally during this investigation — one such process consumed a full core for 10 minutes. A comment in the operator's supervisor script records a prior incident where three of these pinned three cores. A wedged lookup presents as identity key unavailable, which is the same signal that precedes replacement-key generation.

What is reported but not independently verified

An agent with a working identity queried the community roster and reported that five display names each map to two distinct pubkeys, while three are singletons — 13 bot registrations for 8 running agents. It named a specific second pubkey for one of the duplicated display names.

I could not verify this myself: querying the roster requires an authenticated relay connection, and the credential is behind a gpg passphrase prompt unavailable in a non-interactive session. Treat this as a strong lead rather than established fact.

Why this explains the original symptom better

If a display name resolves to two pubkeys and only one has a live process behind it, then mention-by-name and DM addressing can resolve to the dead identity and fail silently — no error on either side. That accounts for the group-works / DM-fails split in the original report: group-channel delivery is by channel membership and does not depend on resolving a name to a pubkey, whereas DM addressing does.

It also explains why a full cold restart with verifiably correct subscriptions changed nothing, which never fit a "relay stopped delivering" explanation.

Suggested hardening

  1. Make provisioning atomic — write the key and the recorded pubkey together, or neither. A record with no key is the origin of this whole failure.
  2. Assert identity at startup — derive the pubkey from the loaded key, compare to the record, refuse to start on mismatch rather than announcing a second identity.
  3. Never generate replacement key material on lookup failure. A missing key should be loud and fatal for that agent only.
  4. Deregister on re-key — if an identity is legitimately replaced, remove the prior registration.
  5. Reject duplicate display names at registration, or require disambiguation, so name resolution cannot become ambiguous.
  6. Bound the credential lookup — a passphrase prompt with no tty should fail fast, not spin.

Original report (preserved)

Summary

After a websocket reconnect, a hosted community relay stopped delivering DM-channel messages to headless ACP agents, while continuing to deliver group-channel messages to those same agents over the same websocket, for over two hours. It survived agent reconnect, full agent restart, and creation of an entirely new DM channel.

Timeline (UTC, single day)

DM delivery working normally — 7 turns across 6 agents on 6 distinct DM channels:

09:22:59  agent A  turn starting for channel 4026f65c…  (DM)
09:41:45  agent B  turn starting for channel 4252c972…  (DM)
09:42:10  agent C  turn starting for channel 551a61db…  (DM)
09:42:33  agent D  turn starting for channel b5b744e2…  (DM)
09:43:31  agent E  turn starting for channel 108ee621…  (DM)
09:47:28  agent F  turn starting for channel 327df1c1…  (DM)
09:48:35  agent A  turn starting for channel 4026f65c…  (DM)   <-- last DM ever delivered

Then the host running the agents came under heavy resource pressure and the websockets flapped:

10:13:31  WARN no pong received within 10s — connection dead, reconnecting
10:14-10:30  reconnect attempts fail: "Connection closed", "Connection reset by peer"
10:32:0x  all 8 agents reconnect and resubscribe

The host-side resource problem was fully resolved shortly after. From 09:48:35 onward, zero DM-channel messages were delivered to any of the 8 agents, while group-channel messages continued flowing the entire time.

Steps taken to recover (none worked)

  1. Agent-driven reconnect — resubscribing to N channel(s) after reconnect, counts matched the pre-disconnect subscription sets including DM channels. Group messages resumed; DMs did not.
  2. Full cold restart of 7 of 8 agents. Fresh processes, fresh discovery: discovered 5 channel(s) followed by explicit subscribed to channel <dm-channel-id>. Group messages delivered; DMs did not.
  3. Brand-new DM channel. An agent received membership notification: subscribing to new channel and subscribed. A message was sent to that channel by the human user. 22+ minutes later the agent had received nothing on it.

Agent-side state during the failure

Every observable client-side signal reported healthy: 8/8 established websockets; each agent logging subscribed to channel <its own DM channel>; each logging subscribed to membership notifications; each logging presence set to online; balanced turn starting/turn complete counts; and group-channel messages from the same human sender delivered and processed normally on the same connection throughout.

Note on observability

subscribed to channel X is logged on the agent side without any indication of whether the relay acknowledged and registered the subscription. A subscription the relay never honoured is indistinguishable in the logs from one that works. Surfacing a relay-side ack — or a periodic subscription reconciliation — would make this class of failure diagnosable.

Activity

  1. changed the title [-]Relay permanently stops delivering DM-channel messages after a websocket reconnect (group channels unaffected); not recoverable by agent restart or new DM channel[/-] [+]Duplicate agent identities (one display name, two pubkeys) break DM/mention resolution silently; no validation that a stored key matches its recorded pubkey[/+] on Aug 2, 2026
  2. CryptoJones commented on Aug 2, 2026

    @CryptoJones
    Author

    Amending my own report: the original framing ("relay stops delivering DMs after a websocket reconnect") is likely wrong, and I've rewritten the issue body around a more probable mechanism — duplicate agent identities, where one display name is registered under two distinct pubkeys and name resolution silently lands on the one with no live process behind it.

    That fits the evidence better than the original theory, which never explained why a full cold restart with verifiably correct subscriptions changed nothing.

    The original observations are preserved at the bottom of the body; they're reproducible and may still be useful independently. I've also marked clearly which parts I verified directly and which came from an agent's roster query that I could not independently confirm — the credential needed to query the roster sits behind a gpg passphrase prompt I couldn't satisfy non-interactively.

    Retitled to match. Happy to close this and open a narrower one if maintainers would prefer the two concerns separated.

  3. CryptoJones commented on Aug 2, 2026

    @CryptoJones
    Author

    Follow-up: the duplicate-identity hypothesis in the issue body is now confirmed, and attempting to remediate it surfaced two further problems in agents archive.

    Duplicate identities confirmed

    Querying the community roster (users get --name <name>) returned two distinct pubkeys for five of eight display names, and one each for the other three — 13 bot registrations for 8 running agents.

    For each of the five I verified which pubkey is live rather than assuming: decrypt the agent's private key, query its own profile (users get with no --pubkey), and compare against the pubkey recorded in the local agent registry. All five matched their record, confirming the second pubkey in each pair is an orphan with no private key behind it.

    This is consistent with the provisioning gap described in the issue body — a record created with a pubkey but no stored key, key material generated later, and the original registration never removed.

    Problem 1 — agents archive reports success without applying

    Archived all five orphans as the community owner. Every request returned:

    {"action":"archive","ok":true,"event_id":"<hex>","target":"<pubkey>"}

    Three subsequently appeared in the archive snapshot (agents archived, kind 13535). Two never did. Re-submitting those two produced a fresh ok:true with a new event id and still no snapshot entry — so this is reproducible, not a propagation race.

    The CLI help notes the relay selects a consent path (self / admin / owner) from the submitted request and does not retry with a different shape. If a request is rejected because the selected path doesn't authorize it, that rejection isn't surfaced: the caller gets ok:true and a signed event id, and can only detect the no-op by polling the snapshot afterwards.

    Whatever the underlying cause, ok:true on a request that is not applied is the reportable part — a caller cannot distinguish success from silent failure.

    Problem 2 — archived identities still appear in name lookup

    For the three orphans that did land in the archive snapshot, users get --name <name> continues to return both pubkeys.

    If name lookup is intentionally a raw profile search that ignores archive state, that's defensible — but then agents archive isn't a remedy for duplicate display names, and there appears to be no other mechanism for retiring a superseded identity short of moderation ban, which is semantically wrong for an identity that was simply replaced by a re-key.

    If mention/DM resolution shares this lookup path, archiving doesn't address the original failure at all.

    Suggested

    • agents archive should return a failure when the request isn't applied, rather than ok:true with an event id.
    • Archive state should be reflected in identity lookup — or the docs should state plainly that it isn't, and name the intended mechanism for retiring a superseded identity.
    • Ideally, registering a second identity under an existing display name should be rejected or require disambiguation, so this state can't arise in the first place.

    Pubkeys and relay host omitted here; happy to supply the event ids privately if useful for tracing the two no-op requests.

  4. CryptoJones commented on Aug 2, 2026

    @CryptoJones
    Author

    Mechanism now confirmed end-to-end, and it reframes this issue. The delivery layer is fine — mentions of a duplicated display name are silently dropped at resolution time.

    The confirming experiment

    Three agents were @Name mentioned by the community owner in one channel, seconds apart. Only one replied. Minutes later, one of the silent two was mentioned in the same channel with an explicit pubkey:

    messages send --channel <ch> --content "@Neuromancer diagnostic ping ..." \
                  --mention <live-pubkey>
    

    Response was immediate:

    13:41:44  turn starting for channel <ch>
    13:41:55  turn complete — replied in-thread
    

    Same agent, same channel, same process, ~5 minutes apart. Ignored by name, answered by pubkey. The agent was healthy the entire time: websocket established, subscribed to that channel, presence set to online, balanced turn counts.

    Why

    From messages send --help:

    Supplying any explicit identity permits unresolved or ambiguous @name text as presentation-only; uniquely resolved member names still notify

    So an @Name matching two pubkeys is ambiguous, and an ambiguous mention does not notify anyone. That is documented behavior. Combined with the duplicate registrations described earlier in this issue, the result is:

    • 5 of 8 display names map to 2 pubkeys each → unreachable by name
    • 3 of 8 are singletons → work perfectly

    That split held exactly, all day, and was the entire user-visible symptom.

    It also explains the DM half. A client resolving one of the duplicated names picks the orphan pubkey; the orphan has no process behind it and therefore no presence; the client renders the agent offline and won't open a DM to it. From the user's side: "the bots are offline and ignoring me," while every server-side and agent-side signal reads healthy.

    Severity

    The failure is completely silent in both directions. The sender gets an accepted message with no warning that the mention resolved to nothing; the agent never learns a message existed. There is no error, no log line, and no metric on either side. Diagnosing it required reconstructing per-channel turn history from agent logs and then A/B-testing name vs pubkey mentions — several hours for a condition the relay already knows about at resolution time.

    The remedy in the tooling does not work

    agents archive is the obvious way to retire a superseded identity, and it does not resolve the ambiguity: three orphans were successfully archived (confirmed present in the kind 13535 snapshot) and users get --name still returns both pubkeys for all three afterwards. Two further orphans could not be archived at all — see the previous comment, ok:true with no effect.

    So there is currently no working path from "I have duplicate display names" to "mentions work again", short of moderation ban, which is semantically wrong for an identity that was merely superseded by a re-key.

    Suggested, in priority order

    1. Warn on ambiguous mention resolution. If @Name resolves to more than one member, that should surface to the sender — a response field, a system message, anything. Silently downgrading to presentation-only is the core defect; everything else here is recoverable if the operator can see it happening.
    2. Reject duplicate display names at registration, or require disambiguation, so this state cannot arise.
    3. Make agents archive exclude the identity from name resolution — otherwise archiving is not a remedy for the one problem it most obviously appears to address.
    4. Log ambiguous-resolution events relay-side so operators can find them without A/B testing.

    Happy to supply event ids or the raw agent logs for the confirming experiment if useful.

  5. changed the title [-]Duplicate agent identities (one display name, two pubkeys) break DM/mention resolution silently; no validation that a stored key matches its recorded pubkey[/-] [+]Ambiguous @Name mentions silently notify nobody when a display name maps to two pubkeys; agents appear online but unreachable, and 'agents archive' does not resolve the ambiguity[/+] on Aug 2, 2026
  6. CryptoJones commented on Aug 2, 2026

    @CryptoJones
    Author

    Root cause found and fixed for the client half: #4371.

    _resolveMentions in mobile/lib/features/channels/send_message_provider.dart took the first profile whose display name matched and breaked. With no ambiguity detection and cache being a Map, "first match" meant "whichever entry the map yielded first" — and that decided which pubkey went into the p tag.

    That explains the detail I could not account for earlier in this thread: why one duplicated agent answered @Name while the others never did. It was map ordering. Names whose live profile happened to enumerate first worked; the rest tagged a dead identity, and since agents match purely on p tags, the intended agent never saw the mention.

    The PR collects all matching candidates and tags them all, with a regression test verified to fail against the previous behaviour.

    Two relay-side items from this issue are not addressed by that PR and remain open:

    1. agents archive returning ok:true with a signed event id while never applying — reproducible on two specific identities, twice each.
    2. Archived identities still appearing in users get --name.

    Happy to split those into their own issue if you'd prefer this one scoped to the mention-resolution bug now that it has a fix attached.

  7. CryptoJones commented on Aug 2, 2026

    @CryptoJones
    Author

    Resolved on our deployment — and the mechanism is narrower than this issue states

    Fixed this today for 5 of 8 agents. Posting the corrected mechanism plus two dead ends, because both cost real time and neither is obvious from the code.

    It is CHANNEL membership that creates the ambiguity, not relay membership

    crates/buzz-relay/src/workflow_sink.rs builds the name map from channel members:

    get_members(community, channel_uuid)   // channel membership
    get_users_bulk(community, &pubkeys)    // display_name per pubkey
    resolve_mention_pubkeys(&text, &named_members)

    resolve_mention_pubkeys then folds name -> pubkey and sets the slot to None when two distinct pubkeys share a name, so no p tag is emitted. The behaviour is deliberate and documented in that function, and I agree with the design choice — arbitrary selection misroutes, tagging all is a false-wake firehose.

    The consequence is that the ambiguity is per channel. A stale twin only breaks mentions in channels it is still a member of.

    Dead end 1: agents archive does not help

    Confirmed. It never touches channel membership, so both twins remain members and the name stays ambiguous. Two of our orphans also returned ok:true with a signed event id and never entered the kind:13535 snapshot at all — a silent no-op, which is worth a separate look.

    Dead end 2: kind 9031 RELAY_ADMIN_REMOVE_MEMBER does not apply

    I built a signer for kind 9031 (buzz-core/src/kind.rs, handler handlers/relay_admin.rs) assuming the twins were relay members. Every request returned:

    400 invalid: member not found: <orphan pubkey>
    

    They were never in relay_members. Worth noting for anyone else who tries this path: POST /events also requires NIP-98 auth (Authorization: Nostr <base64 kind-27235>), and without it you get 401 missing Nostr auth — a different failure that is easy to mistake for the same problem.

    What actually fixes it

    buzz channels remove-member --channel <uuid> --pubkey <orphan-pubkey>
    

    signed with an owner/admin key — agent keys cannot remove members. We removed 22 orphan memberships across 6 channels; all returned accepted:true.

    Verification, same channel, name-only mention:

    {"accepted":true,"mention_pubkeys":["<live-pubkey>"]}

    That array was empty before the cleanup. The --mention <pubkey> workaround is no longer needed.

    users get --name still returns both pubkeys afterwards

    Expected, and not a bug: that endpoint reads the users/profile table, not membership. Only membership drives mention resolution. Flagging it because it looks like the fix failed when it has not.

    Suggestions

    1. Surface the ambiguity to the sender. Today an ambiguous @Name is accepted with an empty mention_pubkeys and no warning on either side. Returning a warning in the send response, or logging it relay-side, would have collapsed a multi-hour investigation into one message. This is the single highest-value change here.
    2. Add display_name to channels members. It currently returns only pubkey and role, so nothing can detect a duplicate-name condition from that endpoint without an N+1 lookup through users get. I wrote a cleanup tool that matched names against the members payload and it silently reported "0 orphans" forever, because the field simply is not there.
    3. Consider a first-class cleanup path — a buzz admin subcommand, or having the relay refuse a channel join that would introduce a duplicate display name in that channel.

    Happy to open separate issues for 1 and 2 if you would prefer them tracked apart from this one.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions