Skip to content

[4.41/C2b] Shared text utilities: bare-mention scanner, code-fence helpers, to_plain_text structural whitespace #193

Description

@patrick-chinchill

Summary

Port upstream's shared text utilities into chat_sdk/shared: a code/URL-aware bare-mention scanner (replace_bare_mentions, mask_code_spans) to replace naive @(\w+) regexes; normalize_code_fences so CommonMark stops swallowing the first line of an incoming fenced block; structural whitespace in ast_to_plain_text (blank line between paragraphs, newline-separated list items/blockquote children, tab-separated cells, empty rows dropped). Utilities only — adapters adopt them later.

Upstream changes

Current Python behavior

  • grep -rn 'replace_bare_mentions\|mask_code_spans\|normalize_code_fences' src/ is empty; src/chat_sdk/shared/ has no mentions.py/code_fences.py.
  • Naive regexes: adapters/teams/format_converter.py:122/:133 re.sub(r"@(\w+)", r"<at>\1</at>", …); adapters/discord/format_converter.py:94/:105 re.sub(r"@(\w+)", r"<@\1>", …); Slack has a URL-only version (adapters/slack/format_converter.py:40 BARE_MENTION_REGEX, :49 _link_bare_mentions_outside_urls) that doesn't skip code.
  • src/chat_sdk/shared/markdown_parser.py:1083-1114 ast_to_plain_text: root/blockquote/list join with "\n", listItem with " ", table/tableRow fall through to "".join. On main:
    • '| Name | Middle | Units |\n| --- | --- | --- |\n| Samsung | | 3 |' → 'NameMiddleUnitsSamsung3'
    • '@test-bot\n\nhi there' → '@test-bot\nhi there'
    • 'a\n\n> q1\n>\n> q2' → 'a\nq1\nq2'
  • The parser already emits tableRow/tableCell; empty cells are {"type":"tableCell","children":[]}. table_to_ascii (markdown_parser.py:1122, caller at :1139) calls ast_to_plain_text per cell — its output must not change.
  • Base extractor BaseFormatConverter.extract_plain_text (shared/base_format_converter.py:170-175) feeds Teams inbound (teams/adapter.py:1213), GitHub inbound (github/adapter.py:524, :554) and Discord _parse_discord_message (discord/adapter.py:1396; fetch/parse paths :1184, :1295). Discord's live gateway path uses raw content (:742) — unaffected; Slack (slack/format_converter.py:180) and GChat (google_chat/format_converter.py:221) override the extractor — unaffected. Telegram uses ast_to_plain_text for outbound/sent-message text (telegram/adapter.py:215, :1954, :2920-2994); so does thread.py:210-234 (thread.post result).

Scope

  • src/chat_sdk/shared/mentions.py, a character-for-character port of mentions.ts at chat@4.41.1: MentionReplacer = Callable[[str, str], str] (full mention, bare name); replace_bare_mentions(text, replacer) -> str; mask_code_spans(text, replacement=" ") -> str; private _is_letter, _is_number, _is_word, _is_host, _is_boundary, _starts_with, _find_url_end, _find_host_end, _find_code_end, _replace_range(…, angles).
  • src/chat_sdk/shared/code_fences.py, a port of code-fences.ts: normalize_code_fences(text, *, convert_text=None, convert_code=None) -> str with the same BLOCK_MARKER_PATTERN, ORDERED_LIST_MARKER_PATTERN, LEADING_WHITESPACE_PATTERN and escaping rules.
  • Export the three functions (and MentionReplacer) from chat_sdk/shared/__init__.py and __all__.
  • Rewrite ast_to_plain_text in shared/markdown_parser.py as a _plain_text_node / _child_plain_text(node, sep, keep_empty=False) pair per the upstream rules (table: "\n".join(r for r in rows if r.strip()); tableRow: "\t".join(cells) keeping empties); keep the public name.
  • Update existing expectations that change, without adding duplicates — e.g. tests/test_thread_faithful.py:823 "hello.\nhow are you?" → "hello.\n\nhow are you?".

Out of scope

Porting notes

  • ASCII classes: upstream isLetter/isNumber use char codes; use explicit range checks, not Unicode-aware str.isalpha()/isdigit().
  • Boundary: isBoundary is char === "<" || char === ">" || char.trim() === "". ch.strip() == "" differs from JS trim() on U+FEFF (JS whitespace, Python not) and U+001C–U+001F (Python whitespace, JS not); match JS explicitly and cover both in a test.
  • Case-insensitive prefix: upstream's startsWith(text, index, value) helper lowercases the slice (so HTTPS:// counts as a URL). Port it as _starts_with; do not substitute case-sensitive str.startswith.
  • Out-of-range indexing: JS text[index - 1] at index == 0 is undefined (isWord(undefined) → false); Python text[-1] wraps. Bounds-check every lookbehind/lookahead; sweep mentions at position 0 and at end of string.
  • Keep: inline spans don't cross newlines; unterminated fences are plain text. Don't regex-ify the scanner — recursive <…> handling (angles=False on re-entry) prevents double-wrapping. Apply SELF_REVIEW.md input sweeps and emit/parse symmetry.

Tests

Fidelity-mapped (scripts/verify_test_fidelity.py MAPPING):

  • packages/chat/src/markdown.test.ts → tests/test_markdown_faithful.py (all 8 missing today): "preserves soft line breaks as whitespace"; "preserves paragraph boundaries as whitespace"; "separates list items with whitespace"; "separates table cells and rows with structural whitespace"; "keeps empty table cells so columns stay aligned"; "drops table rows with no content"; "preserves whitespace after a newline-separated mention"; "preserves whitespace after a paragraph-separated mention".
  • packages/chat/src/chat.test.ts → tests/test_chat_faithful.py: "should call onNewMention for newline-separated GitHub bot mentions"; "should not call onNewMention for concatenated GitHub bot mention text".
  • packages/chat/src/thread.test.ts: update "should preserve double newlines (paragraph breaks) in streamed text" (test_thread_faithful.py:801).

Unmapped, port by exact name:

  • packages/adapter-shared/src/mentions.test.ts → new tests/test_shared_mentions.py: all 17 replaceBareMentions + 6 maskCodeSpans cases (e.g. "does not turn an email address into a mention", "keeps a token after an unterminated fence").
  • packages/adapter-shared/src/code-fences.test.ts → new tests/test_shared_code_fences.py: all 11 cases (e.g. "puts fences on their own lines so the first code line survives").
  • GitHub adapter (5c926f19): "should preserve whitespace in newline-separated issue comment mentions", "should preserve whitespace in newline-separated review comment mentions" (index.test.ts); "should preserve whitespace after newline-separated @mentions" (markdown.test.ts).

Python-specific: index-boundary sweeps (@a at 0, trailing @, a lone backtick); é@x confirming the ASCII word class; HTTPS://host/@x untouched; table_to_ascii output unchanged.

Acceptance criteria

  • Full validation command from CLAUDE.md passes.
  • markdown.test.ts fidelity misses against chat@4.41.1 are 0 for the 8 tests above.
  • docs/UPSTREAM_SYNC.md records the mapping adapter-shared/src/{mentions,code-fences}.ts → chat_sdk/shared/{mentions,code_fences}.py and any boundary-character gaps.
  • CHANGELOG under "Unreleased (4.41 wave)": inbound message.text changes for Teams and GitHub, Discord fetched/parsed messages, and thread.post(...) results — blank lines between paragraphs, newline-separated list items, tab-separated tables. Slack, GChat and Discord live-gateway text unchanged by this PR.

Dependencies

Blocked by #185. Blocks #206, #209, #229, #239, #203.

Metadata

Part of #184.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions