Skip to content

[meilisearch] use the rendered heading anchors in search result URLs - #844

Merged
mishig25 merged 1 commit into
mainfrom
fix-search-anchors
Sep 30, 2026
Merged

mishig25 merged 1 commit into
mainfrom
fix-search-anchors

Conversation

@mishig25

Copy link
Copy Markdown
Contributor

Problem

source_page_url anchors in docs-semantic-search were slugified from the raw heading instead of matching the ids the docs renderer gives headings, so many search results link to anchors that don't exist on the page and open its top instead of the section:

  • headings with a custom anchor: token_to_word[[transformers.BatchEncoding.token_to_word]] gave #tokentowordtransformersbatchencodingtokentoword instead of #transformers.BatchEncoding.token_to_word
  • plain headings: slugify drops hyphens, e.g. Low-rank adaptation gave #lowrank-adaptation instead of #low-rank-adaptation

Changes

  • heading_anchor: the id the renderer gives a heading, i.e. its trailing [[anchor]] as-is (dunders like __call__ included), otherwise the slug check_links already computes for it. Used by both URL builders (process_markdown_file and chunks_to_documents).
  • migrations/fix_source_page_urls.py: backfills existing documents. It processes the docs like populate-search-engine does and gives each document the URL of the chunk with the same ID. IDs only hash text, so they match, nothing is re-embedded and the incremental-update tracker stays valid.
  • clean_heading now only keeps __ markup (dunders), so _Adapters_ becomes Adapters.
  • update_all_documents:
    • waits for the index's pending tasks before scanning: partial updates create missing documents, so documents deleted by a still-queued ingestion run came back as stubs (329 of them had to be deleted after the heading backfill)
    • waits up to a day for its own tasks instead of 10 minutes, since the instance-wide task queue can take hours

Testing

  • Unit tests for anchors (custom, multiple, dunder, hyphenated, numbered headings), clean_heading, and the pending-task wait.
  • Dry run against the live index: 75,534 of 76,442 documents matched by ID, 14,094 URLs would change. Checked samples against the live pages:
    • custom anchors: 98/100 new anchors exist (0/100 old ones did); both misses are lerobot pages that differ from the doc-build dataset
    • other headings: 39/50 new anchors exist (0/50 old ones did); the misses are code comments that split_markdown_by_headings treats as headings

After merging

Run once (no populate-search-engine job running), then re-run clean_headings.py for the _Adapters_-style headings:

uv run python migrations/fix_source_page_urls.py --meilisearch_url <url> --meilisearch_key <key>
uv run python migrations/clean_headings.py --meilisearch_url <url> --meilisearch_key <key>

Not changed here

split_markdown_by_headings splits sections on # lines inside code blocks. Fixing it changes chunk texts, and so IDs, of the affected pages, so it's left for a separate PR.

This PR was generated with an AI coding agent.

🤖 Generated with Claude Code

- `source_page_url` anchors are now the ids the renderer gives headings (`_heading_anchor` from
  check_links), e.g. `#transformers.BatchEncoding.token_to_word` instead of
  `#tokentowordtransformersbatchencodingtokentoword`, which doesn't exist on the page
- `migrations/fix_source_page_urls.py` backfills the existing documents
- `clean_heading` only keeps `__` markup (dunders), so `_Adapters_` becomes `Adapters`
- `update_all_documents` waits for pending index tasks before scanning, so documents deleted meanwhile
  aren't recreated as stubs, and waits for its own tasks for up to a day

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@mishig25
mishig25 merged commit bc117cf into main Sep 30, 2026
4 checks passed
@mishig25
mishig25 deleted the fix-search-anchors branch September 30, 2026 17:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant