Repository navigation
[meilisearch] use the rendered heading anchors in search result URLs - #844
Merged
Merged
Conversation
- `source_page_url` anchors are now the ids the renderer gives headings (`_heading_anchor` from check_links), e.g. `#transformers.BatchEncoding.token_to_word` instead of `#tokentowordtransformersbatchencodingtokentoword`, which doesn't exist on the page - `migrations/fix_source_page_urls.py` backfills the existing documents - `clean_heading` only keeps `__` markup (dunders), so `_Adapters_` becomes `Adapters` - `update_all_documents` waits for pending index tasks before scanning, so documents deleted meanwhile aren't recreated as stubs, and waits for its own tasks for up to a day Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
source_page_urlanchors indocs-semantic-searchwere slugified from the raw heading instead of matching the ids the docs renderer gives headings, so many search results link to anchors that don't exist on the page and open its top instead of the section:token_to_word[[transformers.BatchEncoding.token_to_word]]gave#tokentowordtransformersbatchencodingtokentowordinstead of#transformers.BatchEncoding.token_to_wordslugifydrops hyphens, e.g.Low-rank adaptationgave#lowrank-adaptationinstead of#low-rank-adaptationChanges
heading_anchor: the id the renderer gives a heading, i.e. its trailing[[anchor]]as-is (dunders like__call__included), otherwise the slugcheck_linksalready computes for it. Used by both URL builders (process_markdown_fileandchunks_to_documents).migrations/fix_source_page_urls.py: backfills existing documents. It processes the docs likepopulate-search-enginedoes and gives each document the URL of the chunk with the same ID. IDs only hashtext, so they match, nothing is re-embedded and the incremental-update tracker stays valid.clean_headingnow only keeps__markup (dunders), so_Adapters_becomesAdapters.update_all_documents:Testing
clean_heading, and the pending-task wait.split_markdown_by_headingstreats as headingsAfter merging
Run once (no populate-search-engine job running), then re-run
clean_headings.pyfor the_Adapters_-style headings:Not changed here
split_markdown_by_headingssplits sections on#lines inside code blocks. Fixing it changes chunk texts, and so IDs, of the affected pages, so it's left for a separate PR.This PR was generated with an AI coding agent.
🤖 Generated with Claude Code