[meilisearch] keep the incremental-update tracker in sync with the index - #845
Merged
Merged
Conversation
- wait for stale-document deletions to succeed, and keep the IDs of failed deletions in the tracker so the next run retries them (the tracker used to be saved before deletions were applied) - rebuild the tracker from the main index after `meilisearch-clean --swap`, since full rebuilds didn't update it Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
docs-semantic-searchcontains 259 documents that aren't in the incremental-update tracker (hf-doc-build/doc-builder-embeddings-tracker). Incremental runs only delete tracked IDs, so these stale documents are never removed and keep showing up in search results (with outdated content and URLs). All tracked IDs are in the index, so the drift only goes one way. Two code paths can cause it:_run_incrementalqueued the stale-document deletions without waiting for them, then saved the tracker without those IDs. If a deletion task failed or never ran, its documents stayed in the index while the tracker forgot them.populate-search-enginewithout--incremental, thenmeilisearch-clean --swap) never updated the tracker.Changes
delete_documents_from_dbwaits for its deletion task and raises if it fails (likeadd_embeddings_to_dbalready does)._run_incrementalalways saves the tracker, but keeps the IDs whose deletion didn't succeed, so the next run retries them.meilisearch-clean --swaprebuilds the tracker from the main index after the swap. It now requires--hf_token(orHF_TOKEN) when swapping.Incremental runs now also wait for their deletion tasks, which can take a while when the Meilisearch task queue is busy (they already wait for their additions).
Testing
Unit tests for the deletion wait, a failed deletion batch staying in the tracker, and the tracker rebuild after a swap (with and without
--swap).Existing stale documents
To clean up the 259 documents already in the index: run
migrations/export_meili_ids_to_hf.pyonce to re-bootstrap the tracker from the index. The next incremental run then deletes every document that no current page produces.This PR was generated with an AI coding agent.
🤖 Generated with Claude Code