Repository navigation
fix(indexer): fingerprint chunkers at import time, pass language overrides as memo arg - #299
Merged
Merged
Conversation
…memo arg #287 hashed each chunker's source file from disk on every index run, but the daemon runs the code it imported at startup. Editing a chunker and indexing before a restart stored the old code's chunks under the new fingerprint, and they stayed after the restart. A functools.partial or callable instance fell back to repr(), whose memory address re-indexed everything on every restart. - The daemon records a fingerprint per suffix when it imports a chunker: the "module:attr" spec plus a sha256 of the module file. The digest is kept per loaded module, so a project loaded again in the same daemon (ccc reset, a second project) gets the digest of the code that actually runs. Fingerprints reach process_file through a new CHUNKER_FINGERPRINTS context key with detect_change=True; CHUNKER_REGISTRY is unchanged. - indexer_main builds the suffix -> language map once per run and passes it to process_file, which uses it; the per-file load_project_settings is gone. - Remove chunking_fingerprint, _chunker_source_digest and the unused chunking_config argument. - Move the duplicated test stub embedder into a conftest fixture. - README: any chunking change re-chunks all files; unchanged chunk text keeps its cached embedding; only the chunker's own module file is hashed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Project.create took the chunker functions and their fingerprints as two
dicts keyed by the same suffixes, so a caller could register chunkers
without fingerprints and silently keep stale chunks. It now takes one
{suffix: LoadedChunker(fn, fingerprint)} mapping and provides both context
keys from it; _resolve_chunker_registry returns that mapping directly.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Instead of a separate CHUNKER_FINGERPRINTS context key that process_file read only to record a dependency, CHUNKER_REGISTRY now holds LoadedChunker values and has detect_change=True. LoadedChunker.__coco_memo_key__ returns (spec, module_sha256) and never the function: cocoindex fingerprints a function by module and name only and cannot fingerprint lambdas or closures. CHUNKER_REGISTRY's value type changes from dict[str, ChunkerFn] to dict[str, LoadedChunker]; LoadedChunker is exported from chunking. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On Windows, write_text turns "\n" into "\r\n", so the chunk contents the tests compare exactly came back as "hello\r\n". This already failed the Windows CI job on main after #287. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
georgeh0
force-pushed
the
fix/chunking-config-memo
branch
from
October 6, 2026 00:47
384303b to
8c17404
Compare
ruff format and end-of-file fixes in the daemon socket guard and its test. Every Pre-commit CI job on main failed at the lint step, which also kept pytest from running for any PR. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
It never ran daemon code: it re-implemented the (st_dev, st_ino) comparison on touch()ed files. On ext4 the replacement file reuses the unlinked inode number, so its premise fails and the test fails on ubuntu. #300 removes it too and replaces the guard it was meant to cover. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 6, 2026
georgeh0
added a commit
that referenced
this pull request
Oct 6, 2026
) INDEXING_EMBED_PARAMS was not a change-detected context key, so editing embedding.indexing_params (e.g. adding `prompt_name: passage`) and restarting the daemon left files with unchanged content on their memoized process_file results, embedded with the old params. LiteLLM was already covered because its constructor kwargs are part of the embedder memo key; sentence-transformers was not. - Declare INDEXING_EMBED_PARAMS with detect_change=True. QUERY_EMBED_PARAMS stays as is: it is read only at query time, outside any memoized function. - StubEmbedder in conftest records the kwargs of each embed() call; the duplicate _KwargRecordingEmbedder is gone. - Move the registry index helpers from the chunking test into conftest. - New end-to-end test: adding, changing, or removing a param reprocesses unchanged files with the new kwargs; equal params across a restart (in any key order) leave every file unchanged. Memo entries written before this change don't record the params dependency. Upgrading from v0.2.42 still re-runs every file once, since #299 added a process_file argument, so this must ship in the same release. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #287 (issue #285). Fixes two bugs in how #287 detects chunking-config changes and makes the language-override path use its input directly.
Bugs fixed
ccc indexbefore a daemon restart, the old code ran and its chunks were memoized under the new fingerprint. After the restart the fingerprint was unchanged, so the stale chunks stayed for good.functools.partialor a callable instance,inspect.getsourcefileraisesTypeError, and therepr(fn)fallback contains a memory address.Changes
_resolve_chunker_registryimports a chunker, it wraps it inLoadedChunker(fn, spec, module_sha256): the"module:attr"string fromsettings.ymlplus a sha256 of the imported module's file. The digest is cached per loaded module object. Without that,ccc resetor loading a second project in the same daemon would hash the edited file whileimportlibreturns the module it already loaded, which reproduces bug 1.CHUNKER_REGISTRYis now change-detected (detect_change=True). Its value type changes fromdict[str, ChunkerFn]todict[str, LoadedChunker], andLoadedChunkeris exported fromcocoindex_code.chunking.LoadedChunker.__coco_memo_key__returns(spec, module_sha256)and never the function. cocoindex fingerprints a function by module and qualname only, so a body edit wouldn't register, and lambdas and closures can't be fingerprinted at all.process_filealready reads the registry, so no separate dependency is needed.Project.create(chunker_registry=...)takes the same mapping.process_fileargument.indexer_mainbuilds the suffix → language map once per run, andprocess_fileuses it, so it is part of the memo key. The per-fileload_project_settingscall is gone. Override changes still take effect on the next run without a restart.chunking_fingerprint,_chunker_source_digest, the unusedchunking_configargument, and the__all__entry.Behavior notes
language_overridesre-chunks every file, not only the files that entry applies to. Chunks whose text didn't change reuse their cached embeddings, becauseembedis memoized. I checked this: after a fingerprint change the file is reprocessed butembedisn't called again.module:is hashed. An edit to a helper module it imports isn't detected (documented in the README).process_file's arguments changed, so the firstccc indexafter upgrading re-chunks every file once. Embeddings come from the cache.Tests
All in
tests/test_reindex_on_chunking_config_change.py. The existing two end-to-end tests still pass.test_chunker_change_reprocesses_unchanged_filesnow builds its chunkers with the daemon's_resolve_chunker_registryinstead of passing bare functions.mainbefore this PRfunctools.partialccc resetpath)module:to a different callable in the same unedited module reprocesses the fileThe duplicated
_StubEmbedderfromtest_chunker_registry.pyand this file is now a sharedstub_embedderfixture inconftest.py.Checks:
uv run pytest tests/(325 passed, Docker e2e deselected),uv run ruff check .,uv run ruff format --check ., anduv run mypy .. mypy reports only the 2 errors inscripts/find_best_models.pythat already exist onmain.Also in this PR (CI fixes)
mainby fix(daemon): guard socket unlinking on shutdown by checking bound node identity #291: ruff format and end-of-file fixes indaemon.pyandtests/test_daemon.py. Every Pre-commit job onmainfailed at lint, which kept pytest from running.write_bytes. On Windows,write_textproduced\r\n, which failed fix(indexer): reprocess unchanged files when chunking config changes #287's test onmain.continue-on-error), the in-process restart tests can hitenvironment already open in this program. Garbage collection doesn't free cocoindex's LMDB environment before the test reopens it, and cocoindex has no explicit close. fix(indexer): reprocess unchanged files when chunking config changes #287's restart test already fails this way onmain.Follow-up (out of scope)
INDEXING_EMBED_PARAMSinshared.pyalso lacksdetect_change=True, so changing the indexing embed params (for exampleprompt_name) doesn't re-embed files whose content is unchanged.🤖 Generated with Claude Code