Skip to content

fix(indexer): fingerprint chunkers at import time, pass language overrides as memo arg - #299

Merged
georgeh0 merged 6 commits into
mainfrom
fix/chunking-config-memo
Oct 6, 2026
Merged

georgeh0 merged 6 commits into
mainfrom
fix/chunking-config-memo

Conversation

@georgeh0

@georgeh0 georgeh0 commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

Follow-up to #287 (issue #285). Fixes two bugs in how #287 detects chunking-config changes and makes the language-override path use its input directly.

Bugs fixed

  1. The fingerprint could drift from the running code. fix(indexer): reprocess unchanged files when chunking config changes #287 hashed each chunker's source file from disk on every index run, but the daemon runs the code it imported at startup. If you edited a chunker and ran ccc index before a daemon restart, the old code ran and its chunks were memoized under the new fingerprint. After the restart the fingerprint was unchanged, so the stale chunks stayed for good.
  2. Some chunkers re-indexed everything on every daemon restart. For a functools.partial or a callable instance, inspect.getsourcefile raises TypeError, and the repr(fn) fallback contains a memory address.

Changes

  • Chunkers are fingerprinted when they are imported. Right after _resolve_chunker_registry imports a chunker, it wraps it in LoadedChunker(fn, spec, module_sha256): the "module:attr" string from settings.yml plus a sha256 of the imported module's file. The digest is cached per loaded module object. Without that, ccc reset or loading a second project in the same daemon would hash the edited file while importlib returns the module it already loaded, which reproduces bug 1.
  • CHUNKER_REGISTRY is now change-detected (detect_change=True). Its value type changes from dict[str, ChunkerFn] to dict[str, LoadedChunker], and LoadedChunker is exported from cocoindex_code.chunking. LoadedChunker.__coco_memo_key__ returns (spec, module_sha256) and never the function. cocoindex fingerprints a function by module and qualname only, so a body edit wouldn't register, and lambdas and closures can't be fingerprinted at all. process_file already reads the registry, so no separate dependency is needed. Project.create(chunker_registry=...) takes the same mapping.
  • Language overrides are a real process_file argument. indexer_main builds the suffix → language map once per run, and process_file uses it, so it is part of the memo key. The per-file load_project_settings call is gone. Override changes still take effect on the next run without a restart.
  • Deleted chunking_fingerprint, _chunker_source_digest, the unused chunking_config argument, and the __all__ entry.
  • README: the chunking-config paragraph now describes the actual behavior.

Behavior notes

  • Granularity: any change to a chunker or to language_overrides re-chunks every file, not only the files that entry applies to. Chunks whose text didn't change reuse their cached embeddings, because embed is memoized. I checked this: after a fingerprint change the file is reprocessed but embed isn't called again.
  • Limit: only the module file named in module: is hashed. An edit to a helper module it imports isn't detected (documented in the README).
  • Upgrade: process_file's arguments changed, so the first ccc index after upgrading re-chunks every file once. Embeddings come from the cache.

Tests

All in tests/test_reindex_on_chunking_config_change.py. The existing two end-to-end tests still pass. test_chunker_change_reprocesses_unchanged_files now builds its chunkers with the daemon's _resolve_chunker_registry instead of passing bare functions.

New test On main before this PR
Same config: second index run and daemon restart report all files unchanged (plain function) passes (no bug for plain functions)
… same, functools.partial fails: all 3 files reprocessed after restart
… same, callable instance fails: all 3 files reprocessed after restart
Edit the chunker module, index before restart, restart, index → chunks reflect the new code fails: stale v1 chunks after restart
… same, but the project is reloaded in the same daemon before the restart (ccc reset path) fails: stale v1 chunks after restart. Also fails with the fix if the per-module digest cache is removed
Changing module: to a different callable in the same unedited module reprocesses the file passes (#287 already covered spec changes; this guards the new fingerprint)

The duplicated _StubEmbedder from test_chunker_registry.py and this file is now a shared stub_embedder fixture in conftest.py.

Checks: uv run pytest tests/ (325 passed, Docker e2e deselected), uv run ruff check ., uv run ruff format --check ., and uv run mypy .. mypy reports only the 2 errors in scripts/find_best_models.py that already exist on main.

Also in this PR (CI fixes)

Follow-up (out of scope)

INDEXING_EMBED_PARAMS in shared.py also lacks detect_change=True, so changing the indexing embed params (for example prompt_name) doesn't re-embed files whose content is unchanged.

🤖 Generated with Claude Code

georgeh0 and others added 4 commits October 5, 2026 17:33
…memo arg

#287 hashed each chunker's source file from disk on every index run, but the
daemon runs the code it imported at startup. Editing a chunker and indexing
before a restart stored the old code's chunks under the new fingerprint, and
they stayed after the restart. A functools.partial or callable instance fell
back to repr(), whose memory address re-indexed everything on every restart.

- The daemon records a fingerprint per suffix when it imports a chunker: the
  "module:attr" spec plus a sha256 of the module file. The digest is kept per
  loaded module, so a project loaded again in the same daemon (ccc reset, a
  second project) gets the digest of the code that actually runs.
  Fingerprints reach process_file through a new CHUNKER_FINGERPRINTS context
  key with detect_change=True; CHUNKER_REGISTRY is unchanged.
- indexer_main builds the suffix -> language map once per run and passes it
  to process_file, which uses it; the per-file load_project_settings is gone.
- Remove chunking_fingerprint, _chunker_source_digest and the unused
  chunking_config argument.
- Move the duplicated test stub embedder into a conftest fixture.
- README: any chunking change re-chunks all files; unchanged chunk text keeps
  its cached embedding; only the chunker's own module file is hashed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Project.create took the chunker functions and their fingerprints as two
dicts keyed by the same suffixes, so a caller could register chunkers
without fingerprints and silently keep stale chunks. It now takes one
{suffix: LoadedChunker(fn, fingerprint)} mapping and provides both context
keys from it; _resolve_chunker_registry returns that mapping directly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Instead of a separate CHUNKER_FINGERPRINTS context key that process_file
read only to record a dependency, CHUNKER_REGISTRY now holds LoadedChunker
values and has detect_change=True. LoadedChunker.__coco_memo_key__ returns
(spec, module_sha256) and never the function: cocoindex fingerprints a
function by module and name only and cannot fingerprint lambdas or closures.

CHUNKER_REGISTRY's value type changes from dict[str, ChunkerFn] to
dict[str, LoadedChunker]; LoadedChunker is exported from chunking.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On Windows, write_text turns "\n" into "\r\n", so the chunk contents the
tests compare exactly came back as "hello\r\n". This already failed the
Windows CI job on main after #287.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@georgeh0
georgeh0 force-pushed the fix/chunking-config-memo branch from 384303b to 8c17404 Compare October 6, 2026 00:47
georgeh0 and others added 2 commits October 5, 2026 18:10
ruff format and end-of-file fixes in the daemon socket guard and its test.
Every Pre-commit CI job on main failed at the lint step, which also kept
pytest from running for any PR.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
It never ran daemon code: it re-implemented the (st_dev, st_ino) comparison
on touch()ed files. On ext4 the replacement file reuses the unlinked inode
number, so its premise fails and the test fails on ubuntu. #300 removes it
too and replaces the guard it was meant to cover.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@georgeh0
georgeh0 merged commit 519ce97 into main Oct 6, 2026
3 of 4 checks passed
@georgeh0
georgeh0 deleted the fix/chunking-config-memo branch October 6, 2026 01:45
georgeh0 added a commit that referenced this pull request Oct 6, 2026
)

INDEXING_EMBED_PARAMS was not a change-detected context key, so editing
embedding.indexing_params (e.g. adding `prompt_name: passage`) and
restarting the daemon left files with unchanged content on their memoized
process_file results, embedded with the old params. LiteLLM was already
covered because its constructor kwargs are part of the embedder memo key;
sentence-transformers was not.

- Declare INDEXING_EMBED_PARAMS with detect_change=True.
  QUERY_EMBED_PARAMS stays as is: it is read only at query time, outside
  any memoized function.
- StubEmbedder in conftest records the kwargs of each embed() call; the
  duplicate _KwargRecordingEmbedder is gone.
- Move the registry index helpers from the chunking test into conftest.
- New end-to-end test: adding, changing, or removing a param reprocesses
  unchanged files with the new kwargs; equal params across a restart
  (in any key order) leave every file unchanged.

Memo entries written before this change don't record the params
dependency. Upgrading from v0.2.42 still re-runs every file once, since
#299 added a process_file argument, so this must ship in the same release.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant