chore: add tests for schema/chunking + ruff cleanup - #2
Conversation
- CodeChunk: creation, embedding type flexibility - QueryResult: creation, score range - Chunking exports: Chunk, TextPosition, ChunkerFn, CHUNKER_REGISTRY - 8 new tests, all passing
There was a problem hiding this comment.
Code Review
This pull request introduces a new test suite for the schema and chunking modules and performs minor cleanup in the protocol definitions. Feedback focuses on improving the test suite by removing unnecessary dependencies on numpy and pytest, replacing numpy array usage with standard Python lists for better portability, and avoiding brittle assertions based on internal string representations.
I am having trouble creating individual review comments. Click here to see my feedback.
tests/test_schema_chunking.py (5-6)
The numpy and pytest imports are unnecessary in this file. pytest is not explicitly used, and numpy introduces a heavy dependency that contradicts the 'relaxed compatibility' goal mentioned in src/cocoindex_code/schema.py. Removing these will make the tests more portable and faster to collect in environments where numpy is not installed.
tests/test_schema_chunking.py (24)
To avoid a hard dependency on numpy in the test suite, consider using a standard Python list for the embedding. This is consistent with the Any type hint in the CodeChunk dataclass and the project's goal of maintaining compatibility without requiring numpy as a mandatory dependency.
embedding=[0.0] * 384,
tests/test_schema_chunking.py (81)
Asserting against the string representation of CHUNKER_REGISTRY is brittle as it depends on the internal implementation of ContextKey.__str__. A more robust test would verify the object's type or other public attributes directly.
Housekeeping PR to improve test coverage for untested modules.
New tests for:
schema.py: CodeChunk, QueryResult dataclasseschunking.py: public API exports (Chunk, TextPosition, ChunkerFn, CHUNKER_REGISTRY)@gemini-code-assist review