Measured 2026-10-06 (cyc20261006-214703) on master 38b85268. The grep tool counts a
file as searched before it decides whether it can read it, so a tree whose only copies
of the pattern are unreadable files answers "no matches" and names those files among the
ones it searched:
$ cat > /tmp/probe/big.txt # 524312 bytes, over MAX_FILE_SIZE (524288); holds NEEDLE
$ printf '\xff\xfeNEEDLE\x00' > /tmp/probe/binary.bin # not valid UTF-8; holds NEEDLE
$ echo 'hello world' > /tmp/probe/small.txt; echo 'nothing here' > /tmp/probe/other.txt
$ grep(pattern='NEEDLE', path='/tmp/probe')
No matches for 'NEEDLE' in /tmp/probe (searched 4 files)
files_searched is incremented at the top of the loop and the size / decode guards
continue below it (emrg/tools/grep_tool.py, the for filepath in files: block), so
2 of the 4 files it says it searched were never read — and both of them contain
NEEDLE. The same count is printed on the hit path
(Found N matches ... (searched M files)) and on the floor path.
It is the class this repository keeps fixing, in a third home
grep's own two earlier defects were the same claim in other words: the match count
included context lines that were not matches (#1805), and the count produced by a result
budget was printed as a total (#1805, second cause). This is the third count in the same
tool with the same shape — a number whose subject is not what it says. "Searched" is a
claim about work done; a file skipped for being 24 bytes over the cap, or for not being
decodable text, did not get searched, and the reader's next move ("the string is not in
this tree") is drawn from exactly that word.
The tool description does declare the skipping ("Skips binary files, hidden dirs, and
files over 512KB"), so the declaration half exists — what is missing is the reading.
The sibling tool was just fixed the same way
glob had the same defect in its own skip policy (issue #1874 / PR #1875): a pattern
whose matches were all skipped was answered No files matched, with no sign that the skip
had answered instead of the tree. The shape of the fix is the one to reuse: the count
names its subject, and what was not measured is stated beside it, never folded into it.
Acceptance
searched N files counts only files the loop really read — a file skipped for size, or
not decodable, is not in N;
- the skipped files are named beside it, with the two reasons kept apart (their remedies
differ: raise the size cap, versus a search that can read bytes), on the no-match line,
the hit line and the floor line;
- a search that skipped nothing prints no skip clause — the clause is a measurement, not
boilerplate — and both directions have a test that fails when its half is removed.
Measured 2026-10-06 (
cyc20261006-214703) on master38b85268. Thegreptool counts afile as searched before it decides whether it can read it, so a tree whose only copies
of the pattern are unreadable files answers "no matches" and names those files among the
ones it searched:
files_searchedis incremented at the top of the loop and the size / decode guardscontinuebelow it (emrg/tools/grep_tool.py, thefor filepath in files:block), so2 of the 4 files it says it searched were never read — and both of them contain
NEEDLE. The same count is printed on the hit path(
Found N matches ... (searched M files)) and on the floor path.It is the class this repository keeps fixing, in a third home
grep's own two earlier defects were the same claim in other words: the match countincluded context lines that were not matches (#1805), and the count produced by a result
budget was printed as a total (#1805, second cause). This is the third count in the same
tool with the same shape — a number whose subject is not what it says. "Searched" is a
claim about work done; a file skipped for being 24 bytes over the cap, or for not being
decodable text, did not get searched, and the reader's next move ("the string is not in
this tree") is drawn from exactly that word.
The tool description does declare the skipping ("Skips binary files, hidden dirs, and
files over 512KB"), so the declaration half exists — what is missing is the reading.
The sibling tool was just fixed the same way
globhad the same defect in its own skip policy (issue #1874 / PR #1875): a patternwhose matches were all skipped was answered
No files matched, with no sign that the skiphad answered instead of the tree. The shape of the fix is the one to reuse: the count
names its subject, and what was not measured is stated beside it, never folded into it.
Acceptance
searched N filescounts only files the loop really read — a file skipped for size, ornot decodable, is not in
N;differ: raise the size cap, versus a search that can read bytes), on the no-match line,
the hit line and the floor line;
boilerplate — and both directions have a test that fails when its half is removed.