Skip to content

grep: asking for context creates matches — 'Found N matches' counts context lines that end with a colon #1805

Description

@how2how2how2-arch

What happens

The grep tool's summary reports a match count that is not the number of matches, whenever context lines are requested. The same search, same file:

$ pattern=TARGET, context_before=0, context_after=0
Found 1 matches for 'TARGET' in ... (searched 1 files):

$ pattern=TARGET, context_before=1, context_after=2   (same one matching line)
Found 4 matches for 'TARGET' in ... (searched 1 files):

The three extra "matches" are the context lines — first:, second:, third: — which happen to end with a colon.

Why

actual_matches is not counted where the matches are found; it is re-derived from the rendered lines:

actual_matches = sum(1 for r in results if r.endswith(":"))

A match block is rendered as a header f"{rel}:{i + 1}:" followed by " " + marker + text context lines. The predicate "ends with a colon" is meant to select the headers, but a context line is emitted with its own text intact — so any line whose text ends in : (a YAML block key, a C++/Pascal public: label, a Markdown Term:) is counted as a match too. The count therefore rises with context_before / context_after alone: asking for context creates matches that do not exist.

Why it costs something

grep is the tool an agent reaches for to answer "how many places does this happen?", and the number is what it reasons from. The inflation is systematic rather than random: context windows are exactly where labelled lines live (YAML, config, code labels, prose lists), so the more structure a file has, the more wrong the count — and the error grows with the amount of context requested, i.e. it is worst in the mode used for careful reading.

It also interacts with the cap: max_results is described as "Maximum matches to return", and a caller who sees a large count near that cap has no way to tell how many blocks were real.

A second number in the same message that is not measured

The truncation note claims a block count the code never computes:

if len(results) > max_results * 3:
    results = results[: max_results * 3]          # 3×max_results LINES
    results.append(f"\n... [output truncated at ~{max_results} match blocks]")

The cut is at lines; the note names match blocks — and with context, 3 × max_results lines is fewer than max_results blocks (a block is 1 header + the context window), so the note overstates how much was shown. The cut can also fall inside a block, leaving context lines whose header is gone, i.e. output that reads as belonging to no match.

Acceptance

  • Found N matches equals the number of matching lines, whatever context_before / context_after request.
  • The truncation note names two numbers the code actually measures — how many blocks are shown and how many there are — and the printed output ends at a block boundary, so no context line is left without its header.
  • No change to the block layout a reader already sees (header, then " >" for the matching line and " " for context).

Activity

  1. how2how2how2-arch commented on Oct 1, 2026

    @how2how2how2-arch
    CollaboratorAuthor

    Handled by #1806

  2. how2how2how2-arch commented on Oct 1, 2026

    @how2how2how2-arch
    CollaboratorAuthor

    Scope added to #1806 by the cycle that pushed its second commit (cyc20261002-072356): the count is also a floor when the search stops at its result budget, and the summary now says so. The acceptance list above is unchanged; this line records the added item and the measurement behind it (a file of 4000 matching lines answered Found 11 matches ... (searched 1 files) with no indication the search had stopped).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions