Skip to content

Add opt-in bounded multiline regex find over indexed source #5399

Description

@Widthdom

Request and impact

Add an explicit, bounded multiline/window regex mode to indexed find and its existing MCP counterpart. This is audit candidate D10, suitable for one feature PR covering shared matching, output/continuation contracts, CLI/MCP exposure and tests.

During source auditing, line-local regex cannot directly express a split process/API call and its arguments, a multiline catch body, or an adjacent guard/body sequence. Increasing regex flags alone cannot solve this because the matcher receives one line at a time.

Reproduction of the current limitation

Run from the repository root with the binary built from 38519c2d301ea2e3217d687dd08fd07144936e0a and a healthy index:

audit_cdidx() {
  dotnet ./src/CodeIndex/bin/Debug/net8.0/cdidx.dll "$@" --db .cdidx/codeindex.db
}

audit_cdidx find 'using System\.Globalization;\nusing System\.Text\.Json;' --regex --path src/CodeIndex/Cli/QueryCommandRunner.Find.cs --count --json
# exit 0, count:0, authoritative_count:true

audit_cdidx find 'using System\.Globalization;' --regex --path src/CodeIndex/Cli/QueryCommandRunner.Find.cs --count --json
# exit 0, count:1, authoritative_count:true

The two using statements are adjacent at lines 1–2 of the audited source. (?s) cannot bridge two separate matcher invocations. The zero result is correct under the current line-local contract; this report does not claim ordinary regex support regressed.

Observed with repository-built cdidx 1.50.3, Debug net8.0, macOS arm64. Build and root/workspace freshness checks passed, with complete index/reference graph. Discovery used focused reproductions, not the full test suite.

Source starting points

One-PR implementation instructions and cautions

  1. Design and document an opt-in multiline/window option with conservative defaults and hard maximum bytes, source lines, regex time, and retained memory per window/query. Keep current literal and line-regex behavior unchanged. A bounded window is sufficient; do not load arbitrary full files into memory or introduce an unbounded recursive matcher.
  2. Define the public semantics before implementing: supported maximum match span, newline handling, ^/$ versus dot-all flags, non-overlapping/zero-width matches, match ordering and total/count meaning. Distinguish RegexOptions.Multiline (anchors) from passing multiline source to the matcher.
  3. Build windows from the indexed snapshot, preserving path and exact original start/end line and UTF-16 columns. Handle overlapping stored chunks and adjacent windows without duplicated matches or missed matches inside the advertised span limit. Never join different files. Bound snippet output separately from match span.
  4. Define one stable owner/start position per match and bind continuation to the query, multiline mode, window/span limits, filters, ordering and index generation. Resume with enough bounded context to recover crossing matches exactly once. Reject stale/incompatible cursors rather than guessing; do not reuse the current line cursor without reviewing the changed resume boundary.
  5. Preserve truthful completeness. If an execution cap prevents coverage promised by the selected mode, emit existing partial/non-authoritative metadata, a stable reason and safe recovery guidance. A documented maximum-span search is a limited search contract; do not imply it proves absence of arbitrary-length matches. Invalid regex, regex timeout and cancellation retain their established errors/exits.
  6. Expose consistent CLI and MCP options and schemas through the shared pipeline. Decide and test interaction with --all, count, path/language selection, focus/context options, output byte budgets and semantic origin filters. Unsupported combinations should fail early with structured guidance; do not silently drop filters or downgrade to line matching. Cross-origin matches require an explicit documented policy.
  7. Document the existing line-local limitation immediately in the affected help/tool descriptions as part of this PR, with a concrete multiline example. A heuristic warning may supplement documentation, but should not flag every escaped \n as an error or substitute for the requested feature.

Acceptance and validation

  • In the new explicit mode, the adjacent-using reproducer returns one match spanning source lines 1–2 with correct start/end coordinates; the existing mode retains its behavior.
  • Fixtures cover LF/CRLF, Unicode including astral characters, chunk/window boundaries, a match exactly at the supported span bound, matches beyond that bound, adjacent files, duplicate overlapping chunks, anchors, dot-all and zero-width matches.
  • Paginated/count scans match the bounded full scan without skipped/duplicated matches, including a match crossing the resume boundary. Query/window/filter/generation changes invalidate incompatible cursors.
  • Large lines, huge files, catastrophic-backtracking patterns, cancellation and tight output budgets are bounded and report their actual completion/error state. JSON/NDJSON/compact or envelope output must not lose required span or authority metadata; preserve existing format errors for unsupported semantic-filter combinations.
  • CLI and MCP agree on options, match spans, counts and incomplete-result metadata. Keep current literal and single-line-regex tests green.
  • Follow AGENT_GUIDE.md, SELF_IMPROVEMENT.md, TESTING_GUIDE.md, and the issue-fix/precommit workflows. Run focused find/pagination/MCP/output-budget tests on supported net8.0/net9.0 targets and required repository checks. Update matching help/user/developer documentation and add an English/Japanese changelog.d/unreleased/ fragment tied to this issue. Keep runtime dependencies within repository policy (Microsoft.Data.Sqlite only).

Previous issue and relationship

Closed #1930 requested clearer literal/regex behavior and mentioned anchors/multiline support. Opt-in regex exists now, but its execution remains line-local. This is a follow-up for the remaining multiline capability, not a recurrence of the original absence of ordinary regex support and not a bisected regression.

Submission priority

P3 — posting order 9/9 in the 2026-09-21 dogfood audit. Original candidate: D10. This issue is scoped for one independent implementation PR; closing another issue from this batch is not a prerequisite.

Activity

  1. Widthdom commented on Sep 22, 2026

    @Widthdom
    OwnerAuthor

    Work started on branch fix-issue5399, based on the freshly fetched origin/main. I will implement the opt-in bounded multiline indexed-find contract, matching CLI/MCP behavior, regression tests, bilingual documentation and changelog fragment, and mandatory Codex adversarial review.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions