feat(launchpad): ProjectIndexer producing Symbol records for buzz-core - #213
Conversation
…pter Implements the first two steps of #206's plan (launchpad/plans/2026-08-18-issue-206-project-indexer.md): STEP 1: the Symbol record type (symbol.py), matching launchpad/Research/project-intelligence-layer-design.md's schema exactly. Verified: python3 -m unittest test_symbol (2 tests, both pass). STEP 2 (RUNS HERE): an adapter (indexer.py) that queries RepoQL's own Functions view via the rql CLI for one crate (buzz-core) and maps rows into Symbol records -- kind, qualified_name, defined_at, signature, and a best-effort calls[] scan of the symbol's own source lines. Chose to lean on RepoQL's existing structural index rather than write a second AST parser, since it already exposes qualified names, signatures, declaring types, and line ranges for this exact repo. Location: launchpad/project-intelligence/ (Python), not a new Rust crate in crates/ -- matches this fork's own precedent (launchpad/review-agent/ is substantial Python tooling living outside the Cargo workspace) and its own stated boundary ("we operate Buzz, we do not develop Buzz" -- AGENTS.md #1). Registering a new crate into the upstream-owned root Cargo.toml for cohort-authored agent tooling would have been exactly the kind of "developing Buzz" this fork's own rules say isn't what this repo does. This resolves one of the plan's OPEN items (where this code lives) -- surfaced here rather than decided silently, per the plan-issue skill's own rule. Verified end to end against the real repo: indexing buzz-core produces 453 symbols; the worked-example symbol (is_shared_gated_kind) prints signature "pub fn is_shared_gated_kind(kind: u32) -> bool" and calls: "contains" -- checked directly against RepoQL's own read() of that symbol, whose body is `SHARED_GATED_KINDS.contains(&kind)`. Exact match. Steps 3-8 (called_by inverse index, git_ownership, tests[], config_dependencies[], documentation_links[], final CLI) remain. Signed-off-by: Serina Mcfall <serina.mcfall@gmail.com>
Materializes called_by[] once over the already-indexed symbol set (with_called_by()), matching each symbol's best-effort calls[] entries against the set's own qualified names -- not recomputed per query, per the design doc's requirement. Verified against the worked-example symbol (is_shared_gated_kind): indexer reports called_by = [tests::shared_gated_kinds_membership, is_unshared_gated_event]. Cross-checked by hand: kind.rs:234 calls it inside is_unshared_gated_event, and kind.rs:1076-1083 (test shared_gated_kinds_membership) calls it four times via assert!. Both real, both found. Signed-off-by: Serina Mcfall <serina.mcfall@gmail.com>
…story/blame enrich_git_ownership() shells out to `rql read <uri> => history` and `=> blame` for one symbol and populates GitOwnership(primary_authors, history) -- blame-weighted authors, one summary line per commit. Applied to the worked-example symbol only in the CLI, not eagerly over all 453 indexed symbols: each rql read call costs roughly a second, and batching that for a whole crate is a real performance concern this task's scope (prove the schema once) leaves for later. Verified against the worked-example symbol (is_shared_gated_kind): matches RepoQL's own history/blame output exactly, and independently cross-checked against raw `git log -L 219,221:crates/buzz-core/src/ kind.rs` -- same two commits (114d40d, ab3af82), same order, same messages. Signed-off-by: Serina Mcfall <serina.mcfall@gmail.com>
…dencies[]
STEP 5: tests[] derived as the subset of called_by[] whose qualified
name matches Rust's tests:: convention (or test_/_test), not a second
search pass -- the calling relationship STEP 3 already found is the
same evidence. Verified: is_shared_gated_kind's tests[] correctly shows
tests::shared_gated_kinds_membership.
STEP 6: config_dependencies[] -- a best-effort scan for
env::var(_os)("LITERAL") call sites in the symbol's own body. Refactored
_best_effort_calls to take an already-read body string
(_read_body helper) so calls[] and config_dependencies[] share one file
read instead of two.
buzz-core has no env-var reads at all (grep confirmed), so this step
was verified against buzz-relay instead: service_resource correctly
shows config_dependencies = (OTEL_SERVICE_NAME,) and try_init_tracer
shows (OTEL_EXPORTER_OTLP_ENDPOINT,), both matching real call sites at
crates/buzz-relay/src/telemetry.rs:208 and :239. main() correctly
surfaces 13 real env vars it reads.
Signed-off-by: Serina Mcfall <serina.mcfall@gmail.com>
… full pipeline
STEP 7: documentation_links[] -- word-boundary name match against this
repo's root-level and launchpad/ markdown, read once per crate rather
than per symbol. buzz-core's kind.rs functions have no doc mentions
(expected), so verified against a different real symbol in the same
crate: is_private_ip correctly shows documentation_links =
(ARCHITECTURE.md,), matching real mentions at ARCHITECTURE.md:355,740.
STEP 8: build_index() composes steps 2/3/5/6/7 into one pipeline;
enrich_git_ownership() (STEP 4) stays applied selectively rather than
folded in, per its own docstring on why. _print_symbol() now prints
every field explicitly -- "(none found)" everywhere an earlier step's
placeholder text ("not yet populated") used to be, since those steps
are now real.
All 8 steps of #206's plan
(launchpad/plans/2026-08-18-issue-206-project-indexer.md) are complete.
Full run against buzz-core (453 symbols) prints the worked-example
symbol (is_shared_gated_kind) with every field populated or explicitly
empty, cross-checked by hand against RepoQL and raw git at each step.
Signed-off-by: Serina Mcfall <serina.mcfall@gmail.com>
Codex's independent review of PR #213 confirmed 4 of 5 findings; one (P1, "enrich every symbol with git ownership") was checked against issue #206's own Definition of done ("checked by hand for at least one symbol") and refuted -- not a defect against the actual scoped task, just under-documented. Added an explicit note to build_index()'s docstring so it can't be misread that way again. Confirmed and fixed: - calls[] held bare short names while called_by[] held qualified names -- an internal representation inconsistency for what the schema specifies as symbol references either way. with_called_by() now resolves calls[] to qualified names too, using the same by_name index it already builds for the inverse; an unresolved call (a std-lib method, another crate) keeps its bare name since there's nothing more precise to offer. - GitOwnership.history was a tuple of pre-formatted strings; a consumer wanting just the date had to parse them back apart. Added CommitSummary(hash, date, author, message) and store that instead. - with_documentation_links() only matched a symbol's short name; the design doc's own step 7 wants a match by qualified name OR file. Added file-path matching (verified: is_shared_gated_kind now correctly shows ARCHITECTURE.md, which discusses crates/buzz-core/ src/kind.rs by file, not by this specific function's name). Section-level (#anchor) linking is explicitly left out -- it needs heading-structure parsing, a bigger lift than this task's own done-when asked for. - _print_symbol() never printed symbol_id or kind, so its output couldn't identify the graph node or distinguish functions from methods. Both now print. Also added test_indexer.py: 7 real unit tests for with_called_by() and with_tests()'s pure logic (qualified-name resolution, short-name collision handling -- documented as an accepted limitation, not silently fixed -- deduplication, self-calls, test classification), since test_symbol.py alone (Codex's other correct observation) only covers the dataclass, not the indexer's actual behavior. Verified: all 9 unit tests pass; live re-run against buzz-core (453 symbols) after a RepoQL host lock conflict was resolved (a stale process from a prior session was holding the DuckDB lock -- killed it, then `host restart` to clear the MCP bridge's cached connection to it). Signed-off-by: Serina Mcfall <serina.mcfall@gmail.com>
|
Codex's independent review confirmed 4 of 5 findings — fixed in the follow-up commit:
Also hit and resolved an unrelated RepoQL host lock conflict along the way (a stale process from a prior session holding the DuckDB lock) — noted in the fix commit's message, not something this PR's diff touches. |
benmitchell11
left a comment
There was a problem hiding this comment.
Reviewed the actual code, not just the PR body's claims. Both subprocess.run calls (run_rql_query, _rql_read_json) use list-form arguments with no shell=True, so the security claim ('shells out to rql... no trust-boundary changes') holds up — symbol names and queries flowing into those calls can't reach a shell regardless of their content. re.escape() is used correctly when building regex patterns from symbol names in the documentation-link matcher, which is an easy thing to get wrong and didn't get gotten wrong here.
The per-field verification is real, not asserted — each enrichment (calls/signature against RepoQL's own read(), called_by against manual grep, git_ownership against raw git log -L, config_dependencies and documentation_links against separately-chosen real symbols since buzz-core doesn't exercise those fields) was cross-checked against independent ground truth, not just against the tool's own output. Living outside the Cargo workspace (per AGENTS.md #1, 'we operate Buzz, we do not develop Buzz') is the right call and correctly precedented against launchpad/review-agent/.
tucktuck101
left a comment
There was a problem hiding this comment.
Review: approve
This fulfills #206's Definition of done point by point — Symbol record with every design-doc field, a real run against buzz-core (453 symbols), called_by[] materialized as a true inverse index at index time, git ownership hand-verified against git log -L, and a CLI printing the worked-example shape. The prior review round (calls/called_by representation, structured CommitSummary, file-path doc matching, identity fields, indexer tests) is genuinely addressed in the diff, not just claimed. Scope discipline is good: known imprecision is documented in docstrings rather than hidden, and the expensive per-symbol rql read is deliberately opt-in. Recommend merge; the notes below are follow-up-sized, none blocking.
Suggestions (non-blocking)
-
Sanitize
crate_namebefore SQL interpolation (indexer.pyindex_crate). It comes fromsys.argvand lands insidef"WHERE file LIKE '%crates/{crate_name}/%'"— a quote or%breaks/distorts the query. Are.fullmatch(r"[A-Za-z0-9_-]+", crate_name)guard is one line. (subprocesslist-form already prevents shell injection — good.) -
Unit-test the pure regex helpers.
test_indexer.py's header says the untested functions all shell out torql, but_best_effort_callsand_best_effort_config_depsare pure string functions — and they hold the most fragile logic in the file (_NOT_A_CALLfiltering, method calls like.contains(,env::var_os). A handful of hermetic cases would lock them down cheaply. -
_read_bodyerror handling: onlyOSErroris caught; a non-UTF-8 file raisesUnicodeDecodeErrorand aborts the whole crate index. Considerread_text(errors="ignore"), matching whatwith_documentation_linksalready does. Relatedly, anrqlfailure currently surfaces as a bareCalledProcessErrorwith stderr swallowed bycapture_output=True— worth re-raising with stderr attached. -
Performance nits for the whole-repo follow-up (not this task's scope, just marking the spots):
index_cratere-reads the same source file once per symbol (cache per file), andwith_documentation_linksis O(symbols × docs) full-text regex scans. -
Schema deviations to carry into #207: the design doc types
calls[]/called_by[]assymbol_ids, but this PR stores qualified names, with unresolved externals as bare short names in the same field (distinguishable only by missing::). Documented, and fine as a bridge — but ProjectGraph should normalize tosymbol_idso consumers can join without redoing the short-name mapping. Likewisedocumentation_links[]lacks the#sectionfragment andtemporal_stateis hardcoded"WORKING"; both are stated limitations, just don't let them silently ossify. -
primary_authorsnaming: it currently returns all blame authors sorted by line count. Either truncate (top-N or a threshold) or note in the field docstring that "primary" means "ordered by blame weight", so consumers don't treat a one-line contributor as an owner.
One CI observation: all required checks pass, but no workflow executes the new Python tests — they were run locally per the PR body. If launchpad/project-intelligence/ grows past this proof stage, a small CI job (like whatever launchpad/review-agent/ uses) would keep the suite honest.
…254) * docs(launchpad): plan issue #210 -- SemanticIndex and concept-retrieval pipeline Plans the fifth task under PRD #4's Project Intelligence Layer scope: a 10-step plan implementing ConceptEntry and a two-stage concept -> subsystem -> candidate symbol -> confirmed-reference pipeline, per the design doc's Data Model item 3 and Concept Retrieval reasoning rules. Confirmed before planning: #210 has a REAL stated dependency on #206 (unlike #208/#209) -- "Depends on #206 for the content to embed and summarize" -- so this branches from origin/task/207-project-graph (which carries both #206's and #207's already-merged-locally commits) rather than from `launchpad` directly, since #213/#214 haven't merged into launchpad yet. Key design decision recorded in the plan: "subsystem" is implemented as a real second ConceptEntry level scoped to file (the schema's own stated scope kinds are symbol_id | file | doc_section), not collapsed into a single flat symbol ranking -- so the pipeline's stated shape is actually built as two literal ranking stages, not simplified away. Mechanical checks clean via check-plan.sh. Signed-off-by: Serina Mcfall <serina.mcfall@gmail.com> * feat(launchpad): SemanticIndex STEP 1 -- ConceptEntry and store skeleton Adds ConceptEntry, matching the design doc's schema (scope, embedding, summary), and SemanticIndex, an in-process store keyed by scope. embedding is a tuple of (token, weight) pairs, not a dict, so the frozen dataclass stays genuinely immutable -- same reasoning as #209's MemoryEntry.evidence being a tuple rather than a mutable list. scope accepts symbol_id, file, or doc_section per the design doc's own schema -- this is deliberate groundwork for STEP 4/5's two-level pipeline (per-symbol and per-file "subsystem" entries), not unused generality. Verified: `python3 -m unittest test_semantic_index` -- 3 passed. Signed-off-by: Serina Mcfall <serina.mcfall@gmail.com> * feat(launchpad): SemanticIndex STEP 2 -- summarize_symbol Adds summarize_symbol(): a deterministic natural-language gloss built only from #206's already-extracted Symbol structural facts (qualified_name, kind, signature, calls, tests, config_dependencies, documentation_links) -- generated once, not guessed fresh per query, matching the design doc's own stated constraint. Empty fields are omitted rather than printed blank. Test fixture is a real Symbol from buzz-core, hand-constructed (not built via indexer.build_index(), which shells out to rql and is kept out of this committed hermetic suite, same reasoning as test_indexer.py/test_graph.py) from fields cross-checked directly against crates/buzz-core/src/kind.rs:219-221 and confirmed against ARCHITECTURE.md:142 (which references kind.rs by file path -- the documentation_links match is on file mention, not literal function name, matching #206's with_documentation_links() logic). Verified: `python3 -m unittest test_semantic_index` -- 6 passed. Signed-off-by: Serina Mcfall <serina.mcfall@gmail.com> * feat(launchpad): SemanticIndex STEP 3 -- tokenize, embed_text, cosine_similarity Adds tokenize() (word-boundary plus camelCase/snake_case splitting, so identifiers decompose into the same word tokens a natural-language concept query would use), embed_text() (a bag-of-words frequency vector -- a deliberate, documented lightweight stand-in for a trained ML embedding model, matching #210's own "out of scope: any embedding- model selection process beyond what's needed to demonstrate the pipeline once"), and cosine_similarity() between two such vectors, guarding the zero-vector case rather than dividing by zero. Verified: `python3 -m unittest test_semantic_index` -- 13 passed, including hand-computed (not real-symbol) cases per the plan's STEP 3 done-when: a vector against itself is 1.0, disjoint vocabularies are 0.0, and one fully hand-worked partial-overlap case (a a b / a c c -> 2/5) checked by hand in the test's own comment before being asserted. Signed-off-by: Serina Mcfall <serina.mcfall@gmail.com> * feat(launchpad): SemanticIndex STEP 4 -- from_symbols, two ConceptEntry levels (RUNS HERE) Adds SemanticIndex.from_symbols(), building two levels of ConceptEntry from real Symbol records: one per symbol, and one per file (aggregating that file's symbols' summaries) -- the design doc's own ConceptEntry schema names file as a valid scope kind alongside symbol_id, so this is the schema's own coarser "subsystem" level, not invented machinery. Self-caught bug from live verification (not from a test I wrote in advance): the first version keyed per-symbol entries by qualified_name, which raised ValueError on real buzz-core data -- multiple distinct symbols (e.g. several "build_event" functions in different modules) share one qualified_name, so it is not a safe unique key. Switched to symbol_id (a real, per-symbol-unique RepoQL URI from #206's index_crate()), which the design doc also explicitly names as a valid scope kind. Added qualified_name_for(), since #207's ProjectGraph addresses nodes by qualified_name, not symbol_id -- the pipeline's later confirmation step needs to translate between the two. Verified: `python3 -m unittest test_semantic_index` -- 17 passed, including a synthetic two-symbols-sharing-one-qualified_name regression test for the exact collision this fix addresses. Verified live against ALL 453 real buzz-core symbols (not just the hand-picked fixtures): building the full index raises nothing, and crates/buzz-core/src/kind.rs's file-level entry aggregates both is_shared_gated_kind and is_unshared_gated_event's real content. Signed-off-by: Serina Mcfall <serina.mcfall@gmail.com> * feat(launchpad): SemanticIndex STEP 5 -- two-stage search, embed_symbol identity weighting Adds SemanticIndex.search(): rank file-level "subsystem" entries first, then rank symbol-level entries scoped to the top file(s) -- concept -> candidate subsystem(s) -> candidate symbols, as two literal ranking stages, returned as SearchResult(subsystem, subsystem_score, candidate, candidate_score) sorted by candidate_score. Adds embed_symbol(), weighting a symbol's own identity (kind + qualified_name + signature) 2x over context mentions (calls/tests/ config/docs) in its embedding. Without this, a caller's summary absorbs its callees' name tokens too (since "calls X" contributes X's own identifier tokens), so a caller can outrank the callee it calls for a query about the callee's own behavior -- found empirically verifying this step's own worked example: is_unshared_gated_event (which calls is_shared_gated_kind) initially outranked is_shared_gated_kind itself once the query touched a token unique to the caller's own name ("event"). identity_weight=2.0 was checked empirically against the real worked example, not derived analytically, and is documented as such in embed_symbol()'s docstring. Retrofitted STEP 4's from_symbols() to use embed_symbol() instead of a plain embed_text(summary) for both the per-symbol and per-file embeddings. Verified: `python3 -m unittest test_semantic_index` -- 19 passed, including a hand-checked identity-weighting assertion (event's weight > kind's absorbed-context weight for is_unshared_gated_event) and the corrected two-stage search test (kind.rs wins as subsystem, is_shared_gated_kind wins as candidate, over an unrelated real symbol from crates/buzz-core/src/invite.rs). Verified live against the full real buzz-core index (453 symbols, via indexer.build_index): searching "which function decides if a kind is gated for shared visibility" ranks is_shared_gated_kind first (0.5706) by a clear margin over the next real result (tests::shared_gated_kinds_membership, 0.4300), with kind.rs correctly winning as the top subsystem. Signed-off-by: Serina Mcfall <serina.mcfall@gmail.com> * feat(launchpad): SemanticIndex STEP 6 -- confirm_via_graph Adds confirm_via_graph(): the pipeline's final confirmation step, calling directly into #207's ProjectGraph.edges_from() for tested_by/called_by edges on a candidate symbol -- real structural confirmation, not semantic similarity alone. Returns empty tuples (never hidden) when nothing confirms a candidate, so a caller can see an unconfirmed guess for what it is. Verified: `python3 -m unittest test_semantic_index` -- 21 passed. Cross-checked confirm_via_graph(graph, "is_shared_gated_kind") against #207's own already-proven demo output for the identical symbol (graph.py's __main__, STEP 6): tested_by -> tests::shared_gated_kinds_membership, called_by -> is_unshared_gated_event -- matches exactly. Signed-off-by: Serina Mcfall <serina.mcfall@gmail.com> * feat(launchpad): SemanticIndex STEP 7 -- find_it_for_me full pipeline Adds find_it_for_me(): concept -> subsystem -> candidate -> confirm, tied into one call. Translates the top candidate's scope (symbol_id) back to its qualified_name via SemanticIndex.qualified_name_for() before confirming through #207's ProjectGraph, since the two components address symbols differently. Returns an empty result (candidate/confirmation None) rather than crashing when the index has nothing to rank. Verified: `python3 -m unittest test_semantic_index` -- 23 passed. One test builds both a SemanticIndex and a ProjectGraph from the same three real symbols and confirms find_it_for_me() ties all of it together correctly in one call: is_shared_gated_kind as the resolved qualified_name, kind.rs as the subsystem, and real callers/tests matching STEP 6's own already-verified edges. Signed-off-by: Serina Mcfall <serina.mcfall@gmail.com> * feat(launchpad): SemanticIndex STEP 8 -- positive worked example end to end Adds WORKED_EXAMPLE_CONCEPT, a named module-level constant for STEP 8's own worked concept-search example against this repo's real code (not the design doc's fictional OnboardingMailer example): "which function decides if a kind is gated for shared visibility" -- resolves to is_shared_gated_kind via subsystem -> candidate -> confirmed-reference. Verified: `python3 -m unittest test_semantic_index` -- 25 passed, including an explicit check that the concept sentence contains no contiguous substring match of "is_shared_gated_kind" (nor its underscores-as-spaces form) -- proving this is genuine token/concept overlap, not an accidental literal substring hit, per #210's own Definition of done. Verified live against the FULL real buzz-core index (453 symbols, via indexer.build_index + graph.ProjectGraph.from_symbols): find_it_for_me resolves the exact same concept sentence to is_shared_gated_kind (candidate_score 0.5706, subsystem kind.rs at 0.4184), confirmed by 2 real called_by edges (tests::shared_gated_kinds_membership, is_unshared_gated_event) and 1 real tested_by edge. Signed-off-by: Serina Mcfall <serina.mcfall@gmail.com> * test(launchpad): SemanticIndex STEP 9 -- negative flow-tracing boundary case Demonstrates the documented boundary from #210's own side (#207's graph.py already showed the mirror: a vague description has no symbol_id for reachable() to start from). Poses the SAME 2-hop relationship #207's own reachable() demo already proves (tests::is_unshared_gated_event_author_always_allowed -> is_unshared_gated_event -> is_shared_gated_kind, all real symbols cross-checked against crates/buzz-core/src/kind.rs:997-1007) through THIS pipeline instead. No new production code -- STEPS 1-7 already implement everything this exercises; this step is the issue's own required negative demonstration, not new functionality. Verified: `python3 -m unittest test_semantic_index` -- 27 passed. Confirms structurally (PipelineResult/Confirmation have no hop or path field at all -- checked via dataclasses.fields(), not by reading the source) and behaviorally (confirm_via_graph() only ever returns direct edges; is_shared_gated_kind never appears in a one-hop confirmation for a symbol two hops away) that this pipeline cannot express or verify a multi-hop path, while reachable() answers the identical relationship exactly. Verified live against the FULL real buzz-core index (453 symbols): the same flow-tracing question resolves this pipeline to an unrelated weak match (tests::test_unspecified) with an EMPTY confirmation -- it cannot even find a sensible candidate for a flow-tracing question, let alone a verified 2-hop path -- while reachable() returns the exact real path (tests::is_unshared_gated_event_author_always_allowed -> is_unshared_gated_event -> is_shared_gated_kind). Signed-off-by: Serina Mcfall <serina.mcfall@gmail.com> * feat(launchpad): SemanticIndex STEP 10 -- final CLI, both worked examples Wires __main__ to run both worked examples end to end against the real buzz-core index, matching #206/#207/#208/#209's demo style: the positive concept -> subsystem -> candidate -> confirmation trace, and the negative flow-tracing contrast (this pipeline's result vs ProjectGraph.reachable()'s exact answer), side by side. Named the negative example's constants (NEGATIVE_EXAMPLE_FLOW_QUESTION, NEGATIVE_EXAMPLE_START_SYMBOL) alongside STEP 8's WORKED_EXAMPLE_CONCEPT, and updated STEP 9's test to reference them instead of duplicating the literal strings. Verified live: `python3 semantic_index.py` prints both traces with real results -- is_shared_gated_kind resolved and confirmed for the positive example, and tests::test_unspecified (an unrelated weak match, empty confirmation) for the negative example, contrasted against reachable()'s real verified 2-hop path. Verified: `python3 -m unittest test_semantic_index` -- 27 passed. This completes all 10 steps of #210's plan. Signed-off-by: Serina Mcfall <serina.mcfall@gmail.com> --------- Signed-off-by: Serina Mcfall <serina.mcfall@gmail.com>
Summary
Implements all 8 steps of #206's plan: a ProjectIndexer that maps RepoQL's own
Functionsviewand git history/blame into the
Symbolrecord schema fromlaunchpad/Research/project- intelligence-layer-design.md, for one crate (buzz-core), rather than writing a second parser.Related issue
Closes #206
Issue type
Task
Agent provenance
Objective
Build a
ProjectIndexerthat parses one target crate intoSymbolrecords matching the designdoc's schema, per #206.
Impacted components
launchpad/project-intelligence/symbol.py
launchpad/project-intelligence/indexer.py
launchpad/project-intelligence/test_symbol.py
Approach and rejected alternatives
Chosen: a thin Python adapter over RepoQL's existing structural index (
rql query/rql readCLI), rather than a from-scratch AST parser -- RepoQL already exposes qualified names,
signatures, declaring types, line ranges, and git history/blame for this exact repo.
Chosen:
launchpad/project-intelligence/(Python), not a new Rust crate registered into theroot
Cargo.toml. Rejected the Rust-crate option because this repo's ownAGENTS.md#1 states"we operate Buzz, we do not develop Buzz," and
launchpad/review-agent/is the establishedprecedent for substantial cohort-authored agent tooling living outside the Cargo workspace.
Registering a new crate for this would have been exactly the "developing Buzz" this fork's rules
say isn't what it does, and would touch an upstream-owned file (
Cargo.toml) for no reason tiedto a fork-specific problem.
Chosen:
calls[]/config_dependencies[]as best-effort regex scans of each symbol's own sourcelines, not a real parser -- explicitly scoped this way in the plan (LEFT OUT: full call-graph
precision belongs to #207/ProjectGraph).
Chosen:
git_ownershipapplied selectively (one symbol at a time viaenrich_git_ownership()),not eagerly for all 453 indexed symbols -- each
rql readcall costs roughly a second, andbatching/parallelizing that is a real performance concern out of this task's scope.
Verification
Command run:
Raw output:
Command run:
Raw output:
Every field independently cross-checked by hand at the step that introduced it (see commit
messages for each):
calls/signatureagainst RepoQL's ownread()of the same symbol;called_byagainst a manualgrepfor real call sites;git_ownershipagainst rawgit log -L 219,221:crates/buzz-core/src/kind.rsdirectly (not just RepoQL's own derived view);config_dependenciesagainst a different real symbol (service_resourceinbuzz-relay,verified against
crates/buzz-relay/src/telemetry.rs:208, sincebuzz-corehas no env-varreads);
documentation_linksagainst a third real symbol (is_private_ip, verified againstARCHITECTURE.md:355,740).Not verified
Did not run this against any crate other than
buzz-core(plan's own scope: prove the schemaand pipeline once, not cover the whole repo) or any language other than Rust. Did not measure
performance at the whole-repo scale --
enrich_git_ownership()'s one-call-per-symbol cost isnoted as a concern for later, not solved here. Did not verify behavior on a file the
rqlhosthasn't indexed yet, or on a crate with zero functions.
Security implications
N/A - a read-only tool over already-indexed structural data and git history; no code, config,
network, or trust-boundary changes. Shells out to
rql(already trusted, already on PATH inthis environment) and reads local files under the repo root only.
Escalations
The plan flagged three OPEN items; this PR resolves two of them by implementation choice
(stated above, not silently): where the code lives, and the RepoQL mechanism for finding
callers (a name-match text search, not a
related()-style relationship the plan wasn't sureexisted). The third -- which crate/language to target first -- was resolved by following the
plan's own recommendation (
buzz-core), not an independent decision.