hotfix(citation): drop sampleSize===0 from isInsufficient predicate - #630
Merged
Conversation
Live audit on prod showed /api/citable marking 25 of 26 benches as status=insufficient while /api/stat for the same slugs returned status=live with real leader values. Root cause: the per-bench loader and the aggregator loader compute sampleSize differently; when the aggregator falls back to a draft placeholder due to a cold Prom hit, the bench appears with sampleSize=0 even though per-bench cache holds fresh data. The check b.sampleSize===0 was then mass-flagging these benches as insufficient in /api/citable, /api/llm-context, /api/mcp. The other checks already catch the genuine empty case (liveResults length and p50 finiteness). Dropping the sampleSize check restores the 12 to 14 healthy benches that surfaced pre-regression.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Live regression. /api/citable was returning status=insufficient for 25 of 26 benches while /api/stat returned live values for the same slugs. Root cause is the b.sampleSize===0 check added in #627 which fires whenever the aggregator loader falls back to a draft placeholder, even when the per-bench loader has real data. The other isInsufficient checks already handle the genuine empty case.