feat: Add comprehensive feature upgrades - #8
Conversation
## Quick Wins - Structured JSON output with Pydantic models - Response caching with configurable TTL - Mobile device emulation profiles ## Performance and Speed - Parallel vision analysis with asyncio.gather - Video frame sampling (key frames only) - Browser connection pooling - Streaming responses support ## Quality and Accuracy - Multi-model consensus (Gemini + Ollama) - Visual diff detection (pixel comparison) - Animation FPS analysis (jank detection) ## Hallucination Prevention - Grounding with screenshots evidence - Self-verification prompts - Bounding box annotations - Multi-pass analysis ## New Capabilities - PDF/HTML report generation - Slack/Discord webhook notifications - Baseline comparison for regression detection - Performance metrics (Core Web Vitals: LCP, FID, CLS)
|
Caution Review failedFailed to post review comments Summary by CodeRabbitRelease NotesNew Features
Tests
WalkthroughThis pull request expands the animawatch platform with extensive vision analysis infrastructure including baseline regression tracking, multi-model consensus analysis, device emulation profiles, performance metrics extraction, structured analysis models, intelligent caching, visual diffing, video analysis (FPS and frame extraction), verification workflows, and comprehensive reporting and notification systems. Additionally, the browser recorder gains pooling capabilities and device emulation support. Changes
Sequence DiagramssequenceDiagram
participant Apprentice as Apprentice (Client)
participant Consensus as Consensus Engine
participant Gemini as Gemini Provider
participant Ollama as Ollama Provider
participant Cache as Analysis Cache
Apprentice->>Consensus: analyze_with_consensus(image, prompt)
Consensus->>Cache: get(gemini_key)
Cache-->>Consensus: None
Consensus->>Gemini: analyze_image(image, prompt, structured=true)
Gemini->>Cache: set(gemini_key, result)
Cache-->>Gemini: cached
Gemini-->>Consensus: AnalysisResult (Gemini)
Consensus->>Cache: get(ollama_key)
Cache-->>Consensus: None
Consensus->>Ollama: analyze_image(image, prompt, structured=true)
Ollama->>Cache: set(ollama_key, result)
Cache-->>Ollama: cached
Ollama-->>Consensus: AnalysisResult (Ollama)
Consensus->>Consensus: _findings_similar(gemini_findings, ollama_findings)
Consensus->>Consensus: Merge agreed findings
Consensus-->>Apprentice: ConsensusResult (consensus_score, merged_findings)
sequenceDiagram
participant Apprentice as Apprentice (Client)
participant Store as BaselineStore
participant Disk as Disk Storage
participant Grounding as Grounding Engine
participant Vision as Vision Provider
Apprentice->>Store: save_baseline(name, url, result, screenshot)
Store->>Store: _generate_id()
Store->>Store: _hash_screenshot(screenshot_path)
Store->>Disk: Write baseline.json
Disk-->>Store: Saved
Store-->>Apprentice: Baseline (with id)
Apprentice->>Store: load_baseline(baseline_id)
Store->>Disk: Read baseline.json
Disk-->>Store: baseline data
Store-->>Apprentice: Baseline object
Apprentice->>Grounding: multi_pass_analysis(image, vision_provider, prompt, passes=2)
Grounding->>Vision: analyze_image(image, grounded_prompt, structured=true)
Vision-->>Grounding: AnalysisResult
Grounding->>Grounding: Convert to GroundedFindings (pass 1)
Grounding->>Vision: analyze_image(image, verification_prompt, structured=true)
Vision-->>Grounding: AnalysisResult (verification)
Grounding->>Grounding: apply_verification_result (pass 2)
Grounding-->>Apprentice: list[GroundedFinding]
sequenceDiagram
participant Apprentice as Apprentice (Client)
participant Browser as BrowserRecorder
participant Pool as Context Pool
participant Device as Device Profile
participant Playwright as Playwright
Apprentice->>Browser: __init__(pool_size=3)
Browser->>Pool: Initialize empty pool
Apprentice->>Browser: take_screenshot(url, device="iPhone 15 Pro", use_pool=true)
Browser->>Device: _resolve_device("iPhone 15 Pro")
Device-->>Browser: DeviceProfile
Browser->>Pool: pooled_context(device)
Pool->>Playwright: Create context with device emulation
Playwright-->>Pool: BrowserContext
Browser->>Playwright: Take screenshot
Playwright-->>Browser: Screenshot path
Pool->>Pool: Reuse context in pool
Browser-->>Apprentice: Screenshot path
Estimated code review effort🎯 4 (Complex) | ⏱️ ~60 minutes Possibly related PRs
Poem
🚥 Pre-merge checks | ✅ 2 | ❌ 1❌ Failed checks (1 inconclusive)
✅ Passed checks (2 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing touches
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
|
||
|
|
||
| # JSON schema for prompting vision models to return structured output | ||
| STRUCTURED_OUTPUT_SCHEMA: dict[str, Any] = AnalysisResult.model_json_schema() |
Check notice
Code scanning / CodeQL
Unused global variable Note
Show autofix suggestion
Hide autofix suggestion
Copilot Autofix
AI 8 months ago
General fix: For a global variable that is intentionally exported for external use but unused within the defining module, either (a) add it to __all__ to mark it as a public symbol, or (b) rename it to an “unused” style name if it truly is unused and internal, or (c) delete it if it is genuinely dead code. Here we want to keep existing functionality (exporting the schema), so we should mark it as public.
Best concrete fix: In src/animawatch/models.py, define a module-level __all__ list that includes all the public models and constants, in particular STRUCTURED_OUTPUT_SCHEMA. That way CodeQL recognizes the variable as “explicitly made public by inclusion in the all list” and stops flagging it. We do not need to touch the existing STRUCTURED_OUTPUT_SCHEMA definition line or any imports.
Where and how to change: Add a __all__ = [...] definition near the top of models.py, after the imports and before the class definitions, listing at least:
"Severity""IssueCategory""BoundingBox""Finding""AnalysisMetadata""AnalysisResult""STRUCTURED_OUTPUT_SCHEMA"
No new imports or other helper methods are needed.
| @@ -9,7 +9,17 @@ | ||
|
|
||
| from pydantic import BaseModel, Field | ||
|
|
||
| __all__ = [ | ||
| "Severity", | ||
| "IssueCategory", | ||
| "BoundingBox", | ||
| "Finding", | ||
| "AnalysisMetadata", | ||
| "AnalysisResult", | ||
| "STRUCTURED_OUTPUT_SCHEMA", | ||
| ] | ||
|
|
||
|
|
||
| class Severity(str, Enum): | ||
| """Issue severity levels.""" | ||
|
|
Summary
This PR adds comprehensive feature upgrades to AnimaWatch covering performance, quality, accuracy, hallucination prevention, and new capabilities.
Changes
Quick Wins - Low Effort High Impact
models.py): Force JSON responses for reliable parsing with Pydantic modelscache.py): Cache identical analysis requests with configurable TTLdevices.py): Test on various viewport sizes and device profilesPerformance and Speed
vision.py): Analyze multiple screenshots concurrently with asyncio.gatherframes.py): Skip redundant frames, analyze key frames onlybrowser.py): Reuse browser contexts across recordingsvision.py): Stream analysis results as they are generatedQuality and Accuracy
consensus.py): Run Gemini + Ollama, compare and merge resultsdiff.py): Compare before/after screenshots pixel-by-pixelfps.py): Detect frame drops and jank quantitativelyHallucination Prevention
grounding.py): Always include visual evidence with claimsNew Capabilities
reports.py): Export analysis as PDF with screenshotsnotifications.py): Alert on critical issues via webhooksbaseline.py): Compare against known-good recordings for regression detectionmetrics.py): Extract Core Web Vitals (LCP, FID, CLS) from recordingsTesting
New Files
models.pycache.pydevices.pyframes.pyconsensus.pydiff.pyfps.pygrounding.pyreports.pynotifications.pybaseline.pymetrics.pyPull Request opened by Augment Code with guidance from the PR author