Skip to content

feat: Add comprehensive feature upgrades - #8

Merged
rishitank merged 1 commit into
mainfrom
feature/comprehensive-upgrades
Feb 6, 2026
Merged

rishitank merged 1 commit into
mainfrom
feature/comprehensive-upgrades

Conversation

@rishitank

Copy link
Copy Markdown
Owner

Summary

This PR adds comprehensive feature upgrades to AnimaWatch covering performance, quality, accuracy, hallucination prevention, and new capabilities.

Changes

Quick Wins - Low Effort High Impact

  • Structured JSON Output Mode (models.py): Force JSON responses for reliable parsing with Pydantic models
  • Response Caching (cache.py): Cache identical analysis requests with configurable TTL
  • Mobile Device Emulation (devices.py): Test on various viewport sizes and device profiles

Performance and Speed

  • Parallel Vision Analysis (vision.py): Analyze multiple screenshots concurrently with asyncio.gather
  • Video Frame Sampling (frames.py): Skip redundant frames, analyze key frames only
  • Connection Pooling (browser.py): Reuse browser contexts across recordings
  • Streaming Responses (vision.py): Stream analysis results as they are generated

Quality and Accuracy

  • Multi-Model Consensus (consensus.py): Run Gemini + Ollama, compare and merge results
  • Visual Diff Detection (diff.py): Compare before/after screenshots pixel-by-pixel
  • Animation FPS Analysis (fps.py): Detect frame drops and jank quantitatively

Hallucination Prevention

  • Grounding with Screenshots (grounding.py): Always include visual evidence with claims
  • Self-Verification Prompts: Ask model to verify its own findings
  • Bounding Box Annotations: Have model specify exact coordinates of issues
  • Multi-Pass Analysis: First pass finds issues, second pass verifies

New Capabilities

  • PDF Report Generation (reports.py): Export analysis as PDF with screenshots
  • Slack/Discord Notifications (notifications.py): Alert on critical issues via webhooks
  • Baseline Comparison (baseline.py): Compare against known-good recordings for regression detection
  • Performance Metrics (metrics.py): Extract Core Web Vitals (LCP, FID, CLS) from recordings

Testing

  • All 50 existing tests pass
  • Ruff lint checks pass
  • Mypy type checks pass

New Files

File Description
models.py Pydantic models for structured output
cache.py Response caching with TTL
devices.py Mobile device emulation profiles
frames.py Video frame sampling
consensus.py Multi-model consensus
diff.py Visual diff detection
fps.py Animation FPS analysis
grounding.py Hallucination prevention
reports.py PDF/HTML report generation
notifications.py Slack/Discord webhooks
baseline.py Baseline comparison
metrics.py Core Web Vitals extraction

Pull Request opened by Augment Code with guidance from the PR author

## Quick Wins
- Structured JSON output with Pydantic models
- Response caching with configurable TTL
- Mobile device emulation profiles

## Performance and Speed
- Parallel vision analysis with asyncio.gather
- Video frame sampling (key frames only)
- Browser connection pooling
- Streaming responses support

## Quality and Accuracy
- Multi-model consensus (Gemini + Ollama)
- Visual diff detection (pixel comparison)
- Animation FPS analysis (jank detection)

## Hallucination Prevention
- Grounding with screenshots evidence
- Self-verification prompts
- Bounding box annotations
- Multi-pass analysis

## New Capabilities
- PDF/HTML report generation
- Slack/Discord webhook notifications
- Baseline comparison for regression detection
- Performance metrics (Core Web Vitals: LCP, FID, CLS)
@rishitank
rishitank enabled auto-merge (squash) February 6, 2026 11:14
@coderabbitai

coderabbitai Bot commented Feb 6, 2026 •

Copy link
Copy Markdown

Caution

Review failed

Failed to post review comments

Summary by CodeRabbit

Release Notes

New Features

  • Baseline regression tracking for monitoring visual changes across versions
  • Device emulation profiles supporting mobile, tablet, and desktop testing scenarios
  • Visual diffing and screenshot comparison with highlighted differences
  • Performance analysis including FPS consistency detection and jank measurement
  • Core Web Vitals and detailed performance metrics collection
  • HTML and PDF report generation with optional embedded screenshots
  • Slack and Discord webhook notifications with threshold-based alerting
  • Multi-model consensus analysis combining Gemini and Ollama for improved accuracy
  • Video frame extraction with automatic similarity filtering
  • Intelligent response caching to reduce API calls and improve performance
  • Grounding and verification workflow to reduce hallucinations in AI analysis

Tests

  • Enhanced test isolation with cache reset and temporary file handling

Walkthrough

This pull request expands the animawatch platform with extensive vision analysis infrastructure including baseline regression tracking, multi-model consensus analysis, device emulation profiles, performance metrics extraction, structured analysis models, intelligent caching, visual diffing, video analysis (FPS and frame extraction), verification workflows, and comprehensive reporting and notification systems. Additionally, the browser recorder gains pooling capabilities and device emulation support.

Changes

Cohort / File(s) Summary
Core Analysis Models
src/animawatch/models.py
Introduces Severity, IssueCategory enums and BoundingBox, Finding, AnalysisMetadata, AnalysisResult dataclasses for structured vision analysis output with validation, helper properties, and Markdown rendering.
Vision Provider Enhancements
src/animawatch/vision.py
Adds structured JSON output instruction, per-call result caching via shared AnalysisCache, support for structured flag in analyze_image/analyze_video, new analyze_images_parallel and analyze_image_streaming methods, and internal _parse_structured_response logic across GeminiProvider and OllamaProvider.
Multi-Model Consensus
src/animawatch/consensus.py
Implements analyze_with_consensus to run Gemini and Ollama analyses concurrently, merge findings by similarity, and produce ConsensusResult with agreed/model-specific findings and quantified consensus_score.
Analysis Caching
src/animawatch/cache.py
Provides thread-safe AnalysisCache with TTL, content-hash keys, async get/set/invalidate/clear methods, capacity-based eviction, and hit-rate statistics.
Baseline Regression Tracking
src/animawatch/baseline.py
Adds Baseline and BaselineComparison models; implements BaselineStore for disk-based CRUD operations with screenshot hashing and serialisation; provides compare_against_baseline function to compute regression deltas.
Device Emulation Profiles
src/animawatch/devices.py
Defines DeviceCategory enum, DeviceProfile dataclass with viewport/user-agent/touch properties, DEVICES mapping for iPhone/Pixel/iPad/Desktop variants, and utility functions get_device and list_devices.
Browser Enhancements
src/animawatch/browser.py
Adds context pooling (configurable pool_size), device emulation support (via _resolve_device), new pooled_context generator, and enhanced recording_context/take_screenshot to accept device and use_pool parameters with associated logging.
Grounding & Verification
src/animawatch/grounding.py
Introduces GroundedFinding extending Finding with verification metadata; provides multi_pass_analysis for structured grounding and optional verification passes using vision providers; includes prompt builders and bounding box parsing.
Performance Metrics
src/animawatch/metrics.py
Exposes CoreWebVitals, PerformanceMetrics, MetricsThresholds models; provides extract_performance_metrics, collect_metrics_during_interaction, and generate_metrics_report with LCP/FID/CLS/TTFB rating logic.
Video Frame Analysis
src/animawatch/fps.py, src/animawatch/frames.py
FPS module detects jank events via frame timing extraction with ffprobe; Frames module extracts frames via ffmpeg or imageio with content hashing and similarity filtering, plus cleanup utilities.
Visual Diffing
src/animawatch/diff.py
Implements compare_images to detect pixel-level differences with optional diff image rendering; provides compare_screenshots_batch for multi-image comparison with structured DiffRegion and VisualDiffResult output.
Report Generation
src/animawatch/reports.py
Provides ReportConfig, generate_html_report, save_html_report, and async generate_pdf_report functions with embedded screenshots, findings sorted by severity, and metadata sections.
Notifications
src/animawatch/notifications.py
Introduces NotificationConfig; implements send_notification with service-specific payloads for Slack (colour-coded blocks), Discord (embeds), and generic webhooks; adds notify_on_threshold for conditional triggering.
Server & Test Updates
src/animawatch/server.py, tests/test_vision.py
Server adds structured=False to vision calls with to_markdown conversion for backward compatibility; tests introduce autouse fixture for cache/circuit reset and use real temporary files for video hashing.

Sequence Diagrams

sequenceDiagram
    participant Apprentice as Apprentice (Client)
    participant Consensus as Consensus Engine
    participant Gemini as Gemini Provider
    participant Ollama as Ollama Provider
    participant Cache as Analysis Cache
    
    Apprentice->>Consensus: analyze_with_consensus(image, prompt)
    Consensus->>Cache: get(gemini_key)
    Cache-->>Consensus: None
    Consensus->>Gemini: analyze_image(image, prompt, structured=true)
    Gemini->>Cache: set(gemini_key, result)
    Cache-->>Gemini: cached
    Gemini-->>Consensus: AnalysisResult (Gemini)
    
    Consensus->>Cache: get(ollama_key)
    Cache-->>Consensus: None
    Consensus->>Ollama: analyze_image(image, prompt, structured=true)
    Ollama->>Cache: set(ollama_key, result)
    Cache-->>Ollama: cached
    Ollama-->>Consensus: AnalysisResult (Ollama)
    
    Consensus->>Consensus: _findings_similar(gemini_findings, ollama_findings)
    Consensus->>Consensus: Merge agreed findings
    Consensus-->>Apprentice: ConsensusResult (consensus_score, merged_findings)
Loading
sequenceDiagram
    participant Apprentice as Apprentice (Client)
    participant Store as BaselineStore
    participant Disk as Disk Storage
    participant Grounding as Grounding Engine
    participant Vision as Vision Provider
    
    Apprentice->>Store: save_baseline(name, url, result, screenshot)
    Store->>Store: _generate_id()
    Store->>Store: _hash_screenshot(screenshot_path)
    Store->>Disk: Write baseline.json
    Disk-->>Store: Saved
    Store-->>Apprentice: Baseline (with id)
    
    Apprentice->>Store: load_baseline(baseline_id)
    Store->>Disk: Read baseline.json
    Disk-->>Store: baseline data
    Store-->>Apprentice: Baseline object
    
    Apprentice->>Grounding: multi_pass_analysis(image, vision_provider, prompt, passes=2)
    Grounding->>Vision: analyze_image(image, grounded_prompt, structured=true)
    Vision-->>Grounding: AnalysisResult
    Grounding->>Grounding: Convert to GroundedFindings (pass 1)
    Grounding->>Vision: analyze_image(image, verification_prompt, structured=true)
    Vision-->>Grounding: AnalysisResult (verification)
    Grounding->>Grounding: apply_verification_result (pass 2)
    Grounding-->>Apprentice: list[GroundedFinding]
Loading
sequenceDiagram
    participant Apprentice as Apprentice (Client)
    participant Browser as BrowserRecorder
    participant Pool as Context Pool
    participant Device as Device Profile
    participant Playwright as Playwright
    
    Apprentice->>Browser: __init__(pool_size=3)
    Browser->>Pool: Initialize empty pool
    
    Apprentice->>Browser: take_screenshot(url, device="iPhone 15 Pro", use_pool=true)
    Browser->>Device: _resolve_device("iPhone 15 Pro")
    Device-->>Browser: DeviceProfile
    Browser->>Pool: pooled_context(device)
    Pool->>Playwright: Create context with device emulation
    Playwright-->>Pool: BrowserContext
    Browser->>Playwright: Take screenshot
    Playwright-->>Browser: Screenshot path
    Pool->>Pool: Reuse context in pool
    Browser-->>Apprentice: Screenshot path
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related PRs

Poem

Behold, Apprentice, the Dark Side assembles,
Baselines held firm whilst the rebel code trembles,
Consensus flows strong through the vision's keen sight,
Device emulation bends to your will's mighty might,
Your analyses cached—the Force shows its power bright! ⚡✨

🚥 Pre-merge checks | ✅ 2 | ❌ 1
❌ Failed checks (1 inconclusive)
Check name Status Explanation Resolution
Title check ❓ Inconclusive The title 'feat: Add comprehensive feature upgrades' is vague and generic, lacking specificity about which features or changes are most critical. Consider a more specific title that highlights the primary feature or category, such as 'feat: Add structured output, caching, and multi-model consensus' or 'feat: Add hallucination prevention and performance optimizations'.
✅ Passed checks (2 passed)
Check name Status Explanation
Description check ✅ Passed The description is highly detailed and directly related to the changeset, covering all 12 new modules and their purposes across multiple categories of improvements.
Docstring Coverage ✅ Passed Docstring coverage is 92.86% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch feature/comprehensive-upgrades

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

Comment thread src/animawatch/models.py


# JSON schema for prompting vision models to return structured output
STRUCTURED_OUTPUT_SCHEMA: dict[str, Any] = AnalysisResult.model_json_schema()

Check notice

Code scanning / CodeQL

Unused global variable Note

The global variable 'STRUCTURED_OUTPUT_SCHEMA' is not used.

Copilot Autofix

AI 8 months ago

General fix: For a global variable that is intentionally exported for external use but unused within the defining module, either (a) add it to __all__ to mark it as a public symbol, or (b) rename it to an “unused” style name if it truly is unused and internal, or (c) delete it if it is genuinely dead code. Here we want to keep existing functionality (exporting the schema), so we should mark it as public.

Best concrete fix: In src/animawatch/models.py, define a module-level __all__ list that includes all the public models and constants, in particular STRUCTURED_OUTPUT_SCHEMA. That way CodeQL recognizes the variable as “explicitly made public by inclusion in the all list” and stops flagging it. We do not need to touch the existing STRUCTURED_OUTPUT_SCHEMA definition line or any imports.

Where and how to change: Add a __all__ = [...] definition near the top of models.py, after the imports and before the class definitions, listing at least:

  • "Severity"
  • "IssueCategory"
  • "BoundingBox"
  • "Finding"
  • "AnalysisMetadata"
  • "AnalysisResult"
  • "STRUCTURED_OUTPUT_SCHEMA"

No new imports or other helper methods are needed.


Suggested changeset 1
src/animawatch/models.py

Autofix patch

Autofix patch
Run the following command in your local git repository to apply this patch
cat << 'EOF' | git apply
diff --git a/src/animawatch/models.py b/src/animawatch/models.py
--- a/src/animawatch/models.py
+++ b/src/animawatch/models.py
@@ -9,7 +9,17 @@
 
 from pydantic import BaseModel, Field
 
+__all__ = [
+    "Severity",
+    "IssueCategory",
+    "BoundingBox",
+    "Finding",
+    "AnalysisMetadata",
+    "AnalysisResult",
+    "STRUCTURED_OUTPUT_SCHEMA",
+]
 
+
 class Severity(str, Enum):
     """Issue severity levels."""
 
EOF
@@ -9,7 +9,17 @@

from pydantic import BaseModel, Field

__all__ = [
"Severity",
"IssueCategory",
"BoundingBox",
"Finding",
"AnalysisMetadata",
"AnalysisResult",
"STRUCTURED_OUTPUT_SCHEMA",
]


class Severity(str, Enum):
"""Issue severity levels."""

Copilot is powered by AI and may make mistakes. Always verify output.
@rishitank
rishitank merged commit e12ce55 into main Feb 6, 2026
9 checks passed
@rishitank
rishitank deleted the feature/comprehensive-upgrades branch February 6, 2026 11:21
@rishitank rishitank mentioned this pull request Feb 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants