Skip to content

Add vision-language model text query and whole-track image query - #1992

Open
mattdawkins wants to merge 2 commits into
mainfrom
dev/vlm-query
Open

mattdawkins wants to merge 2 commits into
mainfrom
dev/vlm-query

Conversation

@mattdawkins

Copy link
Copy Markdown
Member
  • Text Query dialog and new Text Query panel: pick SAM3 or a local Ollama vision model (e.g. qwen3-vl:8b)
    • Vision models: Find objects (boxes on this frame or all frames) and Ask (questions about the frame, a track crop, or sampled track frames; answers saved as track/detection attributes)
    • Missing models show how to install them (Add-Ons page for SAM3, Ollama download + ollama pull for vision models)
  • All-frames jobs, with a Track results across frames option: SAM3 tracking/no-tracking pipes, and utility_text_query_ollama_vlm_{tracking,no_tracking}.pipe
  • Image Query panel (was Video Search): new Search from selected track sends up to 6 frames of the track as separate positive exemplars

Needs the companion VIAME changes (interactive_vlm.py, ollama_vlm.py, the two ollama_vlm pipes, and formulate_query exemplars in query_service.py), not yet committed.

🤖 Generated with Claude Code

mattdawkins and others added 2 commits September 27, 2026 16:17
Ollama VLM backends (find objects, ask) in Text Query dialog and panel, with
install guidance; all-frames jobs with optional tracking; Image Query can
search with frames sampled along the selected track.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants