Skip to content

Add correlated V2 runtime latency profiles - #574

Open
ZenAlexa wants to merge 12 commits into
NVIDIA:mainfrom
ZenAlexa:contrib/538-api-user-input-latency-instrumentation
Open

ZenAlexa wants to merge 12 commits into
NVIDIA:mainfrom
ZenAlexa:contrib/538-api-user-input-latency-instrumentation

Conversation

@ZenAlexa

@ZenAlexa ZenAlexa commented Sep 3, 2026 •

Copy link
Copy Markdown

Why

Interactive input latency needs explicit clock and endpoint definitions. A session-relative input timestamp can be compared with runtime observations when the input source supplies its monotonic-clock origin. The runtime can then report where time accumulates between event arrival, UI consumption, and the next host-side window write.

Closes #538 by adding opt-in V2 host-side input profiling through flashdreams-run-v2 --profile-path.

What changed

  • Add TimestampedInputSource at the input-source boundary. Native-window and WebRTC sources provide the monotonic origin for their session-relative event timestamps; a source without a shared origin produces session metadata and empty latency summaries.
  • Add RuntimeProfiler and pass it through ApplicationRunner into run_session. The presenting rank owns the writer, and model workers reject a supplied profiler before opening an artifact. Profile cleanup follows the existing session cleanup and primary-error handling.
  • Record input_to_ui_step_s when the IUILoop claims an event and input_to_window_write_s when the first following IClientWindow.write returns. The current list-of-frames API writes its single presented frame and records completion after that write.
  • Append independent profile segments for replacement sessions. Each session_started record includes the clock origin, concrete window type, frame layout, rates, resolution, presentation mode, backpressure mode, and measurement endpoints. Reset generations prevent a later write from matching an earlier generation's pending input.
  • Keep exact observation counts and maxima. Median and p90 use all observations through 1,024 samples and then a bounded uniform reservoir, with quantile_sample_count and quantiles_approximate in the artifact. Claimed inputs are retained separately until a following write, generation change, or session close.
  • Reject profile-path conflicts with stats, MP4 output, and an enabled chunk-lifecycle trace. Validate CLI paths after resolving the application's actual window mode.

The current head 47fec3fe691ac170419bf9496ce2c03792436120 merges main at fd52f1a1f4076f86f3fcb86cf5a4425497804b65, preserving application preparation, cross-session step budgets, deadlines, and distributed cleanup.

Verification

Current-head checks on macOS arm64, Python 3.12.14, and PyTorch 2.14.1:

Check Observed result
Existing V2 CPU suite 311 passed, 4 deselected
Final application-runner and window-factory regressions after import formatting 54 passed
Two-process CPU Gloo application lifecycle Two sessions completed with a shared total budget of two model steps; each model step executed a real all_reduce
Profile ownership and replacement Only rank zero wrote the artifact; two session headers and four independent summary records were written; worker ownership was rejected before file creation
Retained profile behavior coverage Timestamp correlation, first-following-write matching, reset generations, missing clock origins, bounded summary samples, repeated close, replacement sessions, and path conflicts passed
Ruff 0.12.7 and diff checks Import ordering, formatting, and git diff --check passed
Scoped ty 0.0.53 Three missing-slangpy import diagnostics; the corresponding unmodified main files reproduced the same three diagnostics in this CPU environment

Run the existing suite with uv run pytest -m ci_cpu flashdreams/test_v2. The suite completed with aiortc pending-task teardown diagnostics in the WebRTC tests.

The metrics end at UI claim and host write return: native-window includes the presenter call, and WebRTC includes queue admission when a video track exists; causal frame response, browser/network delivery, physical display, and GPU generation remain outside these measurements. Profiling writes add host overhead, so direct timing comparisons should enable the same profiling configuration on both runs.

Signed-off-by: Ziming Wang <zimingwang945@gmail.com>
Copilot AI lite review requested due to automatic review settings September 3, 2026 16:46
@copy-pr-bot

copy-pr-bot Bot commented Sep 3, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@greptile-apps

greptile-apps Bot commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Medium risk] Adds host-side input latency profiling to the V2 runtime.

The PR does not appear safe to merge until disconnected WebRTC tracks stop admitting frames without an active sender.

Summary

The PR adds opt-in V2 host-side input-latency profiling with timestamp-origin correlation, per-session JSONL segments, bounded quantile sampling, and CLI artifact-path validation.

  • Propagates a runtime profiler through application and session lifecycle management.
  • Measures input-to-UI-claim and input-to-window-write latency.
  • Adds native-window and WebRTC timestamp origins plus regression coverage and documentation.

Diagram

sequenceDiagram
    participant I as Input source
    participant R as Session runner
    participant U as UI loop
    participant W as Client window
    participant P as Runtime profiler
    I->>R: Session-relative timestamped event
    R->>U: Claim event at UI step
    R->>P: ui_step_started(event, generation)
    U-->>R: Composited frame
    R->>W: write(frame)
    W-->>R: write returns
    R->>P: window_write_completed(generation)
    P-->>P: Emit correlated records and summaries
Loading

Reviews (10) · Last reviewed commit: "fix(runtime): adapt input profiling to c..." · Reviewed by Greptile

Comment thread flashdreams/flashdreams/runtime_v2/serving/webrtc_server.py Outdated
Linearize peer availability and frame admission under the track state lock.
Preserve negotiation queuing and reopen admission after peer recovery.

Signed-off-by: Ziming Wang <zimingwang945@gmail.com>

@ArielG-NV ArielG-NV left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment

Comment thread flashdreams/flashdreams/api_v2/client_window.py Outdated
Keep the timestamp clock bridge on a dedicated input-source extension.
Measure IUILoop claim and the first following window write.
Remove transport-specific and duplicated stage instrumentation.

Signed-off-by: Ziming Wang <zimingwang945@gmail.com>
Define WebRTC timing at the existing single-slot sender mailbox write.
Keep active-peer delivery and display timing in matching client telemetry.

Signed-off-by: Ziming Wang <zimingwang945@gmail.com>
@ZenAlexa

ZenAlexa commented Sep 3, 2026

Copy link
Copy Markdown
Author

Good catch to check this race. input_to_window_write_s deliberately ends at the existing IClientWindow.write return; for WebRTC that is the host's single-slot sender-mailbox write. Active-peer delivery, RTP transit, browser decode, composition, and display live beyond this host-side sample and need client telemetry. _sender_available changes transport queuing behavior; client telemetry supplies the presentation timestamp, so that state stays outside this profiling patch.

…r-input-latency-instrumentation

Signed-off-by: Ziming Wang <zimingwang945@gmail.com>

# Conflicts:
#	flashdreams/flashdreams/runtime_v2/application_runner.py
#	flashdreams/flashdreams/runtime_v2/cli.py
#	flashdreams/flashdreams/runtime_v2/session_runner.py
#	flashdreams/flashdreams/runtime_v2/webrtc_client_window.py
@ZenAlexa

ZenAlexa commented Sep 3, 2026

Copy link
Copy Markdown
Author

Pulled #548's multi-session lifecycle into 5a90e4e9 and followed its new session ownership through the profiler path.

Each replacement now gets a fresh clock binding and an independent JSONL segment, with WebRTC's rebased timestamps mapped back to the correct monotonic session origin. The complete V2 CPU suite passes: 211 passed, 3 deselected. The two metrics remain anchored at IUILoop claim and the first following window write (ง •̀_•́)ง

@ZenAlexa

ZenAlexa commented Sep 3, 2026

Copy link
Copy Markdown
Author

I rechecked this against current HEAD and the PR diff. input_to_window_write_s intentionally ends when the existing IClientWindow.write call returns; for WebRTC that records host-side materialization and mailbox admission, including the existing disconnected-state behavior. Active-peer delivery, RTP transit, decode, composition, and display require client timestamps, so they remain a separate telemetry continuation.

This review keeps the transport contract unchanged and keeps #574 scoped to the two host-side perceived-latency checkpoints (•̀ᴗ•́)و

…r-input-latency-instrumentation

Signed-off-by: Ziming Wang <zimingwang945@gmail.com>

# Conflicts:
#	flashdreams/flashdreams/runtime_v2/session_runner.py
Signed-off-by: Ziming Wang <zimingwang945@gmail.com>
@ZenAlexa

ZenAlexa commented Sep 4, 2026

Copy link
Copy Markdown
Author

Synced current main through #584 in 0e007995 and resolved the new UILoopRequests flow at the profiling boundary. The V2 CPU suite completes with 215 passed / 3 deselected; Ruff, focused ty, compileall, and diff checks pass. The WebRTC endpoint description now follows #579's bounded two-frame sender queue.

Comment thread flashdreams/flashdreams/runtime_v2/runtime_profiler.py
Comment thread flashdreams/flashdreams/runtime_v2/runtime_profiler.py Outdated
Comment thread flashdreams/flashdreams/runtime_v2/runtime_profiler.py Outdated

@jmccaffrey-nv jmccaffrey-nv left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the implementation against #538, including native/WebRTC clock bridging, replacement-session ownership, correlation, cleanup, path collisions, security, and profiler overhead.

The PR satisfies the narrowed host-side checkpoints discussed on the issue: event receipt to IUILoop claim, and event receipt to the next window.write return. The latter is not causal or end-to-end perceived response latency: the UI can re-render a held pre-input frame, and WebRTC stops at host queue admission before transport, decode, composition, and scanout. The documentation states these limits; please keep that distinction explicit when closing #538.

Local validation at 0e007995: 215 V2 CPU tests passed; focused Ruff formatting/import checks, ty, compileall, and git diff --check passed. The documentation build reached the new section without a new warning; its warning-as-error run still reports 10 unrelated baseline warnings. GitHub currently shows only the successful Greptile check while NVIDIA runner validation awaits vetting.

I left three inline comments on self-describing profile metadata, unbounded in-memory summary retention, and stable event-type serialization. I found no new code-execution, deserialization, dependency, or credential-handling exposure.

-- reviewed using GPT-5.6 Sol

@jmccaffrey-nv

Copy link
Copy Markdown
Collaborator

/ok to test 0e00799

…r-input-latency-instrumentation

Signed-off-by: Ziming Wang <zimingwang945@gmail.com>
Signed-off-by: Ziming Wang <zimingwang945@gmail.com>
@ZenAlexa

ZenAlexa commented Sep 4, 2026

Copy link
Copy Markdown
Author

Addressed the three review threads in 020fddb4: session segments now include artifact and runtime context; summary quantiles use a bounded 1,024-sample reservoir with explicit approximation metadata; input types use UserInputEvent.get_type_name(). Endpoint docs now cover WebRTC's active-track and no-track write behavior.

Validation: 185 V2 CPU tests passed, 5 skipped in the local optional-dependency environment; 88 focused tests passed, 1 skipped; Ruff, focused ty, compileall, and diff checks passed. Sphinx built all 38 sources with 10 pre-existing warnings; current main reproduces all 10. A 20,000-input JSONL probe completed at a median 6.14 µs per profiled input across three runs, with 1,024 quantile samples retained per metric.

@greptile-apps

greptile-apps Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Want your agent to iterate on Greptile's feedback? Try greploops.

Signed-off-by: Ziming Wang <zimingwang945@gmail.com>
@ZenAlexa

Copy link
Copy Markdown
Author

I've synced the profiling change with the new model-metrics sink on main.

Both outputs keep their session lifecycle, and replacement sessions retain independent profile segments. I've updated the validation section for the current branch.

Signed-off-by: Ziming Wang <zimingwang945@gmail.com>
Signed-off-by: Ziming Wang <zimingwang945@gmail.com>
@ZenAlexa

ZenAlexa commented Oct 8, 2026

Copy link
Copy Markdown
Author

@jmccaffrey-nv, 47fec3fe preserves the two host-side measurement boundaries and carries main's total step budget across replacement sessions. The current two-rank Gloo scenario completes two sessions with exactly two total model steps per rank, independent profile segments, and a single presenting-rank writer; the V2 CPU suite passes 311 cases. Could you recheck the session ownership changes and authorize /ok to test 47fec3fe691ac170419bf9496ce2c03792436120 for the current runner validation?

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[API] User Input Latency Instrumentation

4 participants