Skip to content

Use shared GPU templates for data-loading evals - #6500

Open
chrisknvidia wants to merge 2 commits into
data-loading-bottleneck-skillfrom
fix/christopherk/dali-gpu-template-config
Open

chrisknvidia wants to merge 2 commits into
data-loading-bottleneck-skillfrom
fix/christopherk/dali-gpu-template-config

Conversation

@chrisknvidia

@chrisknvidia chrisknvidia commented Sep 22, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Route Claude Code and Codex evals to their matching shared DALI GPU SandboxTemplates.
  • Copy the baked workload bundle from /opt/eval/workloads before each agent starts.
  • Rename the canonical evals/environment/Dockerfile to the inactive evals/.reference/Dockerfile.dali-gpu-image path. This preserves the recipe for future custom-Dockerfile support while preventing current SkillEvaluator image-builder discovery and runtime projection.

Validation

  • SkillEvaluator 1.5.6 strict Tier 3 contract validation passed.

  • Targeted Tier 1 security validation passed with no critical or high findings.

  • The custom-Dockerfile resolver returns no Dockerfile, and the private evaluator snapshot excludes evals/.reference.

  • The renamed recipe is byte-for-byte identical to the original Dockerfile.

  • Real Harbor smoke using harbor-eval-dali-gpu-claude-code completed both with-skill and baseline attempts: 2/2 scored, no execution, runtime, trial, or job failures.

  • A pre-promotion full shared-image matrix across all three cases, both agents, and both arms completed 12/12 attempts; the final-name Claude smoke above verifies the promoted template name.

Dependencies

  • This PR is intentionally stacked on Add data-loading bottleneck diagnostic skill #6466.
  • Do not merge until both named SandboxTemplates are durably deployed in the NVSkills CI PDX04 environment. This DALI change selects the templates; it does not deploy them.
  • dali-dynamic-mode changes are intentionally out of scope.

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented Sep 22, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@chrisknvidia

Copy link
Copy Markdown
Collaborator Author

/nvskills-ci

@greptile-apps

greptile-apps Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

The repository changes appear safe to merge once the externally managed SandboxTemplates named by the configuration are durably deployed.

Summary

This PR moves the data-loading evaluations from a custom GPU Dockerfile to agent-specific shared GPU SandboxTemplates.

  • Selects dedicated Claude Code and Codex DALI GPU templates.
  • Copies baked workload repositories into the paths expected by the evaluation tasks.
  • Retains the former image recipe as an inert reference artifact rather than an active environment Dockerfile.
  • Requires both external SandboxTemplates to remain durably deployed before merge.

Diagram

%%{init: {'theme': 'neutral'}}%%
flowchart LR
  C[Evaluation config] --> A{Agent}
  A -->|claude-code| CT[Claude DALI GPU template]
  A -->|codex| XT[Codex DALI GPU template]
  CT --> S[Run pre-agent setup]
  XT --> S
  S --> W[Copy /opt/eval/workloads to /workspace/workloads]
  W --> E[Run data-loading evaluation]
Loading

Reviews (2) · Last reviewed commit: "Preserve disabled DALI GPU image recipe"

Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
@chrisknvidia

Copy link
Copy Markdown
Collaborator Author

@mdabek-nvidia : Can you please trigger /nvskills-ci ?
Would like to see how GPU based eval works.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants