Skip to content

CI: move redundant scene tests to daily sweep - #1790

Merged
ChaoZheng109 merged 1 commit into
hw-native-sys:mainfrom
doraemonmj:ci/manual-daily-tests
Aug 13, 2026
Merged

ChaoZheng109 merged 1 commit into
hw-native-sys:mainfrom
doraemonmj:ci/manual-daily-tests

Conversation

@doraemonmj

@doraemonmj doraemonmj commented Aug 11, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Reduce Per-PR scene-test cost by moving redundant smoke, scale-only,
    timing/DFX-observation, P4 collective, and selected duplicate Sim cases
    behind the existing manual selection.
  • Keep --manual exclude as the reusable workflow default for PR CI; the
    separate Daily workflow uses --manual include to retain the full corpus.
  • Extend manual selection to standalone pytest tests and platform-scoped
    markers, and pass the same mode to pre-compilation so only selected test
    classes are compiled.
  • Validate manual_mode as exclude, include, or only before using it in
    shell commands.
  • Keep case parameters, kernels, golden checks, tolerances, runtime settings,
    and workflow timeout thresholds unchanged.

Per-PR coverage policy

Cases are moved out of Per-PR when their value is primarily shallow smoke,
rank/shape scale expansion, performance/marker observation, or a duplicate
execution path with a stronger representative still running. Per-PR retains
production-model paths, fault/recovery/stress coverage, large pressure cases,
and representative algorithm coverage on each supported platform.

Paged-attention shapes are unchanged, and A2/A3 Sim still retains the
500-task BGEMM pressure case.

Cases removed from Per-PR

Single-card and observability cases

Test Per-PR platforms removed Coverage retained in Per-PR
test_hello_worker A2/A3 Sim, A3 Onboard, A5 Sim, A5 Onboard Worker allocation/copy/free and close paths
test_vector_add A2/A3 Sim, A3 Onboard Other runtime vector computation paths
TestMatmulHostBuildGraph::default A2/A3 Sim, A3 Onboard Graph record/replay path exercising the same incore kernels
TestFaninLookupPerf 64x64 cases A2/A3 Sim, A3 Onboard, A5 Sim, A5 Onboard Dependency-generation boundary and representative fanout/fanin cases
TestDummyTask::LongDummyChain A2/A3 Sim, A3 Onboard, A5 Sim, A5 Onboard Smaller dependency-chain and dense fanout/fanin cases
TMR/HBG vector examples A2/A3 Sim, A5 Sim Same cases remain Onboard; stronger computation paths remain on Sim
Graph execution 1D A2/A3 Sim Graph execution 2D remains on Sim; 1D remains Onboard
BGEMM 64 A2/A3 Sim 500-task Case0 remains on Sim; BGEMM 64 remains Onboard
Run-timing and task-timing marker tests Supported Sim platforms Onboard marker coverage and Per-PR vector correctness remain
ffn_tp_parallel A2/A3 Sim, A5 Sim Same two-card matmul + allreduce path remains A3 Onboard
dual_domain_overlap A2/A3 Sim, A5 Sim domain_rank_map remains on both Sim lanes; the full overlap computation remains A3 Onboard

Collective cases

Tests Per-PR platforms removed Coverage retained in Per-PR
Allgather, AllToAll, Broadcast, and ReduceScatter P4 A2/A3 Sim, A3 Onboard, A5 Sim The corresponding P2 collective remains on both Sim lanes and A3 Onboard
Allreduce One-phase, Two-phase, Ring, and Bidirectional Ring P4; Ibing four-rank error A2/A3 Sim, A3 Onboard, A5 Sim One-phase and Ring P2 remain on both Sim lanes; all five P2 algorithms remain on A3 and A5 Onboard
Allreduce Two-phase, Bidirectional Ring, and Ibing P2 A2/A3 Sim, A5 Sim One-phase and Ring P2 remain on both Sim lanes; all five P2 algorithms remain on A3 and A5 Onboard

A5 Onboard is unchanged by the P4 migration because the P4 classes are not
defined for that platform. Daily continues to run every migrated P2 and P4
case on its supported platforms.

CI compatibility fixes

  • The profiling-flags smoke targets a vector example that is now manual on Sim,
    so it explicitly passes --manual include; otherwise pytest deselects the
    only target and exits non-zero.
  • All four reusable Sim/Onboard workflows validate manual_mode through an
    environment variable before shell parsing, then reuse only the validated
    value in pytest, pre-compile, and task-submit --run commands.
  • No-hardware C++ UT builds use four-way CMake build parallelism to avoid the
    macOS serial-build tail that previously exhausted the 20-minute job timeout.

Validation

  • Targeted profiling smoke target passed on A2/A3 Sim and A5 Sim.
  • Collective selection checks verify every P4 class is manual, while the three
    additional P2 cases are manual only on the two Sim platforms.
  • python -m pytest tests/ut/py/test_scene_level_selection.py -q: 8 passed.
  • Targeted pre-commit checks passed for the changed Python, YAML, Markdown, and
    repository policy files.
  • git diff --check upstream/main...HEAD: passed.

@coderabbitai

coderabbitai Bot commented Aug 11, 2026 •

Copy link
Copy Markdown

Review Change Stack

Important

Review skipped

Auto incremental reviews are disabled on this repository.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: c44dc7ff-22a8-4138-851a-24c536f5ca16

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

The PR adds platform-aware manual test selection for standalone pytest tests and scene-test cases. It propagates manual_mode through reusable workflows, adds a daily full-test sweep, and marks selected simulator and hardware tests as manual.

Changes

Manual test coverage

Layer / File(s) Summary
Manual filtering and resource dispatch
conftest.py, simpler_setup/scene_test.py
--manual now validates markers, matches platforms, filters scene cases and standalone tests, and preserves the mode in resource jobs.
Workflow manual-mode propagation
.github/workflows/_st-*.yml, .github/workflows/_st-sim-*.yml
Reusable workflows accept manual_mode and pass it to scene-test compilation and pytest commands. SDMA paths skip only runs.
Daily full-test workflow
.github/workflows/daily.yml
A scheduled and manually dispatchable workflow runs simulation, onboarding, DFX, and pod jobs with manual tests included.
Manual annotations and documentation
tests/st/**, examples/**, docs/**, .claude/skills/testing/SKILL.md
Selected cases and standalone tests receive manual metadata. Testing, CI, CLI, and example documentation describe platform-specific selection.

Estimated code review effort: 3 (Moderate) | ~25 minutes

Sequence Diagram(s)

sequenceDiagram
  participant pytest
  participant conftest
  participant scene_test
  participant CI workflow
  pytest->>conftest: Read --manual
  conftest->>scene_test: Match manual case to platform
  scene_test-->>conftest: Return selection status
  conftest->>pytest: Filter tests and resource jobs
  CI workflow->>pytest: Pass manual_mode as --manual
Loading

Possibly related PRs

Poem

A rabbit flags tests in rows,
For sim platforms, marks it shows.
Daily hops through every scene,
While PR runs stay lean and clean.
--manual guides each testing trail.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 44.44% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description check ✅ Passed The description directly explains moving redundant scene tests to Daily CI and extending manual test selection.
Title check ✅ Passed The title clearly summarizes the primary change: moving redundant scene tests from Per-PR CI to a daily sweep.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.github/workflows/_st-npu-a2a3.yml:
- Around line 27-31: Validate manual_mode via an environment variable and a
shell case statement accepting only exclude, include, or only before any command
uses it. In .github/workflows/_st-npu-a2a3.yml ranges 27-31 and 78-90,
.github/workflows/_st-npu-a5.yml ranges 22-26 and 65-69,
.github/workflows/_st-sim-a2a3.yml ranges 26-30 and 105-107, and
.github/workflows/_st-sim-a5.yml ranges 26-30 and 105-107, replace direct input
interpolation in every pytest and task-submit --run command with the validated
variable.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 83e14adb-d084-4bbc-8faf-727c1af9a0d0

📥 Commits

Reviewing files that changed from the base of the PR and between c5a5874 and 6eb9cd2.

📒 Files selected for processing (32)
  • .claude/skills/testing/SKILL.md
  • .github/workflows/_st-npu-a2a3.yml
  • .github/workflows/_st-npu-a5.yml
  • .github/workflows/_st-sim-a2a3.yml
  • .github/workflows/_st-sim-a5.yml
  • .github/workflows/daily.yml
  • conftest.py
  • docs/ci.md
  • docs/testing.md
  • docs/user/reference/cli.md
  • examples/a2a3/tensormap_and_ringbuffer/README.md
  • examples/a2a3/tensormap_and_ringbuffer/benchmark_bgemm/README.md
  • examples/a2a3/tensormap_and_ringbuffer/benchmark_bgemm/test_benchmark_bgemm.py
  • examples/a2a3/tensormap_and_ringbuffer/vector_example/test_vector_example.py
  • examples/a5/tensormap_and_ringbuffer/vector_example/test_vector_example.py
  • examples/workers/l2/hello_worker/test_hello_worker.py
  • examples/workers/l2/vector_add/test_run_timing.py
  • examples/workers/l2/vector_add/test_vector_add.py
  • simpler_setup/scene_test.py
  • tests/st/a2a3/host_build_graph/graph_execution/test_graph_execution.py
  • tests/st/a2a3/host_build_graph/matmul/test_matmul.py
  • tests/st/a2a3/host_build_graph/vector_example/test_vector_example.py
  • tests/st/a2a3/tensormap_and_ringbuffer/dummy_task/test_dummy_task.py
  • tests/st/a2a3/tensormap_and_ringbuffer/fanin_lookup_perf/test_fanin_lookup_perf.py
  • tests/st/a5/host_build_graph/vector_example/test_vector_example.py
  • tests/st/a5/tensormap_and_ringbuffer/dummy_task/test_dummy_task.py
  • tests/st/a5/tensormap_and_ringbuffer/fanin_lookup_perf/test_fanin_lookup_perf.py
  • tests/st/task_timing/task_timing_slots/test_task_timing_e2e.py
  • tests/st/worker/collectives/allgather/test_allgather.py
  • tests/st/worker/collectives/allreduce/test_allreduce.py
  • tests/st/worker/collectives/broadcast/test_broadcast.py
  • tests/st/worker/collectives/reduce_scatter/test_reduce_scatter.py

Comment thread .github/workflows/_st-npu-a2a3.yml
@doraemonmj
doraemonmj force-pushed the ci/manual-daily-tests branch 3 times, most recently from 1282e85 to d2bc9ca Compare August 12, 2026 06:21
@ChaoZheng109

Copy link
Copy Markdown
Collaborator

Overall LGTM — 机制干净,PR body 的移除清单与实际标记逐条对得上,与 ci-change-detection 规则(schedule 放独立 daily.yml、Per-PR 复用 exclude 默认值)也吻合。两点建议:

1. 新判定逻辑缺少单元测试 (Should-fix)

PR 描述里引用的 tests/ut/py/test_scene_level_selection.py: 8 passed,那 8 个测试全部是 level 选择逻辑,与 manual 无关。本 PR 新增的三个函数:

  • simpler_setup/scene_test.py::is_manual_for_platform
  • conftest.py::_manual_marker_applies
  • conftest.py::_manual_mode_matches

目前没有任何直接的单测覆盖。而这几处恰恰有不少非平凡分支值得钉住:

  • manual 为 True / str / list/tuple/set / None 的各自返回;
  • @pytest.mark.manual 的三种写法(裸标记、位置参数 manual([...])、关键字 manual(platforms=[...]));
  • 混用位置+关键字 platforms、传入未知 kwarg 时应 raise UsageError;
  • exclude / include / only 三态在「按平台 manual」下的组合。

建议补一个针对这三个 helper 的 UT(尤其是错误分支),这样后续改动能有回归护栏。

2. --manual only 会连带筛掉普通独立测试 (Consider)

conftest.py::pytest_collection_modifyitems 里新增的 manual 过滤对所有独立(非 class)pytest 函数生效。在 --manual only 下,判定是「没打 manual 标记的一律 deselect」,于是同一 session 里那些跟 scene-test / manual 机制无关的普通测试也会被一并剔除。

严格说这符合 only 的语义(「只跑 manual」),CI 里也没用到 only(Per-PR=exclude、Daily=include),所以无线上风险。只是本地手动敲 pytest tests/ --manual only 的人可能会困惑「我的普通单测怎么没跑」。建议在 docs/testing.md 的 --manual 说明里点一句这个行为即可,不必改代码。

- Add a scheduled full scene-test workflow for Sim, Onboard, and Pod lanes
- Extend manual selection to standalone tests and platform-scoped cases
- Validate reusable manual-mode inputs before shell execution
- Move redundant single-card, timing, and P4 communication coverage while retaining representative Per-PR paths
- Cover manual selector branches and document session-wide only semantics
- Keep existing case parameters, kernels, golden checks, and timeout thresholds unchanged
@doraemonmj
doraemonmj force-pushed the ci/manual-daily-tests branch from d2bc9ca to 6335147 Compare August 12, 2026 09:30
@ChaoZheng109
ChaoZheng109 merged commit d31395b into hw-native-sys:main Aug 13, 2026
19 checks passed
@doraemonmj
doraemonmj deleted the ci/manual-daily-tests branch August 18, 2026 02:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants