Repository navigation
[Bug] L3 Worker leaves chip child defunct after submit_next_level run #980
Description
Activity
A PyPTO Serving PR that reproduces and works around this Worker lifecycle issue is now available:
- PR: Rewrite non-L3 Qwen3 kernels through L3 worker pypto-serving#22
- Commit: ndleslx/pypto-serving@b2080dc
- Branch:
ndleslx:runner-l3-worker
The PR rewrites the non-L3 Qwen3 generation path to submit compiled chip callables through
Worker(level=3)+orch.submit_next_level(...). In that integration, reusing the same L3 worker after the prefillWorker.run()caused the chip child to become defunct and the parent process to hang before decode. The PR includes a wrapper-level one-shot discard workaround so each submitted non-L3 kernel uses a fresh L3 worker child.Passing reproduction with the workaround and larger ring settings:
cd /data/liuxu/pypto-serving task-submit --device auto --max-time 0 --run \ "PTO2_RING_HEAP=4294967296 PTO2_RING_TASK_WINDOW=1048576 PTO2_RING_DEP_POOL=1048576 \ python examples/model/qwen3_14b/npu_generate.py \ --model-dir /data/linyifan/models/Qwen3-14B \ --prompt 'Huawei is' \ --platform a2a3 \ --max-seq-len 512 \ --max-new-tokens 5"
Observed output:
text: a Chinese company. The token_ids: [264, 8453, 2813, 13, 576] finish_reason: lengthThe smaller/default ring settings still fail during prefill with AICPU
507018, which appears to be a separate resource/timeout issue for this workload:PTO2_RING_HEAP=536870912 PTO2_RING_TASK_WINDOW=131072 PTO2_RING_DEP_POOL=131072
Findings from the pypto-serving multi-program L3-worker investigation:
The child process does not become defunct after prefill. It becomes defunct during the first decode dispatch.
Observed with
SA_L3_DEBUG=1around everyDistributedWorker.run(...)call:[chip_process pid=1548818 dev=4] ready [l3-debug] before dispatch=1 kernel=prefill_fwd chip_pids=['1548818:R:exit=0'] [l3-debug] after dispatch=1 kernel=prefill_fwd chip_pids=['1548818:R:exit=0'] [timing] prefill: fused 40 layers, 19043.06 ms [l3-debug] before dispatch=2 kernel=decode_fwd chip_pids=['1548818:R:exit=0']There was no
after dispatch=2.An external monitor then showed the chip child transition while the parent was still waiting in the decode
worker.run(...):pid=1548818 state=D ppid=1543291 exit_code=0 pid=1548818 state=Z ppid=1543291 exit_code=0So the immediate failure is:
worker.run(decode_fwd) starts first decode -> chip child enters decode execution -> chip child exits before publishing TASK_DONE -> parent waits foreverThe relevant child path is in
simpler/python/simpler/worker.py:cw._impl.run_prepared_from_blob(cid, mailbox_addr + _OFF_ARGS, _MAILBOX_ARGS_CAPACITY, cfg) _mailbox_store_i32(state_addr, _TASK_DONE)
The child died before the
TASK_DONEpublish. Since the zombie exit code was0, this does not look like taskqueue killing it or a signal crash. It looks like the chip child returned/exited normally from the L3 runtime/native path while the parent still expected decode completion.Separate note: one later debug run ended with task
exit=138; that was caused by a manualSIGUSR1probe sent to the parent process after the child had already become zombie. It should not be treated as the root cause.Also separate from this issue: the earlier post-output
exit=137was a lifecycle leak. The serving parent exited without closing the L3 worker, leaving child processes for taskqueue to sweep. Adding explicitexecutor.close()fixed that shutdown symptom, but it is not the same as the first-decode child-exit problem described here.- added 6 commits that reference this issue
on Jul 18, 2026
Platform
a2a3 (Ascend 910B/C hardware)
Runtime Variant
tensormap_and_ringbuffer
Description
A level-3
simpler.worker.Workerused as a single-chip host worker can become unusable after a successfulWorker.run()that submits a chip callable throughorch.submit_next_level(...).In the observed PyPTO serving integration, the prefill kernel completes successfully and
Worker.run()returns to Python, but the chip child process is already left as a defunct process while the parentWorkerstill appears initialized. Reusing the same L3 worker for the next submitted chip task, or closing/switching the worker through the normal wrapper path, hangs instead of cleanly reusing or tearing down the child. The local workaround was to treat this L3 worker as one-shot: after every submitted child task, write_SHUTDOWNto child mailboxes,waitpidthe children, unlink shared-memory mailboxes, and discard the worker state before creating a new Worker for the next kernel.This looks like a Worker lifecycle bug: either
Worker.run()should keep the chip child alive for later runs, orWorker.close()/ post-run cleanup should reliably reap and mark the worker unusable when the child has exited.Related: #824
Steps to Reproduce
Using a PyPTO Serving branch that dispatches non-L3 Qwen3 kernels through an L3 Simpler worker:
The serving-side dispatch shape is roughly:
Expected Behavior
After a successful
Worker.run():Worker.run()on the same chip child, orWorker.close()should not hang after the child has already exited.Actual Behavior
The first submitted prefill task completes:
Immediately afterward, process inspection shows the child process as defunct while the parent Python process remains alive with the Worker still in use:
The parent then makes no progress into the next decode task. In repeated checks it had to be killed manually. Before the one-shot discard workaround, this blocked offline generation after prefill. With the manual one-shot discard/recreate workaround, the same generation completed:
A separate resource-related symptom was also observed with small ring settings (
PTO2_RING_HEAP=536870912 PTO2_RING_TASK_WINDOW=131072 PTO2_RING_DEP_POOL=131072): prefill can fail with AICPU507018. The lifecycle bug above was reproduced with the larger ring settings where prefill itself succeeds.Git Commit ID
293e88a
CANN Version
9.0.0 (
Ascend-cann-toolkit,innerversion=V100R001C10SPC001B250)Driver Version
npu-smireports version26.0.rc1.Host Platform
Linux (aarch64)
Additional Context
The workaround currently used in PyPTO Serving adds a wrapper-level best-effort discard path for one-shot L3 workers:
_SHUTDOWNinto_sub_shms,_chip_shms, and_next_level_shmswaitpidchild PIDs_worker,_orch, child PID/shm lists, and initialized stateThat avoids the hang but relies on Simpler private internals, so the lifecycle should be fixed or exposed as a supported API in Simpler.