Repository navigation
Recover ambiguous DCP creates with configurable timeouts - #20752
Conversation
Confirm timed-out creates before retrying the same named object, bound create recovery, and share readiness-aware request deadlines across client setup and API operations. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
|
🚀 Dogfood this PR with:
curl -fsSL https://raw.githubusercontent.com/microsoft/aspire/main/eng/scripts/get-aspire-cli-pr.sh | bash -s -- 20752Or
iex "& { $(irm https://raw.githubusercontent.com/microsoft/aspire/main/eng/scripts/get-aspire-cli-pr.ps1) } 20752" |
This comment has been minimized.
This comment has been minimized.
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
The cross-cutting retry and timeout lifecycle changes lack live DCP or end-to-end validation.
Review effort: Balanced
Findings: None
What changed in this PR
Adds resilient DCP object creation and configurable Kubernetes API deadlines for slow environments.
Changes:
- Reconciles ambiguous creates before retrying.
- Adds validated API, initialization, and recovery timeout settings.
- Applies Aspire deadlines consistently while preserving established streams.
| File | Description |
|---|---|
src/Aspire.Hosting/Dcp/KubernetesService.cs |
Implements recovery and shared deadline policies. |
src/Aspire.Hosting/Dcp/DcpOptions.cs |
Adds timeout configuration and validation. |
src/Aspire.Hosting/Dcp/DcpKubernetesClient.cs |
Disables the SDK’s independent timeout. |
tests/Aspire.Hosting.Tests/Dcp/KubernetesServiceTests.cs |
Covers recovery, deadlines, cancellation, and streams. |
tests/Aspire.Hosting.Tests/Dcp/ConfigureDefaultDcpOptionsTests.cs |
Covers defaults, configuration, and validation. |
💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.
|
Retrying the failed CI jobs for this pull request from the CI run attempt. The rerun is being tracked in the rerun attempt. |
PR Testing ReportPR Information
Artifact Version Verification
The package repository commit is the normal CI merge commit. Its parents are base The Tests workflow Changes AnalyzedFiles Changed
Change Categories
Environment Setup and Exact CommandsAll five scenario apps were freshly generated with Host runner prefix, with ASPIRE_PR_WORKSPACE="$testDir" ASPIRE_CONTAINER_USER=0:0 \
./eng/scripts/aspire-pr-container/run-aspire-pr-container.sh <command>Inside the runner, installation used the PR's published dogfood command: curl -fsSL https://raw.githubusercontent.com/microsoft/aspire/main/eng/scripts/get-aspire-cli-pr.sh \
| bash -s -- 20752Result: Passed. The installer selected workflow run Resolved prerequisite failureThe first normal AppHost start failed before DCP launched: The runner's base image does not include a .NET SDK. The missing prerequisite was curl -fsSL https://dot.net/v1/dotnet-install.sh \
| bash -s -- --channel 10.0 --install-dir /workspace/.aspire/dotnet --no-path
export DOTNET_ROOT=/workspace/.aspire/dotnet
export PATH="$DOTNET_ROOT:$PATH"Result: Installed SDK Fresh project creationThe following command was executed separately for cli=/workspace/.aspire/dogfood/pr-20752/bin/aspire
"$cli" new aspire-empty \
--name NormalSmoke \
--output /workspace/NormalSmoke \
--source /workspace/.aspire/hives/pr-20752/packages \
--version 17.0.0-pr.20752.g065920e9 \
--language csharp \
--localhost-tld false \
--suppress-agent-init \
--non-interactiveResult: All five fresh C# file-based AppHosts were created successfully. Test Scenarios ExecutedScenario 1: Normal DCP orchestrationObjective: Exercise actual DCP creation, process startup, dependency ordering, Coverage type: Happy path. Status: Passed after installing the missing SDK prerequisite. The AppHost declared three shell executable resources with a dependency chain apphost=/workspace/NormalSmoke/apphost.cs
"$cli" start --apphost "$apphost" --launch-profile http --isolated --non-interactive --format Json
for name in root middle leaf; do
"$cli" wait "$name" --status up --timeout 120 --apphost "$apphost" --non-interactive
done
"$cli" describe --apphost "$apphost" --format Json --include-hidden --non-interactive
"$cli" logs --apphost "$apphost" --tail 10 --non-interactive
"$cli" stop --apphost "$apphost" --non-interactive
"$cli" ps --format Json --non-interactiveObservations:
Evidence: Scenario 2: Resource burst and established stream lifetimeObjective: Exercise concurrent real DCP process creation and confirm that a Coverage type: Happy path and supported configuration boundary. Status: Passed. A fresh AppHost declared 40 shell processes and configured: The test script started the AppHost, launched a native CLI readiness wait for Key commands: apphost=/workspace/BurstSmoke/apphost.cs
"$cli" start --apphost "$apphost" --launch-profile http --isolated --non-interactive --format Json
"$cli" wait worker-00 --status up --timeout 120 --apphost "$apphost" --non-interactive
# The same readiness command was run for every worker through worker-39.
"$cli" describe --apphost "$apphost" --format Json --non-interactive
"$cli" describe worker-00 --follow --format Json --apphost "$apphost" --non-interactive
"$cli" logs worker-00 --follow --format Json --apphost "$apphost" --non-interactive
# After 12 seconds:
"$cli" logs worker-00 --tail 30 --apphost "$apphost" --non-interactive
"$cli" resource worker-00 restart --apphost "$apphost" --non-interactive
"$cli" wait worker-00 --status up --timeout 120 --apphost "$apphost" --non-interactive
"$cli" stop --apphost "$apphost" --non-interactive
"$cli" ps --format Json --non-interactiveObservations:
Evidence: Scenario 3: Invalid timeout configurationObjective: Verify the actual AppHost startup path rejects invalid timeout Coverage type: Unhappy path and bounds. Status: Passed. Three separate fresh AppHosts used the following invalid configurations:
For each AppHost: "$cli" start --apphost "/workspace/$name/apphost.cs" \
--launch-profile http --isolated --non-interactive --format JsonObservations: Each CLI start returned nonzero ( Evidence: Scenario 4: HTTP-backed recovery and options regressionsObjective: Validate failure paths that cannot be deterministically forced in Coverage type: Happy path, unhappy path, cancellation, and timeout boundaries. Status: Passed: 97 tests, zero failures, zero skips. Executed from the source worktree at the exact PR head: MSBUILDTERMINALLOGGER=false ./dotnet.sh test \
--project tests/Aspire.Hosting.Tests/Aspire.Hosting.Tests.csproj \
--no-launch-profile --no-restore /p:SkipNativeBuild=true \
-- \
--filter-class '*.KubernetesServiceTests' \
--filter-class '*.ConfigureDefaultDcpOptionsTests' \
--filter-not-trait 'quarantined=true' \
--filter-not-trait 'outerloop=true'The run compiled and passed in approximately 58 seconds. Coverage includes No partial-response-body stall scenarios were added or run. Evidence: Evidence VerificationThe structured JSON/NDJSON artifacts were checked independently of the CLI's Result: Passed. See Summary
Overall ResultPR VERIFIED for the approved testing scope. Limitations
Report, Evidence, and Cleanup Status
|
Preserve the DCP recovery tests while removing the filesystem warning suppression retired on main. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Tests selector50 / 98 PR test projects · 4 PR jobs, from 5 changed files. Selected PR test projects (50 / 98)
Selected PR jobs (4)
How these were chosen — grouped by what changed
🔧 show 45
🧪 📦 affected project 🧪 Job reasons
Selection computed for commit |
|
Retrying the failed CI jobs for this pull request from the CI run attempt. The rerun is being tracked in the rerun attempt. |
James Newton-King (JamesNK)
left a comment
There was a problem hiding this comment.
No high-confidence blocking issues found. Reviewed recovery design, cancellation and retry edge cases, configuration validation, and regression coverage. Validation at this head: 62/62 focused KubernetesServiceTests passed; controlled stalled-create recovery succeeded; live Windows DCP started a resource with custom timeout settings and shut down cleanly; invalid timeout configuration was rejected. The original macOS Intel HTTP 504 scenario was not reproduced exactly.
|
/backport to release/13.6 |
|
Started backporting to |
Documents the new DcpPublisher:KubernetesApiTimeout, KubernetesInitializationAdditionalTimeout, and KubernetesCreateRecoveryTimeout options added in microsoft/aspire#20752, along with the ambiguous-create recovery behavior they enable. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
Pull request created: #1862
|
|
📝 Documentation has been drafted in microsoft/aspire.dev#1862 targeting Triggered signals: Note This draft PR needs human review before merging. |

Description
A single timed-out or HTTP 504 DCP create currently puts a resource into
FailedToStart, even though the server may have committed the object. The 20-second steady-state client budget can also expire before DCP's own create timeout, and users cannot configure those limits for slower environments.Recover ambiguous creates by checking whether the same kind/namespace/name exists before retrying the POST. Creation has a configurable two-minute recovery budget covering setup, requests, confirmation, and retry delays. Initial conflicts and permanent failures still fail, caller cancellation remains terminal, and successful sibling creates are not replayed.
API operations now share one deadline between client setup and the request/retry work: 40 seconds plus a 20-second initialization allowance until the first successful API operation, then 40 seconds. Each operation selects its deadline once. Watch reconnects reselect the current budget without shortening in-flight requests or limiting established watch/log streams.
KubernetesClient's independent 100-second request timer no longer truncates larger configured budgets. Production requests use Aspire's timeout policy, including the cleanup polling GET that previously bypassed it. No new dependencies are required.
Fixes #20691
User-facing usage
The default timeout settings can now be overridden in AppHost configuration:
{ "DcpPublisher": { "KubernetesApiTimeout": "00:00:40", "KubernetesInitializationAdditionalTimeout": "00:00:20", "KubernetesCreateRecoveryTimeout": "00:02:00" } }The initialization allowance may be zero. API and recovery budgets support one second through ten minutes; the API budget plus initialization allowance must not exceed ten minutes.
Validation
Passed: 97 selected tests, with no failures or skips. Coverage includes ambiguous-create confirmation/retries, permanent failures, caller cancellation, configuration validation, readiness-aware deadlines, watch reconnects, and established stream lifetimes.
No live DCP reproduction, end-to-end validation, or full test-suite run was performed for the final changes.
Checklist
<remarks />and<code />elements on your triple slash comments?