Skip to content

ci : reduce self-hosted server workflow jobs - #24012

Merged
ggerganov merged 1 commit into
masterfrom
gg/ci-reduce-server-self-hosted-jobs
Jun 2, 2026
Merged

ggerganov merged 1 commit into
masterfrom
gg/ci-reduce-server-self-hosted-jobs

Conversation

@ggerganov

@ggerganov ggerganov commented Jun 2, 2026 •

Copy link
Copy Markdown
Member

Overview

Reduce the number of parallel jobs in server-self-hosted.yml by stacking test configurations as sequential steps within a single job, following the pattern from #23927.

  • server-metal: 4 matrix jobs → 1 job with 4 sequential test steps (GPUx1, GPUx1+backend-sampling, GPUx2, GPUx2+backend-sampling)
  • server-cuda: 2 matrix jobs → 1 job with 2 sequential test steps (GPUx1, GPUx1+backend-sampling)
  • server-kleidiai: Removed unnecessary single-entry matrix
  • Removed unused "Setup Node.js" step from server-metal

Total: 7 parallel jobs → 3 parallel jobs.

Additional information

Tests that share the same binaries now run sequentially with different environment variables instead of as separate parallel jobs.

Sample run: https://github.com/ggml-org/llama.cpp/actions/runs/26805755015

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. llama.cpp + local pi + Qwen3.6-27B

Reduce the number of parallel jobs in server-self-hosted.yml by stacking
test configurations as sequential steps within a single job, following the
pattern from #23927.

- server-metal: 4 matrix jobs -> 1 job with 4 sequential test steps
- server-cuda: 2 matrix jobs -> 1 job with 2 sequential test steps
- server-kleidiai: removed unnecessary single-entry matrix
- removed unused Setup Node.js step from server-metal

Total: 7 parallel jobs -> 3 parallel jobs

Assisted-by: llama.cpp:local pi
@ggerganov ggerganov changed the title ci(self-hosted) : reduce server workflow jobs ci : reduce self-hosted server workflow jobs Jun 2, 2026
@github-actions github-actions Bot added the devops improvements to build systems and github actions label Jun 2, 2026
@ggerganov
ggerganov marked this pull request as ready for review June 2, 2026 09:30
@ggerganov
ggerganov requested a review from a team as a code owner June 2, 2026 09:30
@ggerganov
ggerganov merged commit a468b89 into master Jun 2, 2026
6 checks passed
@ggerganov
ggerganov deleted the gg/ci-reduce-server-self-hosted-jobs branch June 2, 2026 10:18
adrianhoehne pushed a commit to adrianhoehne/llama.cpp that referenced this pull request Jul 5, 2026
Reduce the number of parallel jobs in server-self-hosted.yml by stacking
test configurations as sequential steps within a single job, following the
pattern from ggml-org#23927.

- server-metal: 4 matrix jobs -> 1 job with 4 sequential test steps
- server-cuda: 2 matrix jobs -> 1 job with 2 sequential test steps
- server-kleidiai: removed unnecessary single-entry matrix
- removed unused Setup Node.js step from server-metal

Total: 7 parallel jobs -> 3 parallel jobs

Assisted-by: llama.cpp:local pi
zommiommy pushed a commit to zommiommy/llama.cpp that referenced this pull request Aug 18, 2026
Reduce the number of parallel jobs in server-self-hosted.yml by stacking
test configurations as sequential steps within a single job, following the
pattern from ggml-org#23927.

- server-metal: 4 matrix jobs -> 1 job with 4 sequential test steps
- server-cuda: 2 matrix jobs -> 1 job with 2 sequential test steps
- server-kleidiai: removed unnecessary single-entry matrix
- removed unused Setup Node.js step from server-metal

Total: 7 parallel jobs -> 3 parallel jobs

Assisted-by: llama.cpp:local pi
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
Reduce the number of parallel jobs in server-self-hosted.yml by stacking
test configurations as sequential steps within a single job, following the
pattern from ggml-org#23927.

- server-metal: 4 matrix jobs -> 1 job with 4 sequential test steps
- server-cuda: 2 matrix jobs -> 1 job with 2 sequential test steps
- server-kleidiai: removed unnecessary single-entry matrix
- removed unused Setup Node.js step from server-metal

Total: 7 parallel jobs -> 3 parallel jobs

Assisted-by: llama.cpp:local pi
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
Reduce the number of parallel jobs in server-self-hosted.yml by stacking
test configurations as sequential steps within a single job, following the
pattern from ggml-org#23927.

- server-metal: 4 matrix jobs -> 1 job with 4 sequential test steps
- server-cuda: 2 matrix jobs -> 1 job with 2 sequential test steps
- server-kleidiai: removed unnecessary single-entry matrix
- removed unused Setup Node.js step from server-metal

Total: 7 parallel jobs -> 3 parallel jobs

Assisted-by: llama.cpp:local pi
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

devops improvements to build systems and github actions

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants