Skip to content

[SYCL] support sparse FA - #28796

Merged
ggerganov merged 3 commits into
ggml-org:masterfrom
arthw:support_sparse_fa
Sep 25, 2026
Merged

ggerganov merged 3 commits into
ggml-org:masterfrom
arthw:support_sparse_fa

Conversation

@arthw

@arthw arthw commented Sep 12, 2026

Copy link
Copy Markdown
Contributor

Overview

This PR is used to fix the requirement in #28695.

Additional information

It's provided by logari81.
Support long-context decode for qwen4exp on Intel Arc B70.
It's controled by env vars:
GGML_SYCL_SPARSE_FA
GGML_SYCL_SPARSE_FA_DEBUG
GGML_SYCL_SPARSE_FA_MARGIN=256

It's disabled as default.

You need to export GGML_SYCL_SPARSE_FA=1 to enable .

Thank logari81 to contribute the patch!

Requirements

@arthw
arthw requested review from a team and CISC as code owners September 12, 2026 06:16
@github-actions github-actions Bot added documentation Improvements or additions to documentation model Model specific ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language labels Sep 12, 2026
Comment thread docs/backend/SYCL.md Outdated
Comment thread ggml/src/ggml-sycl/ggml-sycl.cpp Outdated
Comment thread ggml/src/ggml-sycl/ggml-sycl.cpp Outdated
GGML_LOG_INFO(" GGML_SYCL_SUPPORT_VMM: no\n");
#endif

//Print the running environment variables for SYCL backend

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

formatting!

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes, update

Comment thread src/models/qwen4exp.cpp Outdated
static const bool sparse_fa = getenv("GGML_SYCL_SPARSE_FA") && atoi(getenv("GGML_SYCL_SPARSE_FA")) != 0;

ggml_tensor * cur = build_attn_mha(q, k, v, nullptr, kq_mask_top_k, nullptr, nullptr,
sparse_fa ? top_k->ne[0] : 0, kq_scale, il);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remove this before merge!

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes, remove them.

@arthw

arthw commented Sep 22, 2026

Copy link
Copy Markdown
Contributor Author

@logari81

Could you help verify this PR?

Thank you!

@logari81

Copy link
Copy Markdown

Just tested the current version of this PR on top of current master and seems to work well.

Until now I have been using the original version of the patch daily with no issues at all.

@arthw

arthw commented Sep 24, 2026

Copy link
Copy Markdown
Contributor Author

OK, it's great!
We will merge it as soon!

Thank you for your feedback and support!

@arthw arthw added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Sep 24, 2026
@ggerganov
ggerganov merged commit cd74ef6 into ggml-org:master Sep 25, 2026
17 checks passed
feal87 added a commit to feal87/myllama.cpp that referenced this pull request Sep 25, 2026
Merge upstream commits:
  - llama: add llama_prec_policy + model-driven W4A4 path (ggml-org#24364)
  - llama: fix tensor split for fused qkv with uneven K/V head sizes (ggml-org#29294)
  - metal: split fa kernels into per-dtype libraries (ggml-org#29329)
  - metal: FWHT kernels for block widths above 512 (ggml-org#29095)
  - CUDA: fuse RMS_NORM + SCALE into one kernel (ggml-org#29393)
  - common: extract shared unicode path/string helpers (ggml-org#29415)
  - common,rpc: simplify fs_create_directory_with_parents() (ggml-org#29432)
  - rpc: include nb in the get_alloc_size cache key (ggml-org#29283)
  - [SYCL] support sparse FA (ggml-org#28796)
  - musa: fix PH1 operator failures and build issues (ggml-org#29193)
  - HIP: bump HIP_VERSION required for fp8 (ggml-org#29231)
  - opencl: add q5_k bin kernel (ggml-org#29401)
  - hexagon: add q5_k quant type support (ggml-org#29123)
  - hexagon: use DMA for contiguous dim1 CONCAT (ggml-org#29404)
  - mtmd: fix mel preprocessor in LFM2 audio (ggml-org#29403)
  - vulkan: fix legacy GLSLC without cooperativeMatrix (ggml-org#29409)
  - gguf-py: ByteLevel processing defaults bos/eos to False (ggml-org#29422)
  - gguf-py: TemplateProcessing has final word on add_special_token (ggml-org#29417)

Assisted-by: Pi
sky-mighty pushed a commit to sky-mighty/llama.cpp that referenced this pull request Sep 26, 2026
* fix conflict

* fix format issue

* rm unused code
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
* fix conflict

* fix format issue

* rm unused code
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
* fix conflict

* fix format issue

* rm unused code

(cherry picked from commit cd74ef6)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. model Model specific SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants