Skip to content

[Feature]: Quantized version of Qwen3.8-27B #769

Description

@Skurios

Since you recently added qwen3.8-mtp:27b, I was wondering whether it would be possible to also provide a quantized version of the model.

The 27B model looks very interesting, especially with the new MTP speculative decoding support. However, I think its memory requirements may be too high for many of the laptops that make up a large part of the FastFlowLM user base.

A lot of Ryzen AI laptops come with only 16–32 GB of system memory, so running the full 27B model alongside the OS and other applications can be difficult or even impossible.

Would it be possible to provide an officially supported quantized version of qwen3.8-mtp:27b, for example a 4-bit variant?

This could make the model accessible to significantly more FastFlowLM users while still allowing us to benefit from the MTP speculative decoding support.

Operating System

windows 11

GPU

AI 9 HX 370 (32 GB)

ROCm Component

Activity

  1. sunjizheng commented on Oct 6, 2026

    @sunjizheng

    The model provided now is the q4 quantized version.

    You need to modify the VRAM size to free up space for the model.

  2. Skurios commented on Oct 6, 2026

    @Skurios
    Author

    You're right, I misremembered that. The model download is about 16,305 MB, so it makes sense that it's already a 4-bit quantized version.

    The problem is that Windows 11 unfortunately limits the shared memory available to the iGPU/NPU to around 50% of the total system RAM, and as far as I know this can't be increased. So on a 32 GB laptop, I can't get the model to load even though the system technically has enough total RAM.

    It would be great if there could be a slightly smaller quantization/version that just fits within the ~16 GB shared-memory limit. I think that would make the model usable on a lot more 32 GB Ryzen AI laptops.

  3. Skurios commented on Oct 6, 2026

    @Skurios
    Author
    Image

    Just to clarify the issue a bit more: with other local LLM software, I can run models larger than 16 GB on my 32 GB system because they can also make use of regular system memory.

    With FastFlowLM/NPU inference, however, the model seems to need to fit entirely into the memory available to the NPU. Since Windows limits this shared memory to roughly 50% of the installed RAM, a 32 GB system effectively has only ~16 GB available — which means the 16.3 GB model doesn't fit once you include the additional memory overhead.

    So unless I'm missing something, I don't think anyone with 32 GB of RAM can currently run this model, which is a bit unfortunate.

    I'm not sure if it's technically possible, but could FastFlowLM use regular system memory for part of the model instead of requiring the whole model to reside in the NPU-accessible memory? Since it's physically the same unified system RAM on these APUs, I was wondering whether some kind of offloading could be possible.

  4. sunjizheng commented on Oct 7, 2026

    @sunjizheng

    It looks like your graphics card is using a lot of memory. You can set it to use less memory in the BIOS when you start up your computer.

  5. Skurios commented on Oct 8, 2026

    @Skurios
    Author

    It looks like your graphics card is using a lot of memory. You can set it to use less memory in the BIOS when you start up your computer.

    Mhh? I was loading up the model..?
    I dont understand.
    Shrared memory is the only memoriy the the NPU can use. Means: when ur system has 32GB the maximum shared would be 16 GB. This is hardcoded in windows. You cant change that. Even if you have enouth ram left, to fit the hole model in, the NPU and iGPU cant load more than 16 GB!!
    Nobody with 32Gb can use the Model 3.8 27B!

    I can run it in LM Studio, becuase while loading a model you can select if the model should "off load on the GPU" --> means: it uses the shared memoey. Everything above shared memory is put into “normal RAM”. Since shared memory and “normal RAM” are essentially the same, you can also set offload to 0% when loading. That makes hardly any difference. But in LM Studio I can’t use the NPU. It only runs on CUDA/llama and Vulkan/CPU. I think it would be really awesome if the developer manages to make sure the model doesn’t use shared memory, or if models under 16 GB are released.

  6. Fir3PL commented on Oct 10, 2026

    @Fir3PL

    3. With FastFlowLM/NPU inference, however, the model seems to need to fit entirely into the memory available to the NPU. Since Windows limits this shared memory to roughly 50% of the installed RAM, a 32 GB system effectively has only ~16 GB available — which means the 16.3 GB model doesn't fit once you include the additional memory overhead.

    I run this on Ubuntu on AI 7 350 24 gb RAM, but i have max 3-4 t/s. On linux you can use all free memory, not just half of it, as you can on Windows.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions