Skip to content

Feature Request: Speculative Prefill #19082

Description

@endyjasmi

Prerequisites

  • I am running the latest code. Mention the version if possible as well.
  • I carefully followed the README.md.
  • I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
  • I reviewed the Discussions, and have a new and useful enhancement to share.

Feature Description

Basically use to improve TTFT. This is very important especially for Ryzen AI Max+ 395 where the prefill stage is slow and unsuitable to be used in Agentic workflow. Hope it can be implemented with

Motivation

This is very important especially for Ryzen AI Max+ 395 where the prefill stage is slow and unsuitable to be used in Agentic workflow. Hope it can be implemented with GLM 4.7 flash

Possible Implementation

https://github.com/[Jingyu6/speculative_prefill](https://github.com/Jingyu6/speculative_prefill)

Activity

  1. github-actions commented on Mar 12, 2026

    @github-actions
    Contributor

    This issue was closed because it has been inactive for 14 days since being marked as stale.

  2. cmp-nct commented on May 27, 2026

    @cmp-nct
    Contributor

    Was closed as stale - would be a huge help with Qwen 27B.
    I don't know about the quality regression when using this method but imho should be on the roadmap

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions