Repository navigation
llama-context : report graph inputs and input tensors during sched reserve - #26625
Merged
Merged
Conversation
ggerganov
force-pushed
the
gg/llama-context-report-graph-inputs
branch
2 times, most recently
from
August 12, 2026 11:12
24496cb to
8e9ed92
Compare
ggerganov
force-pushed
the
gg/llama-context-report-graph-inputs
branch
from
September 15, 2026 12:23
2ab78db to
59bf83c
Compare
…serve - fix the tg (token generation) graph bs label to use n_seqs instead of a hardcoded 1 - report the number of graph inputs from llm_graph_result::inputs for both the pp and tg graphs - report the number of input tensors (nodes and their src tensors flagged with GGML_TENSOR_FLAG_INPUT) - log a warning when an input tensor has an op other than GGML_OP_NONE - log a trace line for each input tensor and the nodes (name and op) that use it Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731
- name the kv-cache idxs input tensors (attn_inp_k_idxs, attn_inp_v_idxs) - name the recurrent state copy idxs input tensor (rs_s_copy) - report the input tensor shape in the sched_reserve trace Assisted-by: pi:llama.cpp/Qwen3.8-27B
Assisted-by: pi:llama.cpp/Qwen3.8-27B
- print nodes, splits, input objects and input tensors in one line - when the pp and tg graphs differ, print each value as 'pp / tg' and annotate the line with the batch sizes used for each graph Assisted-by: pi:llama.cpp/Qwen3.8-27B
ggerganov
force-pushed
the
gg/llama-context-report-graph-inputs
branch
from
September 21, 2026 11:15
f8bbd55 to
e94c26d
Compare
ggerganov
marked this pull request as ready for review
September 21, 2026 11:16
This was referenced Sep 23, 2026
LadislavSopko
pushed a commit
to 0ics-srls/llama.cpp
that referenced
this pull request
Oct 5, 2026
…serve (ggml-org#26625) * llama-context : report graph inputs and input tensors during sched reserve - fix the tg (token generation) graph bs label to use n_seqs instead of a hardcoded 1 - report the number of graph inputs from llm_graph_result::inputs for both the pp and tg graphs - report the number of input tensors (nodes and their src tensors flagged with GGML_TENSOR_FLAG_INPUT) - log a warning when an input tensor has an op other than GGML_OP_NONE - log a trace line for each input tensor and the nodes (name and op) that use it Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731 * cont : count input tensors before reserving the sched * wip * llama-graph : name the unnamed graph input tensors - name the kv-cache idxs input tensors (attn_inp_k_idxs, attn_inp_v_idxs) - name the recurrent state copy idxs input tensor (rs_s_copy) - report the input tensor shape in the sched_reserve trace Assisted-by: pi:llama.cpp/Qwen3.8-27B * llama-context : rename "graph inputs" to "graph input objects" Assisted-by: pi:llama.cpp/Qwen3.8-27B * llama-context : report the sched reserve graph stats on a single line - print nodes, splits, input objects and input tensors in one line - when the pp and tg graphs differ, print each value as 'pp / tg' and annotate the line with the batch sizes used for each graph Assisted-by: pi:llama.cpp/Qwen3.8-27B * cont : pad logs
frostyautumnleaf
pushed a commit
to frostyautumnleaf/llama.cpp
that referenced
this pull request
Oct 5, 2026
…serve (ggml-org#26625) * llama-context : report graph inputs and input tensors during sched reserve - fix the tg (token generation) graph bs label to use n_seqs instead of a hardcoded 1 - report the number of graph inputs from llm_graph_result::inputs for both the pp and tg graphs - report the number of input tensors (nodes and their src tensors flagged with GGML_TENSOR_FLAG_INPUT) - log a warning when an input tensor has an op other than GGML_OP_NONE - log a trace line for each input tensor and the nodes (name and op) that use it Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731 * cont : count input tensors before reserving the sched * wip * llama-graph : name the unnamed graph input tensors - name the kv-cache idxs input tensors (attn_inp_k_idxs, attn_inp_v_idxs) - name the recurrent state copy idxs input tensor (rs_s_copy) - report the input tensor shape in the sched_reserve trace Assisted-by: pi:llama.cpp/Qwen3.8-27B * llama-context : rename "graph inputs" to "graph input objects" Assisted-by: pi:llama.cpp/Qwen3.8-27B * llama-context : report the sched reserve graph stats on a single line - print nodes, splits, input objects and input tensors in one line - when the pp and tg graphs differ, print each value as 'pp / tg' and annotate the line with the batch sizes used for each graph Assisted-by: pi:llama.cpp/Qwen3.8-27B * cont : pad logs
edwardyoon
pushed a commit
to edwardyoon/focus-llama
that referenced
this pull request
Oct 7, 2026
…serve (ggml-org#26625) * llama-context : report graph inputs and input tensors during sched reserve - fix the tg (token generation) graph bs label to use n_seqs instead of a hardcoded 1 - report the number of graph inputs from llm_graph_result::inputs for both the pp and tg graphs - report the number of input tensors (nodes and their src tensors flagged with GGML_TENSOR_FLAG_INPUT) - log a warning when an input tensor has an op other than GGML_OP_NONE - log a trace line for each input tensor and the nodes (name and op) that use it Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731 * cont : count input tensors before reserving the sched * wip * llama-graph : name the unnamed graph input tensors - name the kv-cache idxs input tensors (attn_inp_k_idxs, attn_inp_v_idxs) - name the recurrent state copy idxs input tensor (rs_s_copy) - report the input tensor shape in the sched_reserve trace Assisted-by: pi:llama.cpp/Qwen3.8-27B * llama-context : rename "graph inputs" to "graph input objects" Assisted-by: pi:llama.cpp/Qwen3.8-27B * llama-context : report the sched reserve graph stats on a single line - print nodes, splits, input objects and input tensors in one line - when the pp and tg graphs differ, print each value as 'pp / tg' and annotate the line with the batch sizes used for each graph Assisted-by: pi:llama.cpp/Qwen3.8-27B * cont : pad logs (cherry picked from commit 9655061)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Adds reporting of graph inputs and input tensors to the
sched_reservelog output inllama-context.cpp, and fixes the batch-size label used for the token-generation graph.Specifically:
bslabel for the token-generation graph is now reported asn_seqsinstead of a hardcoded1, matching how the pp graph reportsn_tokens.llm_graph_result::inputs) for both the pp (prompt processing) and tg (token generation) graphs.srctensors) flagged withGGML_TENSOR_FLAG_INPUT.GGML_OP_NONE.Additional information
The graph-input tensors are identified by iterating over the graph nodes and their
srctensors, tracking unique tensors via a hash map so that a tensor shared by multiple nodes is counted once.Requirements