Skip to content

llama-context : report graph inputs and input tensors during sched reserve - #26625

Merged
ggerganov merged 7 commits into
masterfrom
gg/llama-context-report-graph-inputs
Sep 21, 2026
Merged

ggerganov merged 7 commits into
masterfrom
gg/llama-context-report-graph-inputs

Conversation

@ggerganov

Copy link
Copy Markdown
Member

Overview

Adds reporting of graph inputs and input tensors to the sched_reserve log output in llama-context.cpp, and fixes the batch-size label used for the token-generation graph.

Specifically:

  • The bs label for the token-generation graph is now reported as n_seqs instead of a hardcoded 1, matching how the pp graph reports n_tokens.
  • Reports the number of graph inputs (from llm_graph_result::inputs) for both the pp (prompt processing) and tg (token generation) graphs.
  • Reports the number of input tensors: the distinct tensors in the graph (nodes and their src tensors) flagged with GGML_TENSOR_FLAG_INPUT.
  • Logs a warning when an input tensor's op is not GGML_OP_NONE.
  • Logs a trace (debug) line for each input tensor listing the nodes (name and op) that use it.

Additional information

The graph-input tensors are identified by iterating over the graph nodes and their src tensors, tracking unique tensors via a hash map so that a tensor shared by multiple nodes is counted once.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. llama.cpp:DeepSeek-v4-Flash-0731

@ggerganov
ggerganov force-pushed the gg/llama-context-report-graph-inputs branch 2 times, most recently from 24496cb to 8e9ed92 Compare August 12, 2026 11:12
@ggerganov
ggerganov force-pushed the gg/llama-context-report-graph-inputs branch from 2ab78db to 59bf83c Compare September 15, 2026 12:23
…serve

- fix the tg (token generation) graph bs label to use n_seqs instead of a hardcoded 1
- report the number of graph inputs from llm_graph_result::inputs for both the pp and tg graphs
- report the number of input tensors (nodes and their src tensors flagged with GGML_TENSOR_FLAG_INPUT)
- log a warning when an input tensor has an op other than GGML_OP_NONE
- log a trace line for each input tensor and the nodes (name and op) that use it

Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731
- name the kv-cache idxs input tensors (attn_inp_k_idxs, attn_inp_v_idxs)
- name the recurrent state copy idxs input tensor (rs_s_copy)
- report the input tensor shape in the sched_reserve trace

Assisted-by: pi:llama.cpp/Qwen3.8-27B
- print nodes, splits, input objects and input tensors in one line
- when the pp and tg graphs differ, print each value as 'pp / tg'
  and annotate the line with the batch sizes used for each graph

Assisted-by: pi:llama.cpp/Qwen3.8-27B
@ggerganov
ggerganov force-pushed the gg/llama-context-report-graph-inputs branch from f8bbd55 to e94c26d Compare September 21, 2026 11:15
@ggerganov
ggerganov marked this pull request as ready for review September 21, 2026 11:16
@ggerganov
ggerganov requested a review from CISC as a code owner September 21, 2026 11:16
@ggerganov
ggerganov merged commit 9655061 into master Sep 21, 2026
17 checks passed
@ggerganov
ggerganov deleted the gg/llama-context-report-graph-inputs branch September 21, 2026 16:13
LadislavSopko pushed a commit to 0ics-srls/llama.cpp that referenced this pull request Oct 5, 2026
…serve (ggml-org#26625)

* llama-context : report graph inputs and input tensors during sched reserve

- fix the tg (token generation) graph bs label to use n_seqs instead of a hardcoded 1
- report the number of graph inputs from llm_graph_result::inputs for both the pp and tg graphs
- report the number of input tensors (nodes and their src tensors flagged with GGML_TENSOR_FLAG_INPUT)
- log a warning when an input tensor has an op other than GGML_OP_NONE
- log a trace line for each input tensor and the nodes (name and op) that use it

Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731

* cont : count input tensors before reserving the sched

* wip

* llama-graph : name the unnamed graph input tensors

- name the kv-cache idxs input tensors (attn_inp_k_idxs, attn_inp_v_idxs)
- name the recurrent state copy idxs input tensor (rs_s_copy)
- report the input tensor shape in the sched_reserve trace

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* llama-context : rename "graph inputs" to "graph input objects"

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* llama-context : report the sched reserve graph stats on a single line

- print nodes, splits, input objects and input tensors in one line
- when the pp and tg graphs differ, print each value as 'pp / tg'
  and annotate the line with the batch sizes used for each graph

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* cont : pad logs
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
…serve (ggml-org#26625)

* llama-context : report graph inputs and input tensors during sched reserve

- fix the tg (token generation) graph bs label to use n_seqs instead of a hardcoded 1
- report the number of graph inputs from llm_graph_result::inputs for both the pp and tg graphs
- report the number of input tensors (nodes and their src tensors flagged with GGML_TENSOR_FLAG_INPUT)
- log a warning when an input tensor has an op other than GGML_OP_NONE
- log a trace line for each input tensor and the nodes (name and op) that use it

Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731

* cont : count input tensors before reserving the sched

* wip

* llama-graph : name the unnamed graph input tensors

- name the kv-cache idxs input tensors (attn_inp_k_idxs, attn_inp_v_idxs)
- name the recurrent state copy idxs input tensor (rs_s_copy)
- report the input tensor shape in the sched_reserve trace

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* llama-context : rename "graph inputs" to "graph input objects"

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* llama-context : report the sched reserve graph stats on a single line

- print nodes, splits, input objects and input tensors in one line
- when the pp and tg graphs differ, print each value as 'pp / tg'
  and annotate the line with the batch sizes used for each graph

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* cont : pad logs
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
…serve (ggml-org#26625)

* llama-context : report graph inputs and input tensors during sched reserve

- fix the tg (token generation) graph bs label to use n_seqs instead of a hardcoded 1
- report the number of graph inputs from llm_graph_result::inputs for both the pp and tg graphs
- report the number of input tensors (nodes and their src tensors flagged with GGML_TENSOR_FLAG_INPUT)
- log a warning when an input tensor has an op other than GGML_OP_NONE
- log a trace line for each input tensor and the nodes (name and op) that use it

Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731

* cont : count input tensors before reserving the sched

* wip

* llama-graph : name the unnamed graph input tensors

- name the kv-cache idxs input tensors (attn_inp_k_idxs, attn_inp_v_idxs)
- name the recurrent state copy idxs input tensor (rs_s_copy)
- report the input tensor shape in the sched_reserve trace

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* llama-context : rename "graph inputs" to "graph input objects"

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* llama-context : report the sched reserve graph stats on a single line

- print nodes, splits, input objects and input tensors in one line
- when the pp and tg graphs differ, print each value as 'pp / tg'
  and annotate the line with the batch sizes used for each graph

Assisted-by: pi:llama.cpp/Qwen3.8-27B

* cont : pad logs

(cherry picked from commit 9655061)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant