Skip to content

cuda: GGML_CUDA_OP_TIMING=1 per-node GPU time breakdown - #323

Open
professorpalmer wants to merge 1 commit into
PrismML-Eng:prismfrom
professorpalmer:claude/op-timing
Open

professorpalmer wants to merge 1 commit into
PrismML-Eng:prismfrom
professorpalmer:claude/op-timing

Conversation

@professorpalmer

Copy link
Copy Markdown

One of the five pieces of #285, split per bri-prism's request on #317. Single commit on current prism (eaecb50c7), applies without conflicts. Diagnostics only, off by default; happy to drop it if you would rather not carry it.

What it does

GGML_CUDA_OP_TIMING=1 (with CUDA graphs off, GGML_CUDA_DISABLE_GRAPHS=1): records an event before every graph node on the main stream, charges the time to the next event to that node (a fused group to its first node), aggregates by op / type / shape and prints the top 30 every 100 graphs through the ggml log (INFO level, so -lv 5 in llama-server). Launch gaps inflate the small ops; the large mat-vecs and GEMMs read true.

Why it is useful

It is how the per-op cost of a decode step and of a prefill micro-batch was attributed while tuning the 262k serve, and how a recent prefill regression on this card was traced to a demoted buffer rather than a kernel in a few minutes.

With CUDA graphs off (GGML_CUDA_DISABLE_GRAPHS=1), records an event before every graph node on the main
stream, charges the time to the next event to that node (a fused group to its first node), aggregates by
op/type/shape and prints the top 30 every 100 graphs. Launch gaps inflate the small ops; the large mat-vecs
read true. Diagnostics only; off by default.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 3164035)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant