Skip to content

Support superset metric for control_instrumentation - #129

Open
garimadhaked wants to merge 6 commits into
Xilinx:masterfrom
garimadhaked:dtrace_hooks_run_start
Open

garimadhaked wants to merge 6 commits into
Xilinx:masterfrom
garimadhaked:dtrace_hooks_run_start

Conversation

@garimadhaked

@garimadhaked garimadhaked commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Specs and implementation details :
Summary

  • Generate VE2 dynamic-trace control-trace files once per kernel at run construction, where the ELF and its SAVE_TIMESTAMPS locations are available. A single metric set is programmed there; run start does no further selection.

  • Let one design run profile a sequence of inferences. control_instrumentation.profile_runs applies entry N to the Nth profiled inference, start_inference skips leading inferences, and the compute_io_bound shorthand expands to that sequence. The matching control-trace file is selected at run start, so a reused run object or a pool of run objects both follow execution order.

  • Name the files aie_dtrace_ctx_inf.ct. The inference number stays absolute and 1-based. This assumes one kernel per hardware context.

Test plan

  • Single-configuration control_instrumentation (no profile_runs) still profiles every inference with one control-trace file programmed at run construction.

  • An explicit profile_runs array applies a different metric set to each inference, and start_inference leaves earlier inferences unprofiled.

  • "profile_runs": "compute_io_bound" expands to the four-inference sequence.

  • Reusing one run object and creating several run objects up front both walk the sequence in execution order.

  • Control-trace files are named aie_dtrace_ctx_inf.ct and are built once per kernel.

garimadhaked and others added 3 commits September 22, 2026 03:55
In the full ELF flow the hardware context is not fully configured when
run_impl is constructed, so CT generation is moved to the run_start hook
where the ELF and partition data are reliably available.

The hook receives only the run pointer, so the ELF is resolved in the
plugin via get_xdp_kernel_data() and module_int::get_elf_handle() rather
than widening the shared run-lifecycle hook signature.

run_start fires on every submission of a run object whereas the
constructor fired once, so the VE2 implementation now tracks handled run
uids to keep CT generation and dtrace buffer setup once per run.

Signed-off-by: Garima Dhaked <garima.dhaked@amd.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
…start

Producing an IO/compute boundness report needed a separate design run per
metric set. Carry the whole set in one run instead: control_instrumentation
gains "profile_runs", and the Nth profiled inference uses the Nth entry.
A "start_inference" key skips leading inferences, and the super-metric-set
shorthand expands to the same sequence.

CT generation moves back to the run constructor, where the ELF supplying the
SAVE_TIMESTAMPS locations is available, and is memoized per kernel so the
files are built once instead of once per run object.

When a single metric set covers every inference there is nothing to select,
so its CT is programmed at construction as well and the start hook does no
work at all. Only a multi-inference sequence counts inferences at start and
applies the matching CT, which is what lets one reused run object walk the
sequence. The counter lives at start rather than at construction because an
application is free to reuse a single run object or to build a pool up front,
and in neither case does construction order track inference order.

Also makes the plugin safe when runs are created and started concurrently:
the implementation map is guarded and handed out as a shared owner so
teardown cannot free it mid-hook, per-kernel state is serialized, and
configuredOnePartition is atomic.

Co-authored-by: Cursor <cursoragent@cursor.com>
The CT name carried the kernel name so that two kernels sharing a hardware
context would not overwrite each other's files. Nothing consumes that part of
the name, and it makes the files awkward to refer to, so drop it: the files are
now aie_dtrace_ctx_<slot>_inf_<n>.ct. The inference number stays absolute and
1-based, matching start_inference and the log messages.

This assumes a single kernel per hardware context. Two kernels in one context
would now write the same path.

sanitizeForFilename existed only to make the kernel name safe for this filename
and goes with it, along with the cctype include it needed.

Co-authored-by: Cursor <cursoragent@cursor.com>
garimadhaked and others added 3 commits October 7, 2026 00:55
Keep per-inference MetricSelection control-trace generation and carry master's mem tile metric sets (output_channels_details, mm2s_channels_details) on that selection.

Co-authored-by: Cursor <cursoragent@cursor.com>
profile_runs entries after the first can request output_channels_details
while the shared config map still describes inference 1, which has no mem
tile columns. An empty column list means every partition column, so those
inferences get a CT instead of being skipped.

Co-authored-by: Cursor <cursoragent@cursor.com>
The single-metric path and a profile_runs sequence both number the first inference as 0, so the control-trace filename and the run-start selection stay aligned.

Co-authored-by: Cursor <cursoragent@cursor.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants