Skip to content

[Conv] Fix filter iterator pointer offset scaling (issue #3504) - #3623

Open
XFDG wants to merge 1 commit into
NVIDIA:mainfrom
XFDG:fix/conv-filter-pointer-offset-3504
Open

XFDG wants to merge 1 commit into
NVIDIA:mainfrom
XFDG:fix/conv-filter-pointer-offset-3504

Conversation

@XFDG

@XFDG XFDG commented Sep 13, 2026

Copy link
Copy Markdown

Summary

Fix add_pointer_offset byte scaling in five convolution filter tile access iterators:

  • Conv2d fprop analytic
  • Conv2d fprop fixed channels
  • Conv2d fprop few channels
  • Conv3d fprop analytic
  • Depthwise Conv2d fprop direct-convolution optimized

Add a lightweight convolution threadblock unit test that constructs each iterator and verifies its observable pointer displacement.

Rationale

These iterators store their base pointer as char const*, while add_pointer_offset documents its argument in units of Element. The byte displacement must therefore be:

pointer_offset * sizeof_bits::value / 8

The previous inverted expression multiplied by 8 and divided by the element bit width. For fp16, advancing by 8 elements moved 4 bytes instead of 16; fp32 was off by a factor of 16.

Validation

Built the same host/device CUDA probe before and after the change using CUDA 13.1 and -arch=sm_100a, then ran it on an NVIDIA B200.

Before:

  • all five fp16 iterators moved 4 bytes for an 8-element offset on both host and device
  • 10/10 checks failed

After:

  • all five iterators move the expected 16 bytes on host and device
  • 10/10 checks pass and the probe exits with failures=0
  • the new WITHOUT_CUDA unit-test translation unit compiles successfully with the GPU node host compiler
  • git diff --check passes

Fixes #3504

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Five conv filter tile access iterators invert the byte scaling in add_pointer_offset (8/sizeof_bits instead of sizeof_bits/8)

1 participant