Skip to content

Pull requests: NVIDIA/TransformerEngine

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

[PyTorch] Use torch's register_custom_class API for opaque quantizers when available community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3495 opened Sep 8, 2026 by mmarcinkiewicz Contributor Loading…
2 of 13 tasks
[Docs] Add Mixture of Experts guide documentation Improvements or additions to documentation
#3494 opened Sep 7, 2026 by pggPL Collaborator Loading…
[PyT] Disable FA3 for training when head_dim_qk != head_dim_v community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3490 opened Sep 7, 2026 by yuweih205 Loading…
6 tasks done
Add native transport for dynamic context parallelism community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3489 opened Sep 6, 2026 by xiaoyao0115 Draft
Support row-only MXFP8 distributed master-weight casts community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3488 opened Sep 6, 2026 by xiuhu17 Contributor Loading…
5 tasks done
fix(jax): interpolate the tensor-sequence-parallelism warning community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3486 opened Sep 5, 2026 by Anai-Guo Loading…
[PyTorch] Fix NaN expert weight gradients at num_groups == 1 with SReLU community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3485 opened Sep 5, 2026 by GarlGuo Loading…
6 of 13 tasks
[Common] Fix Grouped MXFP8 work mapping and TMA synchronization
#3483 opened Sep 4, 2026 by Oleg-Goncharov Collaborator Loading…
3 of 13 tasks
[PyTorch] Split FusedAttnFunc into single-argument forward/backward helpers 2.20
#3480 opened Sep 4, 2026 by pggPL Collaborator Loading…
6 of 13 tasks
Drop the unsupported window_size argument from attention backend queries community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3479 opened Sep 4, 2026 by Anai-Guo Loading…
[PyT] Linear Attention API 2.20
#3477 opened Sep 4, 2026 by KshitijLakhani Collaborator Draft
13 tasks
[PyTorch] Enable fused activation recompute for ScaledTanhSReLU community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3473 opened Sep 3, 2026 by wanyingw Contributor Draft
13 tasks
[PyTorch] torch.compile support for FusedAttention
#3472 opened Sep 3, 2026 by pggPL Collaborator Draft
7 of 13 tasks
[PyTorch] DeepSeekV3Layer: full MoE transformer layer (MLA + DeepSeek MoE)
#3471 opened Sep 3, 2026 by pggPL Collaborator Draft
7 of 13 tasks
[All] Guard THD learnable dSink on older cuDNN attention bug Something isn't working
#3470 opened Sep 3, 2026 by KshitijLakhani Collaborator Loading…
2 of 13 tasks
[PyTorch] Fix FP8 illegal memory access in single-process multi-GPU execution community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3469 opened Sep 3, 2026 by SuperGoodGame Loading…
Automatically omit unused columnwise primary weights for backward overrides community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3468 opened Sep 3, 2026 by xiuhu17 Contributor Loading…
[Common] row-scaled nvfp4 path: add single-launch group fused amax community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3467 opened Sep 3, 2026 by cael-ling Contributor Loading…
2 of 13 tasks
[JAX] Make MoEBlock aware of Dense TP axes
#3462 opened Sep 2, 2026 by jberchtold-nvidia Collaborator Draft
8 of 13 tasks
[Common] Group NVFP4 Quantize Kernels
#3458 opened Sep 1, 2026 by Oleg-Goncharov Collaborator Loading…
9 of 13 tasks
[JAX] Add SiTU-GLU support to JAX
#3457 opened Sep 1, 2026 by jberchtold-nvidia Collaborator Draft
8 of 13 tasks
[PyTorch] Schedule delayed-scaling updates after backward
#3456 opened Sep 1, 2026 by pggPL Collaborator Loading…
9 of 13 tasks
[Common] row-scaled nvfp4 path: fuse row/col amax into a single TMA-tiled kernel community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3454 opened Sep 1, 2026 by cael-ling Contributor Loading…
2 of 13 tasks
ProTip! Find all pull requests that aren't related to any open issues with -linked:issue.