[Common] Experimental CuTeDSL MXFP4 backend#3223
Draft
janekb04 wants to merge 26 commits into
Draft
Conversation
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
for more information, see https://pre-commit.ci Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
for more information, see https://pre-commit.ci Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
for more information, see https://pre-commit.ci Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
for more information, see https://pre-commit.ci Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
…it__.py Signed-off-by: Kaining Zhong <kainingz@nvidia.com>
for more information, see https://pre-commit.ci
Route the optimized 1D NVFP4 quantize-transpose path through the new CuTe DSL backend, falling back to the existing CUDA kernel when the CuTe DSL backend declines to handle the case. Adds the FP4 dtype mapping to the TVM FFI bridge and an empty stub header for the backend. Signed-off-by: Jan Bielak <jbielak@nvidia.com>
Add the C++ and Python scaffolding for dispatching NVFP4 quantize-transpose to the CuTe DSL backend via TVM FFI, mirroring the general structure and file layout of the MXFP8 backend (quantize_mxfp8_cutedsl.cuh and quantize_mxfp8.py). Kernel instantiation parameters and validation are left as todos. Signed-off-by: Jan Bielak <jbielak@nvidia.com>
Populate NVFP4QuantizeConfig with the kernel instantiation parameters (stochastic rounding, fast math, row-scaled NVFP4 and transpose flags) on both the C++ and Python sides, and add the input/output validation mirroring the CUDA implementation in quantize_transpose_nvfp4.cuh. The config values are derived from the quantization config and output tensor and forwarded across the TVM FFI boundary, and the backend is guarded behind FP4_TYPE_SUPPORTED. Signed-off-by: Jan Bielak <jbielak@nvidia.com>
for more information, see https://pre-commit.ci
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
This PR adds a CuTe DSL NVFP4 quantization backend. It is based on the work in #3137.
In this PR only a CuTe DSL version of
quantize_transpose_tuned_1Dis implemented. It handles the case of 1Dbfloat16to NVFP4 quantization, mirroring the capabilities of the original CUDA implementation inquantize_transpose_nvfp4_tuned_1D.cuh.To be rebased once #3137 merges.
Type of change
Changes
Support for the CuTe DSL backend is implemented in a top-down manner, one commit at a time:
quantize.cuh(done)tvm_ffi_bridge.h) and an empty stub header for the backend.quantize_transpose_nvfp4_cutedsl.cuh, mirroringquantize_mxfp8_cutedsl.cuh.quantize_transpose_nvfp4.py, mirroringquantize_mxfp8.py.NVFP4QuantizeConfigwith the kernel instantiation parameters (stochastic rounding, fast math, row-scaled NVFP4 and transpose flags) on both the C++ and Python sides, mirroringquantize_transpose_nvfp4.cuh:1441. The config values are derived from the quantization config and output tensor and forwarded across the TVM FFI boundary.FP4_TYPE_SUPPORTED.Checklist: