vllm.models.qwen4_exp.nvidia.ops ¶
NVIDIA-only Qwen4Exp kernels; import leaf modules directly.
Modules:
-
hc–NVIDIA HyperConnection kernels for Qwen4Exp.
-
ple–Fused Qwen4Exp PLE kernels.
-
qsa–Triton kernels for Qwen4Exp QSA sparse attention and cache updates.
-
qsa_indexer–Triton kernels for Qwen4Exp QSA index selection.
-
qsa_pre_indexer–Fused QSA pre-indexer kernel for Qwen4Exp.