High-performance implementations of AI kernels for CPUs (x64, AArch64, RISC-V) and Intel GPUs. Powers PyTorch, TensorFlow, OpenVINO, and ONNX Runtime.
-
Updated
Oct 4, 2026 - C++
High-performance implementations of AI kernels for CPUs (x64, AArch64, RISC-V) and Intel GPUs. Powers PyTorch, TensorFlow, OpenVINO, and ONNX Runtime.
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
LLM fine-tuning with LoRA + NVFP4/MXFP8 on NVIDIA DGX Spark (Blackwell GB10)
Patches + recipe to deploy festr2/MiMo-V2.5-Pro-NVFP4-MXFP8-attn-TP8 on 8-node DGX Spark sm_121 (Ray + vLLM, TP=8). Fixes the fused-qkv loader bug that mis-slotted Q values as K/V on 7 of 8 ranks.
ComfyUI custom node for converting diffusion checkpoints to FP8, NVFP4, MXFP8 and INT8 with native per-layer quantization metadata.
A dependency-light, transformers-free toolkit for data-free weight-only quantization (FP8 blockwise, MXFP8, MXFP4), with file-level multi-node support and HF / compressed-tensors output.
Blackwell-optimized llama.cpp bundle with DFlash2, MXFP6/MXFP8/NVFP4, TurboQuant KV, and GPU-resident speculative handoff.
Block-scaled FP8 / FP4 / INT4 tensor primitive with Triton scaled-matmul at FP32 parity on H100. NumPy / PyTorch / MLX / JAX backends.
To associate your repository with the mxfp8 topic, visit your repo's landing page and select "manage topics."