Check platform and hardware support for GPTQModel
mainGPTQModel supports a wide range of platforms and hardware architectures, including NVIDIA GPUs, AMD GPUs, Huawei Ascend NPUs, Intel XPUs, and Apple Silicon. Support varies by device, optimized architecture, and available kernels.
Supported Platforms and Devices
| Platform | Device | Optimized Arch | Kernels |
|---|---|---|---|
| 🐧 Linux | NVIDIA GPU | Turing+ (sm_75+) | Machete, Marlin, Exllama V3 / EXL3, Exllama V2, AWQ GEMM/GEMV, ParoQuant CUDA/Triton, GGUF CUDA/Triton, QQQ, BitBLAS, Triton, BitsAndBytes, Torch |
| 🐧 Linux | AMD GPU | 7900XT+, ROCm 6.2+ | Exllama V2, AWQ GEMM/GEMV, QQQ, FP8 Torch, Torch |
| 🐧 Linux | Huawei Ascend NPU | Ascend 910B, torch-npu / CANN | Native Torch kernels for GPTQ, AWQ, ParoQuant, GGUF, QQQ, and EXL3 |
| 🐧 Linux | Intel XPU | Arc, Datacenter Max | TorchFused, TorchFusedAWQ, FP8 Torch, Torch |
| 🐧 Linux | Intel/AMD CPU | avx, amx | TorchFused, TorchFusedAWQ, TorchAten int4, TorchInt8, GGUF C++, BitsAndBytes, Torch |
| 🍎 macOS | GPU (Metal) / CPU | Apple Silicon, M1+ | Torch, FP8 Torch, MLX via conversion |
| 🪟 Windows | GPU (NVIDIA) / CPU | NVIDIA | Torch |
Key Notes
- NVIDIA GPUs:
Marlinand JIT CUDA kernels supportTuring+(sm_75+) architectures. - Huawei Ascend NPU: Uses native Torch kernels via
torch-npu/CANN. - macOS: Supports Apple Silicon (
M1+) via Torch, FP8 Torch, or MLX (via conversion).