Exact integer matrix products on FP4 Tensor Cores, and the BLAS routines built on them.
Warning
Only GEMM is published so far, and the optimized implementations are not published yet. The code here is the working reference.
Adaptive Weight Encoding (AWE) splits an integer into limbs taken from the set of values FP4 can
store, with freely chosen integer weights, and multiplies linear combinations of limbs. The exact
product needs fewer FP4 GEMMs than the base-13 split of prior work. AWE/Solutions/ is the
catalog of these encodings. Details are in the paper, arXiv:2609.24519.
| Family | Inputs | FP4 GEMMs for one exact product | Ratio | |
|---|---|---|---|---|
| Base-13 limbs (prior work) | AWE | |||
| Direct | INT8 × INT8 | 9 | 6 | 1.5× |
| INT4 × INT8 | 6 | 4 | 1.5× | |
| FP16 significand | 16 | 12 | 1.3× | |
| Residue systems (K = 16,384) | FP64 significand | 75 | 59 | 1.3× |
The radix-13 limb representation itself is published in FP4-is-All-you-Need/Oz-FP4.
The paper's example of INT8 × INT8 in 6 FP4 GEMMs, limb weights (1, 29, 37). The INT8 path of the GEMM library uses the catalog solution with weights (1, 5, 27).
| Path | Content |
|---|---|
AWE/Solutions/ |
Encoding solutions as data, one JSON line per solution. |
AWE/CodeGen/ |
The generator that turns a solution into C constants. |
Platforms/NVIDIA/CUDA/GEMM/ |
GEMM library: INT8 (AWE and radix-13), FP64 (residue systems) and complex FP64 backends, with examples and tests. |
MIT; see LICENSE.
@misc{hayashi2026awe,
title = {AWE: Adaptive Weight Encoding for Exact Integer Matrix Products with Fewer GEMMs on FP4 Tensor Cores},
author = {Hayashi, Shun{-}ichiro and Mukunoki, Daichi and Hoshino, Tetsuya and Katagiri, Takahiro},
year = {2026},
eprint = {2609.24519},
archivePrefix = {arXiv},
primaryClass = {cs.MS},
url = {https://arxiv.org/abs/2609.24519}
}