Skip to content

vllm.v1.attention.ops.ultraquant

UltraQuant 4-bit KV cache: FP4 E2M1 codes + UE8M0 group-of-32 scales.

Production decode uses the FlyDSL D=256 kernel on gfx950, with Triton unified attention as the fallback. Slot size is slot_size(head_dim) (272 B at D=256). Format helpers live in format; import them from there, not this package root.

Modules:

  • format –

    FP4 codepoint + UE8M0 scale constants for the UltraQuant KV cache.

  • reference –

    PyTorch reference for the UltraQuant KV cache format.

  • triton_dequant –

    Full KV dequant for the UltraQuant cache format.

  • triton_store –

    Triton store kernel for the ultraquant KV cache format.

  • triton_unified_attention –

    Unified Triton fallback for the UltraQuant 4-bit KV-cache format.