Skip to content
#

multi-turboquant

Here are 2 public repositories matching this topic...

Native Windows vLLM: stable 0.27.1; 0.29.0 prereleases for Python 3.13/CUDA 13.0 and Python 3.14/CUDA 13.2 (CPU TorchAudio). PyTorch 2.13, FlashAttention/Rust, OpenAI-compatible serving, 10 KV formats, Multi-TurboQuant and experimental CPU/NVMe KV offload. No WSL or Docker.

  • Updated Sep 20, 2026
  • Python

Add this topic to your repo

To associate your repository with the multi-turboquant topic, visit your repo's landing page and select "manage topics."

Learn more