The set-and-forget LLM engine for Pascal and Volta: PXQ codec + kernels, auto-tuned per card, measured against llama.cpp, ik_llama.cpp and 1Cat vLLM
-
Updated
Sep 17, 2026 - C++
The set-and-forget LLM engine for Pascal and Volta: PXQ codec + kernels, auto-tuned per card, measured against llama.cpp, ik_llama.cpp and 1Cat vLLM
Sol-Attn (arXiv 2607.24027) sparse acceleration for MiniMax H3 on V100 — single-node ComfyUI plugin (embedded FP16Safe + keep-or-drop kernel)
Benchmarks and notes for running modern LLMs with vLLM on 8x Tesla V100-32GB in 2026.
ComfyUI custom node: run MiniMax H3 at near-fp16 speed with near-fp32 numerical stability on GPUs without bf16/fp8 hardware (V100 sm_70)
FlashAttention brought back to Tesla V100 — a deep llama.cpp fork: SM 7.0 D256 kernels, SplitKV3, q4_0 KV cache, DFlash2 speculative decoding and multimodal fixes.
Qwen3.8-27B in native NVFP4/FP8 on 2x PCIe Tesla V100-32GB (SM70): the PCIe runbook for v100-skinny + 1Cat-vLLM, with the 3 fixes that make it work without NVLink. 61-74 tok/s decode, MTP speculative decoding, OpenAI-compatible.
Multi-GPU acceleration for MiniMax H3 video generation on NVIDIA V100 (sm_70). Ulysses sequence parallelism as a drop-in ComfyUI custom node — ~19 min to ~7 min on 8x V100.
Found out that using A100 and V100 on Vicuna and Llama2 have a different result, while other model such as Falcon doesn't has such question.
Serve Qwen3.8-Flash-Next (125B MoE, NVFP4) at 262K context on 4x Tesla V100-SXM2-32GB — a Volta (sm70) port of SGLang for agentic coding. Native OpenAI + Anthropic APIs.
Performance of CUDA example benchmark code on NVIDIA A100.
Serve Qwen3.5-397B-A17B (AWQ) on 8x Tesla V100-SXM2-32GB (DGX-1, TP8) for agentic coding & ops — a downstream fork of 1Cat-vLLM.
SunFire V100 10 Inch Compact Case
Manager your private llm server from your desktop
VastLLM: a production-oriented FastLLM fork for native C++ inference, V100/SM70, long context, and Qwen3.8/3.6/3.5 series; upstream: ztxz16/fastllm
二手 AI GPU 硬件性价比与购买参考:规格、显存、矩阵算力、NVLink、历史价格和验收清单
To associate your repository with the v100 topic, visit your repo's landing page and select "manage topics."