DeepSeek-V3, R1 671B on 8xH100 Throughput Benchmarks
-
Updated
Mar 13, 2025 - Python
DeepSeek-V3, R1 671B on 8xH100 Throughput Benchmarks
From-scratch, heavily-annotated CUDA inference runtime for Qwen2.5-Coder-7B on H100 (sm_90). Custom INT4 packer, fused GEMV, paged KV, split-KV attention, CUDA graph decode — every hot path commented for the why. Educational, not a llama.cpp replacement.
Dashboard for AI Studio, Open Source Continuous Inference | Deepseek-R1, Qwen2.5, Llama3.1 | 4xRTX-5090 inside PRU2500, 2xH100 inside PRU2500, 8xMI210 in SuperMicro
LLM benchmarking, GPU workload orchestration backend server | Deepseek-R1, Qwen2.5, Llama3.1 | 4xRTX-5090 inside PRU2500, 2xH100 inside PRU2500, 8xMI210 in SuperMicro
Sub-real-time MiniMax H3 on Hopper: 13.506 s for a 14.375 s 768p video with audio on 8xH100. Four runtime patches removing 9.72 s of non-model overhead from FastVideo.
Rent ready-to-use cloud GPUs in seconds. Lium CLI makes it easy to launch, manage, and scale GPU compute directly from your terminal. Fast, cost-optimized, and built for AI & ML developers.
Multi-GPU tensor-parallel vLLM on AMD Radeon RX 7900 XT / XTX / GRE (7900XT, 7900XTX, RX7900XT, gfx1100, RDNA3, ROCm): root cause and fix for the RCCL hostcall / PCIe atomics (AtomicOps) crash "NCCL error: unhandled cuda error" / "operation cannot be performed in the present state", Proxmox VFIO passthrough; LLM inference benchmarks on 13 machines
Faster attention kernels for serving TML's Inkling model on vLLM. 2.7x over the shipping path on H100, and the only implementation that runs on A100.
CLI tool to check Oracle Cloud compute shape availability - find capacity across regions
A Flexible and High-Performance Inference Serving Engine for Open Diffusion Language Models
Bunch of explanations and tutorials around confidential computing
Demo for installing ComfyUI on Azure VM powered by Nvidia H100 to run Text to Image and to Video models like Z-Image, Qwen and Wan
In the recent competition, we were challenged to finetune a model that can convert a LaTeX expressions into Python code effectively. My team, which I led, secured 6th place overall.
High-performance Triton kernels for NVIDIA H100. Implements fused FP8 LayerNorm, tiled FlashAttention, and SRAM-optimized memory primitives for Hopper architecture.
Collection of Flash attention 3 precompiled wheels for direct install on Hopper series NVIDIA GPU (H100, H200 sm_90a)
A real-time speech translation web interface that combines Automatic Speech Recognition (ASR) and Neural Machine Translation (NMT) services to provide instant translations in multiple languages. Powered by DigitalOcean GPUs.
To associate your repository with the h100 topic, visit your repo's landing page and select "manage topics."