FlashAttention v1 forward pass in CUDA for NVIDIA Turing (SM75)
-
Updated
Jul 27, 2026 - Python
FlashAttention v1 forward pass in CUDA for NVIDIA Turing (SM75)
Compiler MVP that detects Transformer fusion patterns, generates optimized CUDA kernels with WMMA Tensor Cores, and executes them on real GPU hardware — 10.5 TFLOPs on RTX 2070, correctness validated against PyTorch.
Backward pass CUDA kernels for fused GEMM+Bias+GeLU on SM75 (Turing). Float4 vectorized WMMA kernels validated against PyTorch autograd. 3.1x faster than autograd at M=1024. 24/24 tests passing.
Instruction-level Tensor Core micro-kernel engineering on SM75 (RTX 2070), including PTX manual emission, shared-memory staging, and hardware-aware auto-tuning.
Reproducible FreeToken MoE expert-offload compatibility lab and benchmarks for NVIDIA Turing sm_75.
Qwen3.8-27B 双 RTX 2080 Ti 22GB + NVLink 的 vLLM 部署(上游 main 路线):SM75 移植、MTP 下保住 FULL cudagraph(单并发解码 54.1→37.5 ms/步)、256K/500K 上下文与 128K tiered offload profile
To associate your repository with the sm75 topic, visit your repo's landing page and select "manage topics."