#
ninfer
Here are 5 public repositories matching this topic...
Port NInfer, a single-GPU CUDA inference engine, to the NVIDIA L20 (Ada sm_89, 92 SMs, 48 GB): patch set, build tooling, and measured results
-
Updated
Sep 20, 2026 - PowerShell
Measuring proxy + throughput dashboard for local LLM engines (llama.cpp, NInfer): live tok/s, cache-hit rate, TTFT, and history charts. Single-file panel, stdlib-only proxy, MIT.
python dashboard metrics prometheus throughput observability llamacpp llm-inference tokens-per-second ninfer
-
Updated
Sep 20, 2026 - HTML
Qwen3.8-27B on RTX 4090 D (48GB): production deployment of NInfer with MTP7 + E8 KV + NVMe disk cache, 195 tok/s decode, crash forensics for WDDM desktop GPUs
-
Updated
Sep 3, 2026 - C++
Add this topic to your repo
To associate your repository with the ninfer topic, visit your repo's landing page and select "manage topics."