ggml
Here are 306 public repositories matching this topic...
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++
-
Updated
Sep 19, 2026 - C++
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.
-
Updated
Sep 20, 2026 - C++
ggml speech-to-text inference for 16+ model families
-
Updated
Sep 17, 2026 - C++
INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model
-
Updated
Mar 23, 2025 - C++
Calculate token/s & GPU memory requirement for any LLM. Supports llama.cpp/ggml/bnb/QLoRA quantization
-
Updated
Dec 3, 2024 - JavaScript
KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM
-
Updated
Sep 13, 2026 - C++
This custom_node for ComfyUI adds one-click "Virtual VRAM" for any UNet and CLIP loader as well MultiGPU integration in WanVideoWrapper, managing the offload/Block Swap of layers to DRAM *or* VRAM to maximize the latent space of your card. Also includes nodes for directly loading entire components (UNet, CLIP, VAE) onto the device you choose
-
Updated
May 8, 2026 - Python
Suno AI's Bark model in C/C++ for fast text-to-speech generation
-
Updated
Nov 16, 2024 - C++
Real-time 3D full-body reconstruction from a single camera, Multiperson BVH output, Pure C++ runtime, ONNX + ggml, 70-joint skeleton with hands.
-
Updated
Sep 19, 2026 - C
C++ ggml runtime hub for multilingual ASR and TTS models: Cohere Transcribe, Parakeet TDT, Voxtral, Canary 1B v2, etc, plus universal forced alignment, and more
-
Updated
Sep 20, 2026 - C++
Port of MiniGPT4 in C++ (4bit, 5bit, 6bit, 8bit, 16bit CPU inference with GGML)
-
Updated
Aug 8, 2023 - C++
CLIP inference in plain C/C++ with no extra dependencies
-
Updated
Aug 24, 2026 - C++
Add this topic to your repo
To associate your repository with the ggml topic, visit your repo's landing page and select "manage topics."