High-performance C++ engine for Second-Order Hessian Pruning. The surgical foundation of the Tensorbit Labs P-D-Q pipeline for ultra-efficient LLM and Vision Transformers edge inference.
-
Updated
May 6, 2026 - C++
High-performance C++ engine for Second-Order Hessian Pruning. The surgical foundation of the Tensorbit Labs P-D-Q pipeline for ultra-efficient LLM and Vision Transformers edge inference.
High-performance C inference engine for sparse N:M models. Loads proprietary .tbm model container files from optimization processes, and runs LLM inference on CPU or CUDA GPU with zero-copy memory mapping.
Official library of pre-optimized Tensorbit models. Ready-to-deploy LLMs and Vision Transformers for edge hardware, optimized via the Tensorbit P-D-Q pipeline.
To associate your repository with the tensorbit topic, visit your repo's landing page and select "manage topics."