(Realtime) Temporal Convolutions in PyTorch
-
Updated
Apr 7, 2025 - Python
(Realtime) Temporal Convolutions in PyTorch
Streamable Text-to-Speech model using a language modeling approach, without vector quantization
Native-video memory for vision-language-action models, using timestamped visual history and exact streaming inference for long-horizon robot manipulation.
High-performance runtime for real-time multimodal generation and world models—streaming inference, stateful sessions, and distributed GPU execution.
Dual-model speech AI toolkit for speaker verification and speaker-aware diarization, with streaming inference, meeting analysis, long-audio monitoring, and speaker-bank integration.
World's most deployable time series foundation model — 200K-6.5M params, zero-shot forecasting, streaming RNN inference, ONNX edge deployment, runs on Raspberry Pi
Quality-aware adaptive streaming segmentation for volumetric ultrasound research
Zaphira is a hybrid resident+streamed inference engine built for one hardware story: running Mistral's large coding models (Devstral 2 123B, Devstral Small 2 24B) on four consumer RTX 3060s.
Pure PyTorch + 🤗 Transformers reimplementation of Megalodon (CEMA + chunked attention) - readable, hackable, no CUDA kernels required
Lossless AI model compression - ~34% smaller with bit-identical weights; the autopilot profiles your machine, picks the highest fidelity that runs, and streams models bigger than your RAM.
ML systems platform for distributed training, transformer inference, profiling, benchmarking, observability, and model serving.
An end-to-end MLOps pipeline for vehicle diagnostics featuring a temporal LSTM stateful inference engine.
CascadeLUT: Information-Ordered Streaming Inference for Bandwidth-Constrained FPGAs [FPL'26]
Efficient State Space Model layers in pure PyTorch — FFT training, streaming inference, ONNX export for edge deployment
High-performance JAX-to-TensorRT compilation pipeline and decoupled gRPC streaming inference server for quantitative trading architectures.
Don't Think It Twice: Exploit Shift Invariance for Efficient Online Streaming Inference of CNNs
Real-time music-genre classification: spectrogram CNN, ONNX-optimised, served as a streaming/chunked classifier with PyTorch-vs-ONNX benchmarks. Track-aware GTZAN eval.
Real-time voice AI microservice - WebRTC, multi-tenant architecture, STT/TTS, streaming inference
Streaming version of S4ND-U-Net
CPU-native inference runtime. Local-propagation paradigm: the active region pays the cost, not the field. Bit-exact across architectures. Validated for streaming anomaly detection and audio VAD.
To associate your repository with the streaming-inference topic, visit your repo's landing page and select "manage topics."