Skip to content
#

tpu-inference

Here is 1 public repository matching this topic...

Frontier-class open models on a free Kaggle TPU v5e-8: GLM-5.3-Flash 320B MoE (~64 tok/s, our own JAX engine) and Qwen3.8-27B bf16 (~130 tok/s), 262k context, prefix caching. Works with Claude Code, Codex, opencode and pi.

  • Updated Sep 15, 2026
  • Python

Add this topic to your repo

To associate your repository with the tpu-inference topic, visit your repo's landing page and select "manage topics."

Learn more