Frontier-class open models on a free Kaggle TPU v5e-8: GLM-5.3-Flash 320B MoE (~64 tok/s, our own JAX engine) and Qwen3.8-27B bf16 (~130 tok/s), 262k context, prefix caching. Works with Claude Code, Codex, opencode and pi.
kaggle glm pallas tpu jax mixture-of-experts free-gpu openai-api long-context llm-serving vllm llm-inference qwen speculative-decoding coding-agents anthropic-api claude-code qwen3 glm-5-3-flash tpu-inference
-
Updated
Sep 15, 2026 - Python