🦴 paleo — token-saving skills for LLM agents: compress output, trim context, cap budget (Claude Code / Codex / Gemini / Hermes)
-
Updated
Jul 25, 2026 - Python
🦴 paleo — token-saving skills for LLM agents: compress output, trim context, cap budget (Claude Code / Codex / Gemini / Hermes)
An open and practical guide to Edge AI Engineering.
ASAN: A conceptual architecture for a self-creating (autopoietic), energy-efficient, and governable multi-agent AI system.
KodaAI is a performance analytics platform that quantifies the efficiency gains provided by AI-assisted development using Kiro. It bridges the gap between human estimation (Jira) and machine-accelerated output (GitHub/Kiro) to calculate a "Speed Multiplier" for software teams.
A lightweight pre-inference gate for resource-efficient AI agents.
Smart LLM and cloud code optimization tool that cuts token costs by up to 75%. Perfect for AI developers seeking maximum efficiency and lower expenses in 2026.
LLM Token Optimization Toolkit — cut AI agent token cost 60–99%: Decision Distillation, SkillWeaver routing, Quant token playbook. Reduce LLM API bill.
Ternary Quantization for LLMs: Implement balanced ternary (T3_K) weights for 2.63-bit quantization—the first working solution for modern large language models.
Advanced cloud code optimization tool that dramatically reduces token consumption by up to 75% for AI and LLM applications. Perfect for developers and teams seeking maximum efficiency in 2026.
NeuroMem is a memory middleware layer that sits between your client and an upstream model. It retrieves session-scoped memories, injects them as structured context, stores the new turn, and returns the upstream model's answer. Use it when long conversations, large projects,
An open and practical guide to Edge Vision.
Claude Limit Extender is a smart desktop application that helps users significantly extend their productive time with Claude AI. It includes token optimization, session management, and intelligent prompt tools to maximize efficiency within platform limits.
Terminal-native control layer for your AI workflow — decides if AI is needed, then routes tool, skill, context, model, and verification.
AI Performance Engineering Cheatsheet: From Cloud to Edge.
Cut your AI coding agent costs by 60%. Battle-tested efficiency skills for Claude Code, Cursor, and any AI coding agent.
An open and practical guide to Edge Language
Interactive Training Dashboard & CAGS-Operator Verification for JamOne Nano.
Detect avoidable LLM waste: cache breakers, context bloat, and inefficient prompt design
Monitors token usage in real-time and suggests cost-effective model alternatives or prompts optimizations to reduce spending
To associate your repository with the ai-efficiency topic, visit your repo's landing page and select "manage topics."