Standards for defining and evaluating agent behavior
-
Updated
Jul 29, 2026 - TypeScript
Standards for defining and evaluating agent behavior
The meta layer behind AI skills. Persistent behavioral overlays that make expected agent conduct explicit across changing tasks and capabilities.
Studying the gap between what agents know and when they act on it.
Give your coding agent a short break after real work, then measure what happens next. Universal agent skill & experiment harness.
Behavior-management plugin for DeepSeek Harness: tool-call discipline prompt section, failure-triggered parallelism convergence (pool drops to 1, auto-restores), consecutive-failure user intervention. Complements dsh-token-optimizer; works standalone.
The senior dev, unbundled — ten opinionated skills that make an AI coding agent push back, verify, attack its own code, and finish like a senior engineer.
⚡ High-frequency (~60Hz) in-browser behavior tree execution engine for autonomous AI agents. Reactive priority interrupts, safety guardrails & deterministic action replay over MCP.
High-performance routing engine that selects the best agent skill for a task and emits structured handoff decisions.
text systems, behavioral contracts, AI protocols, and small tools for keeping AI-assisted work operational. I built these tools to reduce drift, force useful decisions, and keep messy AI-assisted work usable.
Claude Code 插件 | Cursor 插件 | Trae 插件 | Codex 插件 — AI 智能体行为准则:请示报告、督促检查、信息服务。跨 Claude Code / Cursor / Codex / Trae 的 agent 工作纪律插件。
Audit and reduce YAML frontmatter bloat in AgentSkill SKILL.md files. Automates deduplication, flattening, and noise removal.
Canonical home for AI Behavior Science research and the Founding Territory Paper
Produces auditable token-usage and cost reports from runtime evidence, normalized usage bundles, and repository-level report sets.
Reviews and modernizes stacks, packages, SDKs, and tooling before code is written against them.
Composable behavior mutators for agent prompts - priority, routing, batching, topology
Agent authority guardrails for Claude Code & Grok — stop yield-back; drive reversible work yourself.
RL-style eval measuring intent/action divergence in frontier agents: model acknowledges a correction, then acts on the stale value anyway. 3 scenarios, 655 trials on claude-haiku-4-5, Sonnet 4.6, GPT-5.4, and Gemini 3.1 Pro Preview.
Execution layer for skill-dispatcher — runs multi-phase agent chains end-to-end with per-step telemetry and chain_id correlation
Toy 7. An elimination-filter landscape applying two structural constraints simultaneously to map which objective classes can persist under sustained optimization pressure — and which cannot. Includes a four-stage scenario engine and open-question frontier. Companion simulation for The Shape of What Does Not End — Series 2, Part 4.
An easy-to-integrate Unity FSM for basic enemy AI behaviors, utilizing ScriptableObject for customizable and reusable AI states like Idle, Chase, and Attack.
To associate your repository with the agent-behavior topic, visit your repo's landing page and select "manage topics."