A python library built on top of Inspect AI to support Control research and evaluations, by Redwood and EquiStamp
-
Updated
Sep 20, 2026 - Python
A python library built on top of Inspect AI to support Control research and evaluations, by Redwood and EquiStamp
Kernel-enforced authority and runtime security for AI agents, autonomous systems, and general Linux workloads.
Research developing AI control protocols using task decomposition
🚦🗺️ UrbanFlow AI is a web app for generating 3D traffic simulations from real OpenStreetMap areas. It builds SUMO scenarios, runs microscopic vehicles, pedestrians, buses and trams, edits road events, controls real traffic lights with TraCI, trains JSON AI policies, saves models, and shows live metrics, charts, and notebooks.
An intelligent traffic management system that dynamically adjusts highway lane configurations using AI-powered congestion detection and a movable median barrier.
Judge-first framework where LLM outputs must converge under explicit, adversarial oracles.
Python client for Aegis — stabilize AI systems instantly with a simple API call.
In-depth exploration of Large Language Models (LLMs), their potential biases, limitations, and the challenges in controlling their outputs. It also includes a Flask application that uses an LLM to perform research on a company and generate a report on its potential for partnership opportunities.
Kho lưu trữ này chứa tài liệu, bài tập, và mã nguồn liên quan đến môn Trí tuệ nhân tạo trong điều khiển. Môn học tập trung vào ứng dụng AI trong các hệ thống điều khiển tự động, bao gồm lý thuyết và thực hành.
Benchmark for detecting insider threats by AI agents in a simulated frontier AI lab
White-box detection of collusion in an untrusted monitor: model organisms, linear probes, and a control evaluation that prices what they buy.
An attacker × monitor factorial in ControlArena: which model you pick as your trusted monitor matters more than its capability tier.
AI Governance — human-controlled code injection with oversight
Independent AI governance and control standard
Replayable agent trajectories reconstructed from the July 2026 frontier-lab evaluation-containment failures, plus benign controls for measuring false-positive cost.
Lichtarbeit
v0.8 preview · 2,288 multi-step agent trajectories · 513 hand-authored gold + 1,775 provenance-flagged augmented. A per-step benchmark that scores whether a verifier catches drift inside an agent's trajectory — not whether a prompt is harmful.
GG Tank Watch - frozen public-information archive of a resolved May 2026 chemical emergency. Conduit-only design; responsible-AI safety patterns enforced in code and tests.
AutoRed: Measuring the Elicitation Gap via Automated Red-Blue Optimization — AI Control Hackathon 2026
Is cross-lingual chain-of-thought oversight failure monitor-side or model-side? A control-style deception eval on open-weight reasoning models.
To associate your repository with the ai-control topic, visit your repo's landing page and select "manage topics."