[ICLR 2025] General-purpose activation steering library
-
Updated
Sep 18, 2025 - Python
[ICLR 2025] General-purpose activation steering library
Benchmark evaluation code for "SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal" (ICLR 2025)
We study whether categorical refusal tokens enable controllable and interpretable safety behavior in language models.
🔓 Ablate — directional ablation (abliteration) toolkit for open-source LLMs. Automatic censorship/refusal removal via residual-stream direction ablation, with KL-guided search, an LLM-judge harness, and one-call push to the Hub. pip install ablate-llm
Reproducible, evergreen benchmark for LLM refusal on biological research prompts — 19 models, 141 prompts, 13,389 adjudicated trials
CLI for keeping Claude Fable 5 prompts, skills, and API traffic in shape — lint anti-patterns, canary silent degradation, aggregate refusal analytics. Every rule cites Anthropic docs.
Public Driftmap harness: public-safe CSV suites + rubrics + run logs for drift detection, refusal integrity, injection resistance, and uncertainty tracking.
RAG with verifiable citations and measured refusal — retrieval scored separately (TF-IDF beats embeddings here), citations validated against chunks actually retrieved.
A probe suite that measures which conversation states an LLM cannot leave. Three arms, because two cannot tell obedience from token statistics; a null only counts when the design had the power to see the effect.
中文企业公开报告 Hybrid RAG:Docling 解析 · Qdrant 稠密/稀疏检索 · 查询理解硬过滤 · 带引用生成与拒答 · 文档生命周期与评测看板
Refusal verification surface for Riverbraid fail closed policy boundaries.
Classify response traces with offline refusal/shaping heuristics and compare probe-result labels across saved snapshots.
Training-time defense that redistributes LLM refusal via mean/covariance matching + KD, raising linear-ablation attack rank from K=1 to K≥16 (Llama-3.2-1B-Instruct)
Locating and editing refusal in the J-space workspace with the Jacobian lens: refusal is legible ~10 layers before the first token, and only ~1/3 lives in the verbalizable workspace.
Unified CLI/TUI to abliterate any (V)LLM (Heretic / OBLITERATUS / ErisForge) and benchmark the methods on one schema — G/P/S composite + Pareto front. Pure-stdlib core, no GPU for the core.
An open reproduction of feature-level activation steering with the prompt set released, showing the capability tax that behavioural metrics miss
How does this function's cost scale? Counted, not timed — and UNDETERMINED when no complexity class settles.
Public reference interfaces for proof-gated AI action, refusal, authority, and evidence boundaries.
A document copilot that cites what it says and refuses when the evidence is not there: 33 questions, baseline 12/33 to harness 32/33, control 8/8 both ways.
Mechanistic interpretability of refusal behavior in Qwen2.5 models: sparse feature interventions, residual steering, judge-vs-rule analysis, and 1.5B replication
To associate your repository with the refusal topic, visit your repo's landing page and select "manage topics."