A library for making RepE control vectors
-
Updated
Sep 24, 2025 - Jupyter Notebook
A library for making RepE control vectors
[ICLR 2025] General-purpose activation steering library
Steering vectors for transformer language models in Pytorch / Huggingface
A resource repository for representation engineering in large language models
[🏆 CHI26 Best Paper] CoBRA: Reproducible control of LLM agent behavior via classic social science experiments
KV Cache Steering for Controlling Frozen LLMs
Agents that talk through model internals — activations & J-space — instead of text. No words pass between them. The lab on top of brainscope + hidden-directions.
Lightweight representation engineering dataflow operations for agent developers.
[EMNLP 2026 Main] Steering Geometry: Validating Human Value Geometry in LLM Steering Space.
[🔥 ICLR 2026] - Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots
Activation steering and trait monitoring for HuggingFace transformers
[Under Review] Not All Tokens Are Equally Useful for Steering: Robust Directions and Prefix Steering
Official code for "Activation Steering for Accent Adaptation in Speech Foundation Models" (Interspeech 2026). Parameter-free accent adaptation via mean-shift steering vectors — no weight updates, consistent WER reductions across 8 accents.
CRSM (Continuous Reasoning State Model): An asynchronous "System 2" architecture that implements Hierarchical State Sovereignty within a Mamba backbone. Unlike traditional search wrappers, CRSM uses Forward-Projected Planning and Sparse-Gated Injection to steer latent manifolds in real-time, decoupling strategic reasoning from token generation.
Turn a knob inside a small open model instead of writing a prompt. A reproducible RepE and CAA steering harness with an honest benchmark, including the failures.
Steering vectors with receipts: make one, catch one, deploy a calibrated one. pip install hidden-directions
Mechanistic interpretability experiments on political control circuits, refusal behavior, concept steering, and late-decoder interactions in open LLMs.
Phase-aware LLM activation steering and linear probing. A memory-efficient, practical implementation of Representation Engineering (RepE) for safety research.
Steer2Adapt: data-efficient inference-time LLM adaptation by composing steering vectors via Bayesian optimization over a semantic prior subspace.
To associate your repository with the representation-engineering topic, visit your repo's landing page and select "manage topics."