A curated list of resources for activation engineering
-
Updated
Oct 2, 2025
A curated list of resources for activation engineering
Runtime control of LLM agent behaviors through activation steering vectors. More calibrated than prompting.
Iterative Sparse Matrix Steering: Closed-Form Subspace Alignment for Multi-Layer LLM Control (No SGD required).
A closed-loop control system for Large Language Models that steers internal activation states in real-time to prevent mode collapse and toxicity
🔓 Ablate — directional ablation (abliteration) toolkit for open-source LLMs. Automatic censorship/refusal removal via residual-stream direction ablation, with KL-guided search, an LLM-judge harness, and one-call push to the Hub. pip install ablate-llm
A hands-on, research-oriented journey into transformer architecture, mechanistic interpretability, representation learning, and large language model internals.
How meaning moves through a transformer - found, traced, and tested across four model scales.
Add a description, image, and links to the activation-engineering topic page so that developers can more easily learn about it.
To associate your repository with the activation-engineering topic, visit your repo's landing page and select "manage topics."