Skip to content
View moudrkat's full-sized avatar

Block or report moudrkat

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please donโ€™t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this userโ€™s behavior. Learn more about reporting abuse.

Report abuse
moudrkat/README.md

Hey, I'm Kate ๐Ÿ‘‹

AI Engineer โ€ข Former Particle Physicist โš›๏ธ โ€ข Former Risk Modeler ๐Ÿ“ˆ

I've spent my whole career looking inside systems that would rather stay opaque โ€” first particle collisions, then risk models, now neural networks.

The models talk to us all day. I want to be able to affect them back.

So I build tools that make neural networks less mysterious โ€” because interpretability shouldn't stay in research papers, it belongs in production. (And in zombie games. And in mushroom generators. You'll see.)

Why I'm really doing this โ†’ the manifesto

๐Ÿ’ฅ Come say hi in my collision chamber

My personal site is a chat with a tiny LLM running entirely in your browser, and every answer it generates renders as a real particle collision. Click the event below to fire your own question into the chamber:

One question fired into the chamber: the model answers while its layers, attention heads and logit-lens flips render as a real collision.

๐Ÿš€ My main project is an opensource model interpretability lab - I develop it here and experiment with it on a real production app

The stack that makes a model legible in production โ€” from watching, to diagnosing, to fixing. Watch it think (brainscope) โ†’ diagnose why it did that (causal replay + the lens) โ†’ fix it at the source with a calibrated steering vector, receipts attached (hidden-directions โ†’ hotwire-vllm). Observability first; intervention last, once the instruments have earned it.

Tip

Don't read โ€” just do it. The core of the lab is on PyPI:

pip install hidden-directions brainscope hotwire-vllm

โ€” the vector factory with its eval framework, the live lens server, and the production vLLM steering plugin. Everything below runs on CPU or a free Colab GPU:

  • brainscope --model tiny โ€” a browser view of a model thinking (CPU is fine)
  • make demo in steering-mechanics โ€” real measured figures, no GPU at all
  • point your own OpenAI client at brainscope โ€” watch your app's live traffic

No account, no course โ€” install and look.

The lab runs one pre-registered research question: when does a steering vector generalize from calibration to deployment โ€” and what do steering evals actually measure? The hypotheses were written before the data, and they're allowed to lose.

Click any box to open its repo.

flowchart TD
    hd["๐Ÿงญ <b>hidden-directions</b><br/>behavior โ†’ vector"]
    bs(["๐Ÿง  <b>brainscope</b><br/>watch a model think<br/>(on your running app)"])
    st["๐Ÿ•น๏ธ <b>steeropathy</b><br/>agents talking through activations"]
    tm["โš–๏ธ <b>in-two-minds</b><br/>agent hesitating between tools"]
    hw["๐Ÿ”ฅ <b>hotwire-vllm</b><br/>steering in production vLLM"]
    sm["๐Ÿงช <b>steering-mechanics</b><br/>how steering actually works"]

    hd -->|"vectors"| bs
    bs -->|"hosts & captures"| st
    bs -->|"hosts & captures"| tm
    st -.->|"steers with"| hd
    hd -->|"vector + passport"| hw
    bs <-.->|"same spec: lab โ†” prod"| hw
    bs -->|"causal replay"| sm
    hw -.->|"vector under study"| sm

    click hd "https://github.com/moudrkat/hidden-directions"
    click bs "https://github.com/moudrkat/brainscope"
    click st "https://github.com/moudrkat/steeropathy"
    click tm "https://github.com/moudrkat/in-two-minds"
    click hw "https://github.com/moudrkat/hotwire-vllm"
    click sm "https://github.com/moudrkat/steering-mechanics"

    classDef engine fill:#1f6feb,stroke:#1158c7,color:#ffffff;
    classDef exp fill:#8957e5,stroke:#6e40c9,color:#ffffff;
    class bs,hd,hw engine;
    class st,tm,sm exp;
Loading

The blue boxes are the instrument. brainscope hosts any Hugging Face model and streams its internals to the browser; hidden-directions makes the steering vectors โ€” auto-calibrates them (Optuna, with a KL damage guard), bakes them into weights, then audits for the bake; hotwire-vllm takes those vectors to production โ€” steering inside vLLM's CUDA graphs, per request, steered speed = vanilla vLLM. All three speak one steering spec: a vector calibrated under the lens deploys unchanged, and a misbehaving production conversation replays back under the lens.

The purple boxes are experiments run under that lens. steeropathy wires agents together through activations instead of text; in-two-minds catches an agent hesitating between tools before it commits; steering-mechanics asks how steering vectors actually work inside the model.


๐Ÿค What I'm looking for

Collaborators and users โ€” not a job (see the manifesto). If you build on LLMs and want to see inside your model, or you work on steering / interpretability and want to compare notes โ€” or run SteerBench against your own method โ€” open an issue on any repo and say hi. The single best thing you can do: pip install, try it, and tell me where it breaks.


๐Ÿ”ฌ Also on the bench โ€” smaller, self-contained ways to look inside
  • ๐Ÿ“œ paper-remembers โ€” Hopfield's 1982 paper, running live: rub out any part of the page and watch it rebuild itself
  • ๐ŸŽญ sixteen-voices โ€” how a tiny transformer encodes writing style, through LoRA adapters and attention heads
  • ๐Ÿ‘๏ธ show-me-your-attention โ€” attention maps and neuron activations over your own prompt
  • ๐Ÿ’ฅ detektor โ€” the collision chamber above, open source (SmolLM2 in your browser, no server)
  • ๐Ÿ–ผ๏ธ jepa-demo โ€” I-JEPA & V-JEPA 2 hands-on, no GPU needed, with a visual deep-dive article
  • ๐Ÿ„ Mushroom-generator โ€” a VAE growing mushrooms, with latent-space walks and the decoder taken apart layer by layer
  • ๐ŸŽ Applepear โ€” apples vs pears in a tiny CNN, activations and grad-CAM included
  • โš™๏ธ Minimize_me โ€” race TensorFlow optimizers across loss landscapes
๐Ÿƒ And off the bench
  • ๐ŸŽจ personal-rembrandt โ€” you can't build a personal brand, so build a personal Rembrandt: paste your bio, GPT-2 reads it in your browser, and its activations repaint his 1659 self-portrait.
  • ๐Ÿ›๏ธ go-to-damn-bed โ€” a Claude Code skill that sends you to bed like a mom sends naughty children: it saves your work into TOMORROW.md, then counts to three. It never says what happens at three
  • ๐Ÿ‘‘ KingOfDiamonds โ€” the King of Diamonds game from Alice in Borderland, played by LLMs in character, recursive strategic thinking and all
  • ๐Ÿ—จ๏ธ paralel-discordverse โ€” your company's Discord gets a parallel universe, populated entirely by fictional colleagues
  • ๐Ÿงฎ least-squares-method โ€” code archaeology: a printed Pascal listing, photographed page by page and revived on Turbo Pascal 5.5

None of it is perfect. That's kind of delightful.

Pinned Loading

  1. brainscope brainscope Public

    OpenAI-compatible server with a live view into any HF model's residual stream. pip install brainscope

    Python 40 8

  2. hidden-directions hidden-directions Public

    Steering vectors with receipts: make one, catch one, deploy a calibrated one. pip install hidden-directions

    Python 1

  3. steeropathy steeropathy Public

    Agents that talk through model internals โ€” activations & J-space โ€” instead of text. No words pass between them. The lab on top of brainscope + hidden-directions.

    Python 21 2

  4. hotwire-vllm hotwire-vllm Public

    CUDA-graph-safe per-request activation steering plugin for vLLM. pip install hotwire-vllm

    Python 1