AI researchPython · PyTorch · Transformers
VEMR: verified LLM model surgery
Verified Extract → Modify → Reinsert. A transactional protocol for editing a live language model's weights.
The problem
Existing tools only cover half of model editing. Activation-patching frameworks change a model for one forward pass and then the change is gone. Weight-editing tools make permanent changes but never check whether the change broke the model, and there is no way to undo it.
What I built
VEMR backs up a layer or a single neuron, applies any edit to a scratch copy, loads the edited weights tentatively, and re-runs a fixed probe set. If the KL divergence stays under a calibrated threshold, the edit is committed. If not, the original weights are restored bit-for-bit. The model always ends up in exactly one of those two states.
It ships as a Python CLI and session API. It covers layer and neuron extraction, a logit lens, attention and neuron viewers, circuit discovery and benchmarking. It detects the architecture of any causal LM it loads, both GPT-2 style and Llama style models. I wrote it up as a research paper.
- 1Extractsnapshot weights + checksum
- 2Modifyedit a scratch copy
- 3Reinsertload tentatively
- 4VerifyKL divergence on probe set vs τ
- Δ ≤ τ → commit Δ > τ → rollback
- 100%
- correct commit/rollback decisions across 40 labelled edits on 4 layers
- 16 / 16
- severe edits caught. Without verification, perplexity rose to 1.1×10⁷ (baseline 61)
- 0
- of 12 mild edits wrongly rejected, so the check does not block safe changes
- L8 · N802
- single GPT-2 neuron found, characterised and edited surgically
