Official code for the paper: "Simple LLM Baselines are Competitive for Model Diffing"
-
Updated
Feb 13, 2026 - Python
Official code for the paper: "Simple LLM Baselines are Competitive for Model Diffing"
ActDiff: diagnose, repair, and prevent persistent domain priors in narrowly finetuned vision-language models.
Supplementary code & results for "Variant-specific crosscoder features are seed-stable but not detectably task-causal in a GRPO-LoRA math setting" (ICML 2026 Mech Interp Workshop, Spotlight)
Sealed, preregistered benchmark for black-box model-diffing agents: five LoRA finetunes of Qwen3.5-9B (one null, three planted behaviours, one dropped backdoor), audited blind by Neel Nanda's diffing-agent recipe and four cheaper conditions. The recipe fails by not asking, and the auditor itself is a failure mode. MATS 12 application.
Code and artifacts for The Convergence Gap: when instruction-tuned models settle on next-token predictions.
Code and artifacts for Same Targets, Different Computation: how post-training divides work across model layers.
Research workspace for model diffing between pretrained and post-trained language models.
To associate your repository with the model-diffing topic, visit your repo's landing page and select "manage topics."