Software Development Engineer at Amazon and independent AI safety researcher.
I study whether model-auditing and model-diffing methods can recover the conditional behaviour introduced through fine-tuning—not merely identify the training topic.
My work uses controlled fine-tunes, blinded evaluations, causal activation interventions, preregistration, adversarial controls, and reproducible evaluation pipelines.
Testing whether model-diffing methods recover a fine-tuned model's actual conditional behaviour. Includes eight controlled Qwen3-1.7B fine-tunes, a 720-response blinded benchmark, causal activation interventions, matched controls, and published negative results.
- Blindfolded Expert Iteration — Preregistered evaluation-awareness study with hash-locked validation and stopping rules.
- Sandbagging Evaluations — Black-box and activation-based detection of strategic underperformance.
- Local LLM Fine-Tuning Lab — MLX/QLoRA training with versioned datasets, frozen evaluations, and checkpoint gates.
- Toy Models of Superposition — Mechanistic-interpretability and sparse-autoencoder reproductions.
- VeriTraceMix — Formal verification of traceable mixnets using Dafny.
- B.Tech (Hons.) in Electrical Engineering, IIT Delhi
- Software Development Engineer at Amazon
- Research experience in verifiable electronic voting, ProVerif, and Dafny
- BlueDot AI Safety course alum
- Independent study using the ARENA mechanistic-interpretability curriculum
- LinkedIn: omung-agarwal