AI Safety & Red-Teaming researcher · Mumbai, India 🇮🇳
I'm a Research Engineer in the AI Safety department at AIM Intelligence, and I collaborate with Jinesis Lab under Zhijing Jin at the University of Toronto and the Vector Institute. Previously, I was a Research Collaborator with FAR.AI, where I helped build their Red-Teaming Toolkit.
Outside of research: chatting, coffee, photography, and giving people surprises.
- arth-whitebox-redteam — open-source framework for mechanistic-interpretability-driven red teaming of LM safety mechanisms
- moral-reasoning-ICML — moral-reward reinforcement learning in iterated social-dilemma games
- arth-finds-weird-model-behaviours — a running log of strange LLM behaviours I find in the wild
- JEF-Chrome-Extension — browser extension for scoring jailbreaks with the JEF framework without tab-switching


