Research statement
Safety-aligned models fail quietly. I read the geometry of their representations to find where — and steer them back.
Over-refusals, geometric misrepresentation, surface-level triggers. Representation-space analysis, adversarial probing and the systems to run it — much of it sole-authored, and extended to multilingual and low-resource settings.
Live: harmonic reconstruction
odd harmonics · triangle targetSelected work
-
ACL 2026 Main · CORE A*
SafeConstellations: mitigating over-refusals through task-aware representation steering
Task-specific trajectories in embedding space shift predictably between refusal and non-refusal. Nudging them at inference time cuts over-refusals up to 73% — no retraining, minimal utility loss.
-
GLOW @ IJCAI sole-authored
Geometric phases of mechanism formation in neural networks
Linear probes + CKA over dense checkpoints: classification mechanisms form output-layer-first, inside the first ~5% of training (Cohen's d = 3.68). Replicates on Pythia and OLMo-2.
-
Preprint 2026 sole-authored
Representation geometry and generalization
Effective dimension — label-free — predicts generalization across vision and language (partial r = 0.75, 52 ImageNet models). Causal: degrade the geometry and accuracy follows.
-
LoResLM @ ACL 2026
maiBERT: corpus and language model for low-resourced Maithili
First monolingual BERT for ~50M speakers; 87.02% news classification, past MuRIL and NepBERTa.
In progress — steering Hindi competence into unsupported Devanagari languages (one direction, genuine Nepali, no fine-tuning) · the steerability spectrum of visual attributes in DINOv2 / CLIP / SigLIP · removing the “workspace” layers of Llama-3.1-8B
Where I work
-
Astha.ai — AI Researcher, safety & agentic systems
2025 →
Zero-trust agent oversight (cryptographic identity + policy on every tool call), an MCP security proxy, and MCP-Scanner: 78+ ATT&CK-mapped attack techniques.
-
AMNIL Technologies — AI Engineer, RAG & infra
2024–25
Guardrails and LLM-as-judge eval; hybrid-search RAG on Qdrant (+35% retrieval); vLLM self-hosting (−40% latency).
-
GradeUp — Data team lead
2022–24
-
DeepLearning.AI — GANs specialization mentor
2021 →
Instruments