Sumit Yadav

Mech interp · alignment & safety · robustness

Research statement

Safety-aligned models fail quietly. I read the geometry of their representations to find where — and steer them back.

Over-refusals, geometric misrepresentation, surface-level triggers. Representation-space analysis, adversarial probing and the systems to run it — much of it sole-authored, and extended to multilingual and low-resource settings.

Live: harmonic reconstruction

odd harmonics · triangle target
RMS err 0.0000 drag plot to scrub buffers — — fps

Selected work

  1. ACL 2026 Main · CORE A*

    SafeConstellations: mitigating over-refusals through task-aware representation steering

    Task-specific trajectories in embedding space shift predictably between refusal and non-refusal. Nudging them at inference time cuts over-refusals up to 73% — no retraining, minimal utility loss.

  2. GLOW @ IJCAI sole-authored

    Geometric phases of mechanism formation in neural networks

    Linear probes + CKA over dense checkpoints: classification mechanisms form output-layer-first, inside the first ~5% of training (Cohen's d = 3.68). Replicates on Pythia and OLMo-2.

  3. Preprint 2026 sole-authored

    Representation geometry and generalization

    Effective dimension — label-free — predicts generalization across vision and language (partial r = 0.75, 52 ImageNet models). Causal: degrade the geometry and accuracy follows.

  4. LoResLM @ ACL 2026

    maiBERT: corpus and language model for low-resourced Maithili

    First monolingual BERT for ~50M speakers; 87.02% news classification, past MuRIL and NepBERTa.

In progress — steering Hindi competence into unsupported Devanagari languages (one direction, genuine Nepali, no fine-tuning) · the steerability spectrum of visual attributes in DINOv2 / CLIP / SigLIP · removing the “workspace” layers of Llama-3.1-8B

Where I work

  • Astha.ai — AI Researcher, safety & agentic systems

    2025 →

    Zero-trust agent oversight (cryptographic identity + policy on every tool call), an MCP security proxy, and MCP-Scanner: 78+ ATT&CK-mapped attack techniques.

  • AMNIL Technologies — AI Engineer, RAG & infra

    2024–25

    Guardrails and LLM-as-judge eval; hybrid-search RAG on Qdrant (+35% retrieval); vLLM self-hosting (−40% latency).

  • GradeUp — Data team lead

    2022–24

  • DeepLearning.AI — GANs specialization mentor

    2021 →

Instruments

activation steeringlinear probes / CKAeffective dimension nnsight / NDIFPyTorchvLLMPEFT / LoRA Model Context Protocolprompt-injection defenselow-resource NLP QdrantDocker / K8s