Sumit Yadav

Publications & writing

Publications 8

Peer-reviewed papers, preprints and reports — mechanistic interpretability, AI safety and low-resource NLP. Much of the interpretability work is sole-authored.

  1. ACL 2026 Main · CORE A*

    SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering

    Utsav Maskey, Sumit Yadav, Mark Dras, Usman Naseem

    Task-specific trajectories in embedding space separate refusal from non-refusal. Steering only over-refusal-prone tasks at inference time cuts over-refusals up to 73% — no retraining, minimal utility loss.

  2. Preprint 2026 sole-authored

    On the Relationship Between Representation Geometry and Generalization in Deep Neural Networks

    Sumit Yadav

    Effective dimension — unsupervised and label-free — predicts generalization across vision and language (partial r = 0.75 over 52 ImageNet models, 13 architecture families). Causality runs both ways: degrade the geometry with noise and accuracy follows (r = −0.94).

  3. GLOW @ IJCAI-ECAI 2026 sole-authored · best-paper candidate

    Geometric Phases of Mechanism Formation in Neural Networks

    Sumit Yadav

    Linear probes and CKA over dense checkpoints: classification mechanisms form output-layer-first, inside the first ~5% of training (Cohen's d = 3.68). The same deep-first pattern holds in the first ~200M tokens of LLM pretraining and reproduces on Pythia and OLMo-2.

  4. LoResLM @ ACL 2026 Rabat, Morocco

    MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili

    Sumit Yadav, Raju Kumar Yadav, Utsav Maskey, Gautam Siddharth Kashyap, Ganesh Gautam, Usman Naseem

    The first monolingual BERT for Maithili (~50M speakers), pre-trained on a newly built corpus. 87.02% on news classification — 5–7% over NepBERTa and HindiBERT — and open-sourced for downstream NER and sentiment work.

  5. J. Bus. Econ. Stud. 2024

    Revolutionizing Currency Security: A YOLOv8-Based Approach for Detecting Counterfeit Nepali Banknotes

    Sumit Yadav et al.

    A YOLOv8 detector for counterfeit 1000-rupee notes: true-positive recall of 0.82 on the front face and 0.986 on the back, portable across hardware platforms.

  6. B.E. Thesis 2024 Tribhuvan University (IOE, Pulchowk)

    Evaluating Auto-Encoding Transformer Language Models for Maithili Text Classification

    Sumit Yadav, Raju Kumar Yadav

    A Maithili masked language model built by transfer learning and fine-tuned on a curated news-classification set — the precursor work to maiBERT.

  7. Technical Report 2023 Tribhuvan University

    Machine Learning Analysis of Tirhuta Lipi

    Sumit Yadav, Raju Kumar Yadav

    Character recognition for the endangered Tirhuta script: MobileNet embeddings with logistic regression reach 0.97 accuracy, opening OCR and translation for Maithili.

  8. 2023

    Support Vectors Are a Better Way of Text Classification for Imbalanced Data

    Sumit Yadav et al.

    A TF-IDF + n-gram support-vector pipeline over 100+ imbalanced classes that beats neural baselines and retrains incrementally. Consolidates two 1st-runner-up competition submissions (LOCUS 2021, 2023).

Blog LLM safety & interpretability

Write-ups from my own runs. Each one opens on sumityadav.com.np, where the rest of the writing lives.

More writing — the full archive on sumityadav.com.np