NeurIPS 2026 · Workshop on Interpretability for Discovery

Your Intervention Was Too Narrow

Mechanism Width as a Precondition for Interpretability-Driven Discovery

Sumit YadavBasanta Joshi

Pulchowk Campus, Institute of Engineering, Tribhuvan University, Nepal

When an intervention on one direction, one layer, or one attention head changes nothing, the usual conclusion is that the part is not involved. Across three unrelated systems, that test recovers only about a tenth of the effect it is looking for. A width-1 null is an underpowered test, not evidence of absence.

Top: a mechanism drawn as a bundle of k-star strands; intervening on one strand captures 10 percent of the effect, on all k-star strands 95 percent, and when the reader is disconnected no width helps. Bottom left: width sweeps for GPT-2 heads, Llama-8B layers and DINOv2 orientation directions rise together from about 0.1 at width 1 to 0.95 at their own k-star. Bottom right: in PaliGemma the representation gap closes to 81 percent while the language model's readout stays near zero.
Figure 1. (a) Width sweeps at three granularities, normalized by each mechanism's own k*eff. Passing through (1, 0.95) holds by construction; the evidence is everything left of it, and where the intervention-free k*repr (triangles) lands. (b) The readout boundary: PaliGemma's representation moves while its answer does not.

The standard unit recovers about 10%

8.6%
one attention head, IOI circuit sufficiency
GPT-2 small · k*eff = 8.6 heads
9.6%
one decoder layer removed
Llama-3.1-8B · k*eff = 13 layers
11.1%
rank-1 steering of orientation
DINOv2 · k*eff = 53 directions

Widening each intervention to its mechanism's width k* recovers 95–100% of the achievable effect.

Abstract

Interpretability-driven discovery rests on interventions: a part is credited with a function when perturbing it changes behavior, and dismissed when nothing happens. We show that the field's standard intervention unit (one direction, one layer, one attention head) recovers only ≈10% of the achievable causal effect in three unrelated testbeds, one task each: rank-1 steering in vision encoders (0.111), single-layer ablation in Llama-3.1-8B (0.096), and single-head sufficiency in GPT-2's IOI circuit (0.086). The mechanism's width k* can be measured before intervening, and widening the intervention to k* recovers 95–100%. The effect is about width itself: a 12-layer band does 7.5× the additive damage prediction, random subsets track contiguous bands, and pushing past k* destroys object identity. Across 12 language models, whether a single-unit test works is predicted by a cliff in the writer-score ladder (ρ = 0.81). Width is necessary, not sufficient: at a documented readout boundary, 81% of a representation gap closes while the consumer's output moves ≈0, though it reads the attribute when prompted. A negative result at standard width is an underpowered test, not evidence of absence; a cheap protocol (sweep width, compare k*repr to k*eff, run matched controls) tells the difference before a discovery claim is made.

Two widths, three verdicts

k*repr is read from the representation alone, with no intervention. k*eff is the smallest intervention reaching 95% of the achievable effect. Their relation classifies a failed intervention:

Width-limited

k*eff exists and sits near k*repr. The standard test was simply too narrow; widening fixes it. DINOv2 orientation: 49 vs 53.

Readout-limited

The representation moves, the consumer does not. No width helps; change the handle. PaliGemma: 81% of the gap closed, captions invariant.

Basis-limited

Small k*repr but no saturating sweep: a curved, nearly one-dimensional manifold a linear handle grips poorly. DINOv2 size peaks at 0.40.

Varying all four analytic thresholds over 81 settings never moves a system between these classes (paper, Appendix B).

Role criteria are width choices too

Counting heads that pass a role criterion (here, previous-token score ≥ 0.5 × max) is the obvious representation-side width for attention heads. Selecting heads with no role score, by greedily adding whichever upstream head does the most damage jointly with those already chosen, shows the count is not reliable in either direction.

ModelCriterion headsk*eff (joint greedy)OverlapCriterion set's shareWidth 1
GPT-2 small2102/20.510.34
GPT-2 medium6216/60.430.03
GPT-2 large7145/70.320.03
GPT-2 XL9220/90.200.01
Pythia-410m524/50.980.80
Pythia-1.4b484/40.850.24
Llama-3.2-1B783/70.700.33

In GPT-2 the criterion finds the right heads first but stops at 20–51% of the mechanism; in GPT-2 XL it shares none of the first nine. Measured against this criterion-free mechanism, the writer-ladder cliff predicts width-1 success with ρ = 0.89 (p = 0.007, n = 7).

Before reporting an interventional null

  1. Measure k*repr from the representation, with no intervention.
  2. Sweep width past it and locate k*eff, the width reaching 95% of the achievable effect.
  3. Run matched controls: random units of equal width and magnitude, plus a prompted capability check on the consumer.
  4. Classify: saturates near k*repr → width-limited; representation moves but the consumer does not → readout-limited; linear handles fail on a low-stable-rank manifold → basis-limited. Only then interpret the null.

More results

Citation

@inproceedings{yadav2026narrow,
  title     = {Your Intervention Was Too Narrow: Mechanism Width as a
               Precondition for Interpretability-Driven Discovery},
  author    = {Yadav, Sumit and Joshi, Basanta},
  booktitle = {NeurIPS 2026 Workshop on Interpretability for Discovery},
  year      = {2026},
  note      = {Non-archival}
}