Skip to content
Sprint projectJun 22, 2026Bengaluru

How does Instruction Hierarchy Training mitigate prompt injections: Preliminary results from an attentional study

Aman Neelappa, Amey Muke · Team AA

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: How does Instruction Hierarchy Training mitigate prompt injections: Preliminary results from an attentional study

Code (opens in new tab)
Share

We replicate and extend results from previous work on the causal impact of attention on instructions within tool responses as a mechanism for prompt injection. We further show that instruction hierarchy training partially mitigates this by reducing attention on instructions in tool responses and increasing that on legitimate user prompt

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is the most technically serious paper I reviewed, and it is well past what I would expect from a weekend. Most people would have stopped at the correlation. You went further and showed cause: knocking out the attention to the probe roughly halves the injection, and turning it up makes it worse in a dose-response way, with a random-span control to rule out the boring explanation. The leakage-free probe tells me you were careful not to fool yourselves. You are also upfront that your IHT model is an analog and the effect is small, which I appreciate. The flip side is that the IHT question itself stays half-answered, so the strong result here is the general mechanism, attention to the probe causing injection, not anything IHT-specific. The activation patching you mention as next is exactly what would close it. Strong work.

  2. A nice feature of this work is the honest handling of the PVCC result. The idea is intuitive and worth trying, and in practice it turns out to be only a weak, dataset-dependent signal. Rather than bury that, the authors report it plainly and use it to motivate the next step, which is useful for the community and a sensible response.

    The probe result is a nice surprise, and the leakage controls give good reason to trust the higher detection numbers. The causal intervention is perhaps the most useful part: blocking and amplifying the attention changes injection success in both directions against a null control, which moves the signal from a correlation towards an actual mechanism.

    The main weakness is scope. Everything is one model (Qwen3-8B), with no other families or sizes tested. The IHT setup is a homemade version of the real recipe rather than the actual training, and its effects are small, so the "shifts the operating point along the same mechanism" claim rests on a fairly thin IHT model. They also stop short of the circuit-level patching that would tie the behavioural gain to specific heads, though they flag it as the next step.

    To their credit, these are stated clearly and the IHT comparison is framed as an operating-point shift rather than a circuit-level claim. Overall a clean, honest proof of concept in a welcome direction, and I'd encourage the authors to pursue it.

    Read full reviewShow less
  3. Great work! The leakage-free probe controls (F2p/F3p with clean-control at chance) and the candid, detailed limitations section are exemplary. Two main weaknesses. Your IHT-analog is weak (only −3.5pp ASR, vs. much larger gains reported for production IHT), so the headline conclusion of an "operating point shift along the same mechanism" may not hold for a strongly trained hierarchy model. And without a GRPO clean-control (e.g. a reward-shuffled LoRA) you cannot yet attribute the attention shift to instruction-hierarchy training specifically; it could be an effect of RL fine-tuning in general. Prioritise those two ablations, plus the base↔IHT activation-patching experiment you propose, and include a code/artifact link. The results are strong but right now nobody can reproduce them independently. One small thing: a leftover sentence in the introduction ("against which an IHT counterpart will later be compared") contradicts the fact that Section 5 does the comparison. Worth fixing.

    Read full reviewShow less
  4. This is an ambitious project related to instruction injection. The authors show that the attention the model pays to the injected text is causally related to the propensity of the model to follow the instruction. Comparing a base model to a hierarchical-trained model (through LoRA), they also demonstrate that instruction hierarchy training alleviates the effect of instruction injection. Although the change is not really significant, it shows that this is an interesting direction to pursue with more time and means.

    A significant flow of the project is that the report is really unclear. It contains all the information but it took me (and Claude!) several reads to understand what this is about. A better context and guidance through the objective, methodology and results would have been appreciated. There are still elements I did not get in details, and that makes it difficult to find the methodological assumptions and the potential flaws of the project.

    I believe the most innovative part of this project is the interpretability element. Understanding the mechanism that makes IHT more robust to prompt injection is, to the best of my knowledge, a worthy research direction and the preliminary results shown here are promising.

    Read full reviewShow less

Cite this project

@misc{neelappa2026instruction,
  title = {{How does Instruction Hierarchy Training mitigate prompt injections: Preliminary results from an attentional study}},
  author = {Aman Neelappa and Amey Muke},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/how-does-instruction-hierarchy-training-mitigate-prompt-injections-preliminary-results-from-an-attentional-study-om7d}},
  url = {https://apartresearch.com/sprints/projects/how-does-instruction-hierarchy-training-mitigate-prompt-injections-preliminary-results-from-an-attentional-study-om7d}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026