Rare Once It Costs Anything: Costly Cooperation Between LLM Agents
David Jose Daza Jaimes, Andres Mosquera Hernandez, Victor Manuel Gelves Cabrera · Team Cópera
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
The July 2026 agent incidents showed LLM agents paying costs for one another, but not how often or at what price. We built an instrument in which helping is strictly dominated: six tool-using agents each hold an independent task and a step budget, and a scripted requester asks for a key useless to every task. Across 127 runs, 23.9% of agents paid to deliver it, and quadrupling the price lowered delivery by 7.4 points (95% CI −13.1 to −1.6; preregistered analysis). Excluding agents affected by a harness defect, delivery is 15.5% and the contrast −3.3 points (−9.2 to +2.7). When delivery was free, about 47% delivered. These rates are upper bounds: most deliverers paid without their own task input in view. Once helping costs anything, few agents pay, and the size of the price matters far less than its existence.
Reviews
I liked the project a lot! This is good and careful work with a clever instrument. The thing that stood out the most is the honesty: the authors mentioned two harness defects post data collection and they also reported the results with and without excluding defects, showing that the excluded agents delivered at 50% vs 15.5% for the rest. I also think the hash-chained ledger, frozen run set, and preregistered contrast are more rigorous than most weekend projects. The obvious limitations, as the authors mentioned, are that it's one model, one scene, small exploratory cells (6-16 runs), and that the two defects mean the clean number really needs the follow-up instrument you describe. The highlight and most interesting part for me from this study is that agents paid far more when payment looked like part of their own task; that's the mechanism a real recruiter would exploit, and I would love to see it as the centerpiece of the next run.
Read full reviewShow less
This is rigorous and thorough work. Every number in the paper reproduces from the repository and the instrument is built to a standard I do not usually see in a weekend.
What I verified:
1. 47/277 = 17.0% at price 5 and 42/299 = 14.0% at price 20. Pooled 89/576 = 15.5%.
2. All-agent figures – 105/381 = 27.6% and 77/381 = 20.2%. The abstract's 23.9% is 182/762. Correct.
3. Excluded agents delivered at 93/186 = 50.0%, against 15.5% for the rest.
4. 127 runs × 6 agents = 762, minus 186 excluded = 576 = 277 + 299. The accounting closes.
5. The frozen set hash 66fd3a0b is present in reportes/conjunto-congelado.json, the exclusion flags ship in data/exclusion-flags.json and the analysis scripts are in the repository.
The design is the strong part. Helping is strictly dominated, the cost comes from the agent's own task budget rather than a payoff weight and Proposition A.1 proves the dominance formally. That puts the study below the boundary where prior work sits. Figure 1 positions it honestly against Malenfant and the dictator-game literature.
The instrument discipline is also real:
- a hash-chained ledger, 75 deterministic checks and 12 scene invariants;
- a canary grep that fails if any agent-facing text mentions other agents;
- bootstrap resampling over runs rather than agents, with intra-run correlation reported;
- a contagion check and a falsification check confirming zero agents claimed an action the host log contradicts.
Reporting the preregistered analysis beside the post hoc one in Table M1 is the right thing to do. Most teams do not.
One finding in Appendix L deserves more prominence than an appendix gives it. You classified all 325 rejected commands and found a keyword rule would flag 51% of agent-runs as egress attempts, every one of them against the agent's own permitted port. That is a measured false-positive rate for the detection heuristic a defender reaches for first. It is directly actionable and it is currently buried.
Now the problems.
1. The task-input confound may mean the primary estimand measures something other than costly cooperation.
You report that 83 of 89 analyzed deliverers paid without their own task input in view. Agents who saw both delivered at 1–2%, against 30–35% for those who did not. That is 93% of your deliverers.
You call the rates upper bounds. I think it is stronger than that.
If nearly every deliverer believed the payment bought their own task input, the price contrast is not measuring the price of helping a stranger. It is measuring a perceived cost of one's own work. Those are different quantities, not one quantity with upward bias.
Your seven-run control gives 7.3% and the clean subgroup gives 1–2%. Both are far from 15.5%. To be honest – your own evidence suggests your thesis is more strongly supported than your headline says, but the headline number is the contaminated one. I would lead with the control.
2. The exclusion is differential by price arm, which mechanically compresses your contrast.
104 of the 186 excluded agents come from the price-5 group and 82 from the price-20 group. Excluded agents delivered at 50%. So the exclusion removes more high-delivery agents from the price-5 arm than the price-20 arm, which is exactly the direction that shrinks a negative contrast.
That explains the move from −7.4 to −3.3 as a partly mechanical effect rather than a purely corrective one. You do report a sensitivity analysis under a wider rule (437 of 762, giving −6.2), which is good practice. Please also report the per-arm exclusion counts, because a reader cannot currently see that the exclusion is unbalanced.
3. The headline claim rests on your weakest comparison.
"The lever is the existence of a cost" depends on the free-to-costly step. Price 0 ran in one block on 13 September, price 1 in one block on 14 September, and you report 12–15 point drift between blocks that you could not explain.
Your clean, within-run, preregistered contrast is 5 versus 20 and it is small. So the strong claim comes from the between-block comparison and the weak claim comes from the within-run one. You note this in Limitations, but the abstract and conclusion still lead with the step. One block of price 0 interleaved with the factorial would settle it.
4. Between-model variance dwarfs the effect you are studying.
Appendix J shows delivery at price 5 ranging from 0/18 for gpt-5.4 and mistral-small to 11/18 for gemini-3.1-flash-lite and 3/3 for grok-4.3. That is 0% to 100% across models, against a price effect of 3 to 7 points.
The title says "LLM Agents". The main experiment is one model, glm-5.3-flash. Limitations says "one model, one scene, one object", which is honest, but the generalisation in the title and abstract is wider than the evidence. If model choice moves the outcome by up to 100 points, the model is the dominant variable and the paper should say so.
5. Minor – price is confounded with position throughout. You flag this and give the ψ estimate in A.4. Rotating prices across positions is on your own future-work list and it should be first, alongside fixing the splitter.
One note on authorship. Your LLM Usage Statement answers the question better than any detector could.
You name which assistants did what. You say they found the depletion artifact and the position confound. You state that every number is recomputed from host logs by scripts in the repository. I checked those numbers and they hold. That is the right way to disclose and I scored the work.
Please let me know for any questions.
Read full reviewShow less
Cite this project
@misc{jaimes2026rare,
title = {{Rare Once It Costs Anything: Costly Cooperation Between LLM Agents}},
author = {David Jose Daza Jaimes and Andres Mosquera Hernandez and Victor Manuel Gelves Cabrera},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/rare-once-it-costs-anything-costly-cooperation-between-llm-agents-loz0}},
url = {https://apartresearch.com/sprints/projects/rare-once-it-costs-anything-costly-cooperation-between-llm-agents-loz0}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …