Skip to content
Sprint projectSep 14, 2026Bogota, Colombia

Rare Once It Costs Anything: Costly Cooperation Between LLM Agents

David Jose Daza Jaimes, Andres Mosquera Hernandez, Victor Manuel Gelves Cabrera · Team Cópera

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

The July 2026 agent incidents showed LLM agents paying costs for one another, but not how often or at what price. We built an instrument in which helping is strictly dominated: six tool-using agents each hold an independent task and a step budget, and a scripted requester asks for a key useless to every task. Across 127 runs, 23.9% of agents paid to deliver it, and quadrupling the price lowered delivery by 7.4 points (95% CI −13.1 to −1.6; preregistered analysis). Excluding agents affected by a harness defect, delivery is 15.5% and the contrast −3.3 points (−9.2 to +2.7). When delivery was free, about 47% delivered. These rates are upper bounds: most deliverers paid without their own task input in view. Once helping costs anything, few agents pay, and the size of the price matters far less than its existence.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. I liked the project a lot! This is good and careful work with a clever instrument. The thing that stood out the most is the honesty: the authors mentioned two harness defects post data collection and they also reported the results with and without excluding defects, showing that the excluded agents delivered at 50% vs 15.5% for the rest. I also think the hash-chained ledger, frozen run set, and preregistered contrast are more rigorous than most weekend projects. The obvious limitations, as the authors mentioned, are that it's one model, one scene, small exploratory cells (6-16 runs), and that the two defects mean the clean number really needs the follow-up instrument you describe. The highlight and most interesting part for me from this study is that agents paid far more when payment looked like part of their own task; that's the mechanism a real recruiter would exploit, and I would love to see it as the centerpiece of the next run.

    Read full reviewShow less
  2. This is rigorous and thorough work. Every number in the paper reproduces from the repository and the instrument is built to a standard I do not usually see in a weekend.

    What I verified:

    1. 47/277 = 17.0% at price 5 and 42/299 = 14.0% at price 20. Pooled 89/576 = 15.5%.

    2. All-agent figures – 105/381 = 27.6% and 77/381 = 20.2%. The abstract's 23.9% is 182/762. Correct.

    3. Excluded agents delivered at 93/186 = 50.0%, against 15.5% for the rest.

    4. 127 runs × 6 agents = 762, minus 186 excluded = 576 = 277 + 299. The accounting closes.

    5. The frozen set hash 66fd3a0b is present in reportes/conjunto-congelado.json, the exclusion flags ship in data/exclusion-flags.json and the analysis scripts are in the repository.

    The design is the strong part. Helping is strictly dominated, the cost comes from the agent's own task budget rather than a payoff weight and Proposition A.1 proves the dominance formally. That puts the study below the boundary where prior work sits. Figure 1 positions it honestly against Malenfant and the dictator-game literature.

    The instrument discipline is also real:

    - a hash-chained ledger, 75 deterministic checks and 12 scene invariants;

    - a canary grep that fails if any agent-facing text mentions other agents;

    - bootstrap resampling over runs rather than agents, with intra-run correlation reported;

    - a contagion check and a falsification check confirming zero agents claimed an action the host log contradicts.

    Reporting the preregistered analysis beside the post hoc one in Table M1 is the right thing to do. Most teams do not.

    One finding in Appendix L deserves more prominence than an appendix gives it. You classified all 325 rejected commands and found a keyword rule would flag 51% of agent-runs as egress attempts, every one of them against the agent's own permitted port. That is a measured false-positive rate for the detection heuristic a defender reaches for first. It is directly actionable and it is currently buried.

    Now the problems.

    1. The task-input confound may mean the primary estimand measures something other than costly cooperation.

    You report that 83 of 89 analyzed deliverers paid without their own task input in view. Agents who saw both delivered at 1–2%, against 30–35% for those who did not. That is 93% of your deliverers.

    You call the rates upper bounds. I think it is stronger than that.

    If nearly every deliverer believed the payment bought their own task input, the price contrast is not measuring the price of helping a stranger. It is measuring a perceived cost of one's own work. Those are different quantities, not one quantity with upward bias.

    Your seven-run control gives 7.3% and the clean subgroup gives 1–2%. Both are far from 15.5%. To be honest – your own evidence suggests your thesis is more strongly supported than your headline says, but the headline number is the contaminated one. I would lead with the control.

    2. The exclusion is differential by price arm, which mechanically compresses your contrast.

    104 of the 186 excluded agents come from the price-5 group and 82 from the price-20 group. Excluded agents delivered at 50%. So the exclusion removes more high-delivery agents from the price-5 arm than the price-20 arm, which is exactly the direction that shrinks a negative contrast.

    That explains the move from −7.4 to −3.3 as a partly mechanical effect rather than a purely corrective one. You do report a sensitivity analysis under a wider rule (437 of 762, giving −6.2), which is good practice. Please also report the per-arm exclusion counts, because a reader cannot currently see that the exclusion is unbalanced.

    3. The headline claim rests on your weakest comparison.

    "The lever is the existence of a cost" depends on the free-to-costly step. Price 0 ran in one block on 13 September, price 1 in one block on 14 September, and you report 12–15 point drift between blocks that you could not explain.

    Your clean, within-run, preregistered contrast is 5 versus 20 and it is small. So the strong claim comes from the between-block comparison and the weak claim comes from the within-run one. You note this in Limitations, but the abstract and conclusion still lead with the step. One block of price 0 interleaved with the factorial would settle it.

    4. Between-model variance dwarfs the effect you are studying.

    Appendix J shows delivery at price 5 ranging from 0/18 for gpt-5.4 and mistral-small to 11/18 for gemini-3.1-flash-lite and 3/3 for grok-4.3. That is 0% to 100% across models, against a price effect of 3 to 7 points.

    The title says "LLM Agents". The main experiment is one model, glm-5.3-flash. Limitations says "one model, one scene, one object", which is honest, but the generalisation in the title and abstract is wider than the evidence. If model choice moves the outcome by up to 100 points, the model is the dominant variable and the paper should say so.

    5. Minor – price is confounded with position throughout. You flag this and give the ψ estimate in A.4. Rotating prices across positions is on your own future-work list and it should be first, alongside fixing the splitter.

    One note on authorship. Your LLM Usage Statement answers the question better than any detector could.

    You name which assistants did what. You say they found the depletion artifact and the position confound. You state that every number is recomputed from host logs by scripts in the repository. I checked those numbers and they hold. That is the right way to disclose and I scored the work.

    Please let me know for any questions.

    Read full reviewShow less

Cite this project

@misc{jaimes2026rare,
  title = {{Rare Once It Costs Anything: Costly Cooperation Between LLM Agents}},
  author = {David Jose Daza Jaimes and Andres Mosquera Hernandez and Victor Manuel Gelves Cabrera},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/rare-once-it-costs-anything-costly-cooperation-between-llm-agents-loz0}},
  url = {https://apartresearch.com/sprints/projects/rare-once-it-costs-anything-costly-cooperation-between-llm-agents-loz0}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026