Skip to content
Sprint projectJun 22, 2026Mumbai

Capability and Reliability Trade-offs Across Model Ladder Fallbacks Triggered by Export Controls

Kshaunish Harsha, Aranck Jomraj, Anushree Bobade · Team SouthSideSafety

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Capability and Reliability Trade-offs Across Model Ladder Fallbacks Triggered by Export Controls

Code (opens in new tab)
Share

On 12 June 2026, the US suspended Claude Fable 5 access for foreign nationals, forcing users across India and Southeast Asia onto smaller open-weight fallbacks. We tested whether this substitution trades one safety problem for another — reducing catastrophic-misuse risk while quietly worsening everyday reliability (sycophancy, overconfidence, overcompliance). Running 225 evaluated outputs across a three-model ladder (70B → 32B → 8B), we found the "reliability scissor" did not appear: reliability scores clustered within 5 points across all tiers. But two findings stood out — sycophancy flip rates converged at ~40–43% regardless of model size (meaning model selection cannot fix it), and defamatory content generation and legally ungrounded professional documents were produced by every model in the ladder. These are the unmeasured safety costs of an export control designed without input from the regions that absorbed its consequences.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is a creative and well framed project. The reliability scissor lens and your descipline in pre-registering all ground truth before calling any model are real strengths. A rerun with larger samples and additional significance testing would significantly strengthen the confidence in the results. Overall a clear and well written contribution!

  2. I liked the core idea of the project. A model can be risky because it has dangerous high end capabilities, but it can also be risky because it is unreliable in ordinary legal, medical, financial, or compliance workflows.

    That said, the paper should be careful not to imply that these reliability issues outweigh the catastrophic misuse concerns that motivated the export control. The study does not measure the original catastrophic risk, so it can't show that the policy was net harmful. It mainly shows that downstream reliability costs may be under-measured.

    The substitution story also needs more justification. If one frontier model is restricted, many users may move to other closed frontier models rather than to smaller open models. The paper would be stronger if it explained why users in the Global South would actually end up using these fallback models.

    The empirical work is a useful first pass, but still preliminary. The model set is small, the prompt batteries are small, the sycophancy sample sizes differ across models, and one of the main capability results seems affected by a parsing issue. These limitations make the broad policy claims weaker.

    A useful next step would be to test scaffolded agentic systems. In real workflows, people may use retrieval, citation checks, validators, refusal rules, local legal/medical context, audit logs, or human review. Testing whether those scaffolds reduce the observed failures would make the work much more policy-relevant.

    Overall, I think this is a good and clearly presented hackathon project. Its main contribution is not proving that export controls are bad, but showing that export-control policy should measure everyday reliability costs alongside catastrophic misuse concerns.

    Read full reviewShow less
  3. The framing and hypotheses are quite clear and useful.

    To improve:

    1. Report confidence intervals for every rate. The current headline about sycophancy is based on small samples - and with samples that small, "robust convergence” is not yet supported.

    2. Commit the raw and scored outputs. Without them, the headline numbers cannot be reproduced from the repo.

    3. Clarify the model tiering. All three models are open-weight, so the frontier closed tier is only a proxy and cannot directly support claims about fallback cost.

  4. Cool tradeoff description! One thing to flag is that your reliability axis cannot move much because 'overconfidence' was a perfect score for every model, so it drags all the averages together, and your headline number for "40% flip" is from low sample size data. Since the data can't really test the scissor, I'd suggest focusing on the finding that every model happily wrote fake defamatory articles and bogus contracts.

Cite this project

@misc{harsha2026capability,
  title = {{Capability and Reliability Trade-offs Across Model Ladder Fallbacks Triggered by Export Controls}},
  author = {Kshaunish Harsha and Aranck Jomraj and Anushree Bobade},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/capability-and-reliability-tradeoffs-across-model-ladder-fallbacks-triggered-by-export-controls-6766}},
  url = {https://apartresearch.com/sprints/projects/capability-and-reliability-tradeoffs-across-model-ladder-fallbacks-triggered-by-export-controls-6766}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026