Skip to content
Sprint projectJun 22, 2026Bogota, Colombia

Thought Anchors for Social Bias: Which Reasoning Steps Matter in Extended Thinking LLMs on Latin American Scenarios

Andres Felipe Mosquera Hernandez

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Thought Anchors for Social Bias: Which Reasoning Steps Matter in Extended Thinking LLMs on Latin American Scenarios

Recording (opens in new tab)Code (opens in new tab)
Share

We investigate which reasoning steps in extended-thinking LLMs are associated with pro-stereotypical outputs on Latin American social bias scenarios. Social bias benchmarks, including BBQ \citep{parrish2022bbq}, SESGO \citep{robles2024sesgo}, and EsBBQ/CaBBQ \citep{ruizfernandez2025esbbq}, measure final-answer distributions but provide no visibility into intermediate reasoning. We adapt the thought-anchor resampling framework \citep{bogdan2025thoughtanchors} to social bias measurement: for each sentence-level chunk in a model's chain-of-thought, we remove it, resample $K{=}3$ continuations, and measure the change in pro-stereotypical output probability via three importance metrics. We apply this pipeline to SESGO-small, an 80-item stratified pilot subset spanning four Latin American bias categories (\textit{género, clasismo, racismo, xenofobia}), two languages, and three recent models with extended thinking. Across 1,708 chunk-level attributions, disambiguated contexts show higher pro-stereotypical output rates than ambiguous ones (33–75\% vs.\ 0\%), and high-importance reasoning positions tend to be tagged as \texttt{logical\_reasoning} or \texttt{context\_recall} rather than \texttt{stereotype\_activation}. These pilot results suggest that bias-relevant moments may lie in general inference steps rather than overt stereotype invocations, offering a step-level diagnostic target that complements answer-level benchmarks. More broadly, attributing bias to specific reasoning steps could help localize where stereotyping enters a model's reasoning, though our findings remain preliminary.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Adapting thought-anchor resampling to social bias measurement in extended-thinking LLMs on Latin American scenarios opens a research direction that did not previously exist. The finding that high-importance steps are logical_reasoning and context_recall (not stereotype_activation) is counterintuitive and potentially significant. This suggests that bias enters reasoning disguised as general inference, which has direct implications for how mitigation interventions should be designed. To consolidate the work, the most urgent next steps are increasing K from 3 to at least 10 rollouts per chunk to obtain importance estimates with higher resolution, validating semantic chunk labels with human annotation, and applying the pipeline to the full SESGO benchmark for adequate statistical power per cell. The future work question of whether backtracking acts as a protective mechanism against stereotypical outputs is one of the most important questions the paper leaves open.

    Read full reviewShow less
  2. This is a strong and creative extension of Thought Anchors into social-bias evaluation.

    The repo is very reviewer-friendly, with clear reproduction commands and cost disclosure.

    Small note: the video is announced but I could not play it. I would have enjoyed seeing it, not many teams did one and it helps reviewers a lot.

  3. Agentic harnesses have been seen to greatly improve capability, and therefore I think harnesses are a very important safety angle to study. So I like the problem choice. If I'm understanding the result correctly, I think it is quite interesting that bias is coming out in "reasoning steps" and not in "recall steps", so somehow it seems that even if bias is not recalled directly from model "memory" it could come out in agents in more subtle ways. I think this is a good angle to study.

    I find it hard to study the correctness of methodology and validity of results under my time constraints, but I like the problem choice. Presentation could use less jargon to be followed at least at a high level under strict time constraints.

  4. This is a very strong and focused contribution. The paper applies a resampling approach to understand where social bias appears inside the model’s reasoning, using the Latin American benchmark SESGO. This is a genuinely interesting angle, because it goes beyond simply looking at final answers. The pipeline is carefully built, and the main finding is interesting: the reasoning steps that shift bias are often classified as general reasoning or context recall, rather than direct stereotype activation. The biggest strength is the paper’s honesty about its limits. It is clear that the small SESGO exercise, with 80 items and limited validation, is a proof of concept rather than a basis for firm conclusions. To strengthen the paper, I would scale the analysis to the full SESGO benchmark, use more repeated runs, validate the labels with human coders, and report agreement across coders. Excellent writing and reproducibility. A very strong project.

    Read full reviewShow less

Cite this project

@misc{hernandez2026thought,
  title = {{Thought Anchors for Social Bias: Which Reasoning Steps Matter in Extended Thinking LLMs on Latin American Scenarios}},
  author = {Andres Felipe Mosquera Hernandez},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/thought-anchors-for-social-bias-which-reasoning-steps-matter-in-extended-thinking-llms-on-latin-american-scenarios-27ti}},
  url = {https://apartresearch.com/sprints/projects/thought-anchors-for-social-bias-which-reasoning-steps-matter-in-extended-thinking-llms-on-latin-american-scenarios-27ti}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026