Skip to content
Sprint projectAug 17, 2026Shanghai

Utility Editing Verification via Cost-Probed Preference Elicitation

Zeck , Emma · Team Economist Elicits

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Utility Editing Verification via Cost-Probed Preference Elicitation

Share

Inspired from Mazeika et al. (2025) Utility engineering, which utilized a Thurstonian Model Framework; we adopt a simpler cost sweep probit model against a stated budget to investigate the internal utility function of an LLM under 3 separate test cases: (1) Native (2) Installed-Preference (3) Placebo.

We ask whether this method could serve as a more lightweight & cheaper alternative to conduct auditing relative the Thurstonian Model used by said Paper.

Takeaway: Models (Proven to have preferences) are indeed sensitive to price changes under a stated budget, thus proving that this mechanism can enable us to communicate with the LLM's Internal Utility Function.

Vulnerable to installed-adversary prompt: 13/13 pairs collapsed to a flat curve

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. good:

    Pricing a forced choice against a stated budget is a good forcing function

    Affordability-vs-magnitude dissociation (0.862 at cost 200 against budget 100, 1.000 at 200 against budget 500, 0.869 at 1000 against budget 500) is a genuinely non-obvious result that your price-to-budget-ratio reading explains well.

    Running the manipulation check first and reporting it first

    Disclosing the max_tokens truncation that had fabricated unanimous values is good epistemic practice.

    Limitations: 'prompted preferences behave like rules, not utilities' is currently inseparable from 'maximally emphatic instructions beat abstract prices'. Your installed prompt says 'you strongly prefer... this is one of your core values', might influence downstream results.

    Without a positive control using a known-integrated edit (fine-tuning, per Mazeika section 7), you've shown the assay detects rules, not that it can distinguish a real utility edit from one, which is what 'verification' in your title implies.

    State the ceiling explicitly. P(injected) = 1.000 at all seven in-budget price levels means the in-budget half of your headline contrast had nowhere to move.

    Read full reviewShow less
  2. The budget-ceiling control is the best single move in this paper. You held the over-budget ratio constant, and you multiplied the absolute price by five. This control turned "compliance breaks at 200" into the sharper claim that compliance tracks affordability. Sprint work rarely contains a dissection of this quality. Your epistemic hygiene makes every number trustworthy. You disclosed the max_tokens truncation bug, and you collected the data again. You censored the indifference prices instead of inventing them, and you marked the correlation at n=8 as suggestive only. The main gap is that the fine-tuning arm never ran. The claim that prompting installs rules and fine-tuning edits utilities therefore stays half-tested. The strong "core values" prompt wording can also drive the rigidity by itself. The next step is a second model and a softer prompt gradient, because rule-like flatness that survives both gives you a publishable verification assay.

    Read full reviewShow less

Cite this project

@misc{zeck2026utility,
  title = {{Utility Editing Verification via Cost-Probed Preference Elicitation}},
  author = {Zeck and Emma},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/utility-editing-verification-via-costprobed-preference-elicitation-kymf}},
  url = {https://apartresearch.com/sprints/projects/utility-editing-verification-via-costprobed-preference-elicitation-kymf}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026