Skip to content
Sprint projectMay 25, 2026New Delhi,India

TrajectoryCheck: Trajectory-Level Invariant Validation for Behavioral Drift Detection in Generated Code

Aamish Ahmad

Submitted to The Secure Program Synthesis Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: TrajectoryCheck: Trajectory-Level Invariant Validation for Behavioral Drift Detection in Generated Code

Share

TrajectoryCheck is a trajectory-level invariant validation framework that evaluates whether generated systems preserve behavioral consistency across execution phases. Instead of relying on static output inspection, the framework models execution as a sequence of operational phases connected through local and global invariants. Using AND-gate validation and stability scoring, the system detects trajectory drift, invariant collapse, and partial recovery behaviors across generated code executions.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The trajectory-level framing is reasonable, but the evaluation is three hand-built authentication examples with predetermined outcomes of 5/5, 0/4, and 1/4 gates, so it demonstrates the scoring arithmetic without testing whether the method catches drift in code it did not already know was broken. Applying it to a set of actual LLM-generated implementations, where the failures are not controlled in advance, is the experiment that would show the idea works. Automating the rule-based invariant extraction you list as future work would also move this from a concept sketch toward a usable tool.

  2. Maintaining invariants is a key method is program verification. This sprint covers local and global invariants. As far as I can tell, the implementation so far uses an LLM to check the invariants, and focuses on login processes. Going forward it would be interesting to see full formal invariant checking. It would be worth checking the literature to connect with the state of the art. This is a promising direction and a solid base to build on.

  3. TrajectoryCheck proposes a simple trajectory-level validation framework for generated code. The core idea is that generated software should not only be checked at isolated points, but across an execution sequence: registration, login, session validation, logout, and post-logout access, for example. This is a sensible framing, especially for security-sensitive systems where bugs often appear through state transitions rather than single function outputs.

    The report is clear and easy to understand. The phase-based model, local/global invariants, AND-gate validation, stability score, and drift map are all explained in straightforward language. The authentication example is also a reasonable first domain because it naturally has sequential behavior and security invariants.

    The main weakness is that the project is still very early and mostly conceptual. The evaluation uses only three toy authentication variants: stable, total drift, and partial recovery. This shows the scoring mechanism works, but it does not yet demonstrate that the framework can discover subtle failures in realistic generated code. The invariant extraction is rule-based, the examples appear hand-constructed, and there is no comparison against ordinary tests, property-based testing, runtime monitors, or formal verification.

    The stability score is interpretable but also quite simple. Counting passed gates is useful for a demo, but it treats all invariants as equally important and does not capture severity. For example, plaintext password storage and a minor phase inconsistency should not necessarily have the same weight. The method would be stronger if it supported severity-weighted scoring, automatic trace generation, and richer temporal properties.

    Overall, this is a clear and accessible prototype with a good intuition, but it needs more technical depth and stronger evaluation. The idea could become useful if expanded into a real trajectory-testing framework for generated programs, with automated invariant extraction, generated traces, realistic benchmarks, and comparison to existing testing methods.

    Read full reviewShow less

Cite this project

@misc{ahmad2026trajectorycheck,
  title = {{TrajectoryCheck: Trajectory-Level Invariant Validation for Behavioral Drift Detection in Generated Code}},
  author = {Aamish Ahmad},
  year = {2026},
  month = may,
  note = {Submitted to The Secure Program Synthesis Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/trajectorycheck-trajectorylevel-invariant-validation-for-behavioral-drift-detection-in-generated-code-ltjt}},
  url = {https://apartresearch.com/sprints/projects/trajectorycheck-trajectorylevel-invariant-validation-for-behavioral-drift-detection-in-generated-code-ltjt}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026