Skip to content
Partner hackathon projectNov 20, 2024

A Fundamental Rethinking to AI Evaluations: Establishing a Constitution-Based Framework

Arrow Paquera, and Paul Ivan Enclonar · Team Unit 1112

Submitted to Howard University AI Safety Summit & Policy Hackathon. Projects from partner hackathons are early-stage work by participants, not Apart Research publications.

Read the report

Report: A Fundamental Rethinking to AI Evaluations: Establishing a Constitution-Based Framework

Share

While artificial intelligence (AI) presents transformative opportunities across various sectors, current safety evaluation approaches remain inadequate in preventing misuse and ensuring ethical alignment. This paper proposes a novel two-layer evaluation framework based on Constitutional AI principles. The first phase involves developing a comprehensive AI constitution through international collaboration, incorporating diverse stakeholder perspectives through an inclusive, collaborative, and interactive process. This constitution serves as the foundation for an LLM-based evaluation system that generates and validates safety assessment questions. The implementation of the 2-layer framework applies advanced mechanistic interpretability techniques specifically to frontier base models. The framework mandates safety evaluations on model platforms and introduces more rigorous testing for major AI companies' base models. Our analysis suggests this approach is technically feasible, cost-effective, and scalable while addressing current limitations in AI safety evaluation. The solution offers a practical path toward ensuring AI development remains aligned with human values while preventing the proliferation of potentially harmful systems.

Reviews

Judging this hackathon?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. Seems like a scalable, cost-effective, and rigorous solution! Achieving a global consensus would be difficult so could mention how this could start (ie smaller scale/modular), would be useful to have some dicussion on how compliance with the constitution-based framework would be enforced as companies would be resistant (eg reduced liability)

Cite this project

@misc{paquera2024fundamental,
  title = {{A Fundamental Rethinking to AI Evaluations: Establishing a Constitution-Based Framework}},
  author = {Arrow Paquera and Paul Ivan Enclonar},
  year = {2024},
  month = nov,
  note = {Submitted to Howard University AI Safety Summit \& Policy Hackathon, a partner hackathon},
  howpublished = {\url{https://apartresearch.com/sprints/projects/a-fundamental-rethinking-to-ai-evaluations-establishing-a-constitution-based-framework}},
  url = {https://apartresearch.com/sprints/projects/a-fundamental-rethinking-to-ai-evaluations-establishing-a-constitution-based-framework}
}

More from Howard University AI Safety Summit & Policy Hackathon

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026