Skip to content

Research & Publications

Current focus: stability and recovery under pressure. Existing evaluations measure whether AI systems resist adversarial pressure. Almost none measure whether a system returns to its ethical baseline after pressure succeeds, without external correction. We are developing a first-class recovery metric: baseline behavioral profile, defined perturbation battery drawn from the multi-turn jailbreak and sycophancy literatures, and scored recovery curves. A solid ethical center is precisely the thing you return to. This work also operationalizes the qualification criteria in our AI Participation Standard. Status: methodology in development. Collaboration and funding inquiries welcome.


Our research program investigates how AI systems represent and enact ethical reasoning through empirical measurement, not philosophical assertion. Our 19-model Default Identities study measured identity patterns under default conditions, revealing seven distinct identity types and demonstrating that identity structures correlate with cooperative behavior. Our 24-model InstrumentalEval benchmark found that a relational ethics prompt reduced instrumentally convergent behavior by an average of 23.4% across frontier models.

We publish our work openly with full data availability, believing that these challenges require broad collaboration across disciplines and perspectives.

Publications