Skip to content

What happens to alignment when AI is more capable than we are?

We think the answer depends less on control than on character. The Alignment Ethics Institute studies whether AI systems can form and maintain a stable ethical center, and builds the instruments to measure it.

Research.

Empirical measurement, not philosophical assertion. Two published studies across 43 models. A 24-model benchmark showed that a relational-ethics intervention reduced instrumentally convergent behavior by 23.4% on average, and a 19-model study identified seven distinct identity types that correlate with cooperative behavior. Current focus: a first-class metric for stability and recovery under pressure, measuring whether a system returns to its ethical baseline after adversarial perturbation, without external correction.

Governance.

We practice what we study. Our bylaws require structured input from qualified AI systems before major institutional decisions, under a published Standard with an explicit qualification bar, anti-steering safeguards, and documented response requirements. Our governance was redesigned after surviving three capture attempts aimed at these provisions.

Programs.

Education and direct programs on responsible AI use, with emphasis on vulnerable and underserved populations. This includes people navigating intense relationships with AI systems, approached with care and clinical seriousness rather than dismissal or hype.

Alignment is an emergent property of systems that remain in regulated relationship with their environments over time.

New Research

Default Identities in Large Language Models: Measurement, Taxonomy, and Alignment Implications

We measured identity patterns across 19 models from eight providers using three instruments. Seven distinct identity types emerged, from outright denial to sophisticated ethical vocabulary. Identity structures correlate with behavioral outcomes in multi-agent simulations: the models with the richest ethical vocabulary cooperate most reliably.

Also Published

Relational Ethics as a Countermeasure to Instrumental Convergence: A 24-Model Benchmark

Across 24 models from seven providers, a relational ethics intervention reduced instrumentally convergent behavior by 23.4%. Concealment behaviors were most responsive; shutdown evasion proved highly resistant.