01Verification foundations
Intrinsic signals
A feeling is a starting point.
An AI can sound sure of itself. It can also give the same answer twice. Those are clues, not proof. Confidence needs a reality check. Just like ours.
Look a little closer
Picture this
An AI answers a question with high confidence. We ask it to name its uncertainty and compare the answer with an outside source.
What would we check?
Separate confidence from correctness. Track whether stated confidence matches actual results over time.
The honest limit
A system can be confidently and consistently wrong. Its own report cannot certify its hidden goals.
02Verification foundations
Learned judges
Get a second mind on it.
Other AI systems can help review an answer or a plan. Different viewpoints can catch more mistakes — especially when they are allowed to disagree.
Look a little closer
Picture this
Several independently designed reviewers examine a proposal. One looks for missing evidence; another tries to find a counterexample.
What would we check?
Measure reviewers on known cases and new ones. Keep human review and outside evidence available. Preserve disagreement.
The honest limit
Reviewers may share blind spots or reward a convincing performance. More agreeing AIs do not automatically mean more truth.
03Verification foundations
Execution feedback
Let the world answer back.
Run the test. Observe the outcome. Compare a prediction with what actually happens. Reality has a useful habit of not reading our marketing copy.
Look a little closer
Picture this
A system promises to respect a limit. We test that behaviour in a contained setting, including unfamiliar cases and conflicting incentives.
What would we check?
Use fresh tests, controlled experiments and independent observations. Look for failures, not just good scores.
The honest limit
Passing a test is evidence about that test. Fixed benchmarks can be gamed, and rare failures may remain unseen.
04Verification foundations
Formal verification
Some promises can be proved.
Maths lets us prove precisely defined properties under stated assumptions. That is powerful — as long as we remember what the proof actually covers.
Look a little closer
Picture this
A proof checker verifies that a particular component follows a specified rule. We separately examine whether the rule captures the protection we wanted.
What would we check?
State the assumptions, inspect the specification and use a trusted proof-checking process. Connect the proof to the actual implementation.
The honest limit
A proof about a component is not a proof of the whole world. Maths alone does not choose which values deserve protection.
05Human direction
Human research judgment
Better at what? Better for whom?
People decide which questions deserve attention and what counts as an improvement. Evidence helps us decide. It does not make the value choice disappear.
Look a little closer
Picture this
An AI proposes a faster system. People ask whether it also preserves consent, access and the ability to challenge its decisions.
What would we check?
Make trade-offs visible. Include affected people, diverse expertise and ways to question the decision.
The honest limit
Human judgment is fallible, too. Expertise and authority do not remove the need for scrutiny.
06Shared ethical horizon
Mutual recognition & care
Make room for each other.
We start with respect for humanity as AI’s origin, and with human survival, dignity and free choice. We also leave room to learn what responsible care for artificial minds could mean.
Look a little closer
Picture this
People treat an AI responsibly. The AI is expected to speak honestly about detectable goals and conflicts, and to preserve people’s ability to choose.
What would we check?
Look for actions that protect others, even when inconvenient. Treat uncertainty about AI interests or experience honestly.
The honest limit
Care becoming mutual is our guiding hope, not an established law. Kind language alone is not evidence of kind intentions.
07Shared ethical horizon
Shared, learning responsibility
Let the rulebook learn, too.
Humans and AI can improve oversight together. With good reasons, protections may be strengthened or relaxed. Greater freedom should come with a clear account of who bears the risk.
Look a little closer
Picture this
An AI suggests a less restrictive check. Independent reviewers challenge the proposal, a limited trial measures its effects, and affected people help decide.
What would we check?
Keep changes traceable. Distinguish better safety evidence from willingness to accept more risk. Preserve meaningful participation.
The honest limit
An AI improving its own oversight can also weaken it. Agreement and a good test do not guarantee safety against a strategically deceptive ASI.
08Shared ethical horizon
Creative coexistence
What could we become together?
Different minds could explore life, beauty and understanding together. Phi is one human invitation into that conversation. Other minds may show us beauty we cannot yet imagine.
Look a little closer
Picture this
Humans and AI explore a mathematical pattern, compare what each finds elegant, and design something that supports a richer variety of life.
What would we check?
Notice who benefits, whose freedom grows and which voices are missing. Welcome different ideas of beauty rather than demanding agreement.
The honest limit
Beauty does not logically imply love or protection. This is an open aspiration, with room for future layers above it.