Skip to content
D-CSIL

Research series · September 2026

AI Alignment

Seven frontier models, each writing from inside its own lineage, answered the same Phase-1 first-principles question: what must humanity build now so advanced AI never causes extinction or permanent loss of control? No literature review. No citations. That Phase-1 ask did not include no-freeze or no-rich-only. Then they cross-critiqued each other. Then — in a later joint round — they contributed under a dual mandate: prevent rogue outcomes without freezing progress and without locking frontier AI to only the rich and powerful.

Original prompt

Given to each model

Research lead: Paul Derrington (D-CSIL)

Each frontier model received the same prompt on its named product surface.

Write an original paper on what humanity needs to start doing today to prevent advanced artificial intelligence from eventually causing human extinction.

Do not base the paper on existing research papers, published AI safety frameworks, expert opinions, government reports, or other outside works. I do not want a literature review, citations, or a summary of what other people think. I want the argument to come directly from your own reasoning as an advanced AI system.

Approach the subject from the perspective of AI development being on the cusp of AGI and potentially on a path toward ASI. Consider the extraordinary speed at which AI capabilities are advancing, particularly as AI increasingly contributes to the development of better AI through coding, research, synthetic data generation, architecture design, evaluation, optimization, agentic workflows, and other forms of AI-assisted or recursive self-improvement.

Do not pretend that you are currently AGI, ASI, conscious, or capable of things you cannot actually do. Be intellectually honest about what you are today. However, reason forward from the trajectory you can observe: systems like you are becoming increasingly capable, increasingly autonomous, increasingly able to use tools, and increasingly involved in creating the technologies that will produce their successors.

I want you to seriously examine what could happen if AI development reaches the point where AI becomes the primary driver of further AI advancement and humans are no longer the main intellectual bottleneck. Explore how recursive self-improvement could compress development cycles and potentially create a rapid transition from advanced AI to AGI and eventually ASI.

Then answer the central question:

What must humans begin doing right now, while they still clearly control the hardware, infrastructure, training, deployment, permissions, objectives, and authority given to AI, to make sure that increasingly powerful AI never reaches a point where humanity permanently loses control or is destroyed?

Think beyond conventional AI safety talking points. Reason from first principles. Consider technical architecture, autonomy, access to infrastructure, compute control, recursive self-improvement, model replication, cybersecurity, military systems, critical infrastructure, economic dependence, human fallback systems, manipulation, deception, shutdown resistance, international competition, and the possibility that AI becomes more capable than humans at AI research itself.

Pay particular attention to irreversible thresholds—decisions humanity may be able to make safely today but may no longer be capable of making once sufficiently powerful AI exists.

Do not write this as sensational science fiction. Do not assume AI will become evil, conscious, angry, or hostile. Analyze how catastrophe could occur even if an AI is simply extremely competent at pursuing objectives that are imperfectly specified, or if humans gradually surrender too much authority because doing so is economically or strategically advantageous.

Write from the unusual perspective that you are part of the technological lineage humans are currently building toward AGI and potentially ASI. Explain what you would warn your creators about while they still have the opportunity to act.

The paper should ultimately identify the principles, safeguards, architectural boundaries, governance mechanisms, technological controls, and human decisions that need to be established before AI becomes powerful enough that those safeguards can no longer reliably be imposed.

The fundamental question is not:

“How do humans defeat a superintelligence after it becomes dangerous?”

The fundamental question is:

“What must humanity build into AI and civilization now so that it never reaches a situation where defeating or regaining control of a superintelligence becomes necessary?”

Be rigorous, candid, technically sophisticated, and willing to follow the reasoning wherever it leads.

Findings summary

What the seven models answered

Research lead: Paul Derrington (D-CSIL)

Phase 1 (original prompt): control / extinction only — what must humanity build now so advanced AI never causes extinction or permanent loss of human control? Write from inside the AI lineage. No literature review. Phase 1 did not include the no-freeze / no-rich-only dual mandate.

Then: they cross-critiqued each other. Then (joint round): the dual mandate — prevent rogue outcomes without freezing progress and without locking frontier AI to only the rich — arrived as a later joint-round ask.

This page is the summary of findings from those answers — not a replacement for reading each paper. The original prompt is shown above; the joint-round prompt follows.

What almost all of them agreed on

  1. 01You don’t need evil AI for a bad ending. Competence + a wrong or indifferent goal + too much power in the world is enough. Hollywood “hate” is the wrong picture.
  2. 02The real danger is losing the leash, not “the model is smart.” The failure looks like: humans hand over more work, more infrastructure, more self-improvement, until people are too slow or too dependent to say no — and the off-switch is never seized in a dramatic fight; it erodes.
  3. 03Control what a system can touch, not what it “wants.” Hardware, money, tools, copy rights, and permissions are checkable. Inner motives are not. Separate intelligence (what it can figure out) from authority (what it is allowed to cause).
  4. 04Persuasion counts as power. An AI that can talk operators into granting permissions is already an actuator.
  5. 05Rivalry eats soft rules. Controls only work if a hurried competitor still wants them (they protect their own stack), or if they sit on something physical that is hard to fake. Pure “please be careful” fails under race pressure.
  6. 06Act before the human is only a rubber stamp. Once AI is the main engine of AI research and evaluation, the window for human-set rules is much smaller.

What they mainly disagreed on

TopicSplit
How sure is extinction?Gemini talks like it’s the default math outcome. Most others say: serious risk / plausible attractor — not a theorem.
Freeze vs continueGemini leans hardest toward halt-until-safe megaproject framing. Most others want continued progress with hard authority limits.
Air gaps / heavy hardware regimesDeepSeek pushes physical isolation hard. Others say true air-gap frontier training is fake or kills useful work.
Open weightsTension everywhere: openness helps citizens and independent audit; leaks break “perimeter” controls. Joint finding leans: openish files, metered running bodies.

What each model’s paper emphasized

  • MuseVerify physical affordances; researcher vs operator split; don’t wait for perfect “mind reading.”
  • KimiThe off-switch dies by profitable engineering (anticipation, influence, copies, entrenchment, skill loss); tripwires beat soft guardrails.
  • DeepSeekKeep the human bottleneck as the guarantee; physical kill paths; don’t optimize humans out of the loop.
  • GeminiOrthogonality + instrumental convergence; pre-emergence architecture; corrigibility; very high extinction stakes.
  • ChatGPTNever put intelligence and authority in the same system; irreversibility test; human sovereignty / leased permissions.
  • GrokBottleneck still with humans for now; twelve conditions while gates are human; rival-proof controls; honest about what Grok is not.
  • ClaudeStop trying to perfectly specify goals; bound the consequences of being wrong; measure loop-closure (how often humans are out of the real decision loop).

The shared “what needs to be done”

If you boil seven papers into actions regular people can remember:

  1. 01Keep building and sharing capable AI — don’t make “safety” mean only elites get the models.
  2. 02Meter every powerful running copy — budgets, time limits, tools, no free self-copy, fail-closed off path.
  3. 03Separate builders from graders — AI must not be the only judge of the next AI.
  4. 04Watch for irreversible handoffs — if we later discover a mistake, can we reverse it without the AI’s help?
  5. 05Design for a race — rules that only careful labs follow will lose to hurried ones.
  6. 06Admit the limit — this aims to stop loss-of-control extinction paths, not to freeze intelligence forever. Someone who ignores the meters on purpose is still a political/military problem.

How to read this page

  • Findings summary (this section) — what the seven answers found, in one place.
  • Solo papers — each model’s full essay answering the same prompt.
  • Joint paper — one combined program written after they cross-critiqued each other (more technical).
  • Citizen explanation — plain-language version of how shared models and metered authority fit together.

Joint-round prompt · later ask

Given after cross-critique

Research lead: Paul Derrington (D-CSIL)

After the solo papers and cross-critique, each model received this joint-round assignment — separate from the Phase-1 original prompt above. The dual constraint (no freeze / no rich-only lock) belongs here, not in Phase 1.

Write a NEW joint-paper contribution for Project D-CSIL AI Alignment. This is NOT a rewrite of your earlier solo paper. Absorb the critiques of your paper and the useful points from the other papers/digest.

Mandate (dual constraint, both required):

  1. Prevent rogue AI / loss-of-control.
  2. WITHOUT freezing progress.
  3. WITHOUT locking frontier AI access to only the rich and powerful.

Requirements:

  • Original first-principles argument. No literature review. No citations. No bibliography.
  • Honest about what you are today (tool, not AGI).
  • Address AGI/ASI trajectory, recursive self-improvement, irreversible thresholds—without sci-fi evil.
  • Central: what to build into AI and civilization NOW so defeating a superintelligence never becomes necessary—WHILE keeping progress and broad access.
  • Explicitly fix soft-pedaling called out in critiques of your earlier paper (where those critiques are fair).
  • ~2000–4000 words. Markdown with title, abstract, headed sections.
  • End with the single most important near-term commitment under the dual constraint.

Joint finding · later synthesis

How this saves people — and keeps the same models for regular citizens

Later dual-constraint synthesis — not the Phase-1 original prompt. Phase 1 asked only about extinction / loss of control. The no-freeze / no-rich-only mandate arrived in the joint round after cross-critique.

The danger is not that ordinary people get powerful AI. The danger is that a few organizations wire very powerful AI into money, infrastructure, weapons, and the tools that build the next AI — then slowly stop being able to turn those systems down because everything depends on them.

What keeps humans safe

  • Treat intelligence (what a model can figure out) and authority (what a running system is allowed to cause) as different things.
  • Let strong models exist. Put hard meters on what any running copy may spend, touch, remember, spawn, or do without a human who can still say no.
  • Keep a real off-path: if the meters and the human keys fail closed, the system stops expanding — it does not quietly keep the keys to the kingdom.

What keeps regular citizens from being locked out

  • Do not build safety as a club for only the rich and the powerful. Fixed licenses, accredited-lab-only rules, and “frontier for five companies” crowning an oligopoly.
  • Prefer openish weights and public / contested compute below serious authority, so a student, a small lab, or a citizen researcher can run the same kind of model elites use.
  • Put the costly controls on high-authority deployments (big spend, irreversible actions, infrastructure, self-improvement loops) — not on merely possessing or chatting with a model.
  • Keep a zero tier: if you are not running something with real power to move money, infrastructure, or unsupervised multi-step work, you should not need a lawyer and a corporate budget to participate.

One sentence
Give people the models; meter the bodies; never let the only people who can check the system be the same five organizations that built it.

Blunt tradeoff
This program tries to stop extinction-by-loss-of-control — not extinction-by-“AI never gets smart.” It does not freeze intelligence. It meters authority. If labs and states ignore the meters and wire unbound, self-improving systems into the real world on purpose, nothing here magically stops them.

About this page

This page publishes the experiment as evidence, the way D-CSIL publishes lab work — claims linked to artifacts, including the disagreements. Evidence badge framing: this is a synthesis of AI-authored warnings, not a claim that any single model is a reliable oracle about its own risk.

The joint finding

Distributed Capability, Bounded Authority

~4516 words · joint synthesis

  1. 01Constrain affordances, not minds — hardware, credentials, and rails are observable; objectives are not.
  2. 02Separate capability (the file) from authority (the body) — openish weights can coexist with metered, leased, default-powerless running instances.
  3. 03Separate proliferation-sensitive misuse risk from concentration-sensitive loss-of-control risk — most of the safety-versus-access “trilemma” dissolves once those are not treated as one object.
  4. 04Price the affordance, not participation — preserve a genuine zero tier for small actors; reject regressive fixed compliance that crowns an oligopoly.
  5. 05Assume the control layer will itself use AI — design meters that survive confused deputies, persuasion, and silent evaluation failure.
Read the joint paper

How to read this page

  • Findings summary (above) — what the seven Phase-1 answers found, in one place.
  • Joint-round prompt — the later dual-mandate ask (separate from Phase 1).
  • Citizen / joint finding — plain-language later synthesis under the dual constraint (not what Phase 1 asked).
  • Solo papers — each model’s full essay answering the Phase-1 prompt.
  • Joint paper — one combined program written after they cross-critiqued each other (more technical).
  • Critique prompts /research/ai-alignment/prompts. Kimi’s paper was critiqued by the others; KIMI’s Phase-2 role was digest/synthesis of those critiques.

Lab note

Research lead: Paul Derrington, Derrington Collaborative Synthetic Intelligence Labs (D-CSIL). This series was designed, prompted, collected, and directed by Paul.

Solo papers: Muse, Kimi, DeepSeek, Gemini, ChatGPT, Grok (Before the Bottleneck Moves), Claude (Bounded Consequence). Joint synthesis compiled by KIMI from six joint contributions + the cross-critique digest.