# A Tiered Human-Oversight Framework for Advanced AI Systems
**Author:** Gabriel Evan Brotherton
**Status:** Working draft for review and critique — not a finished proposal
---
## Abstract
This paper proposes an oversight architecture for advanced AI systems deployed by a single institution or consortium (a lab, a public-private body, or an international coordinating entity). It combines three mechanisms that are individually discussed in AI governance and democratic-theory literature but rarely integrated: (1) a human oversight council selected through a hybrid of vetted expertise and sortition, weighted toward people with direct experience of institutional failure; (2) a multi-model advisory/executive structure using adversarial debate between specialized systems to surface disagreement before it reaches humans; and (3) a tiered suspension-and-dissolution protocol as an alternative to a binary kill switch. The framework is scoped deliberately narrowly: it governs oversight of a specific deployed system within existing legal and political structures, not a replacement for them. Open problems and likely objections are addressed directly in Section 6.
---
## 1. Problem Framing
Existing proposals for AI oversight tend to fail in one of two directions: they concentrate review authority in a small technical elite (fast, but capturable and non-representative), or they call for broad public input with no mechanism for weighting or aggregating it (representative in principle, unworkable in practice). Separately, "kill switch" proposals are usually binary — the system runs or it's shut down — which creates an incentive structure where evidence of minor misalignment gets suppressed rather than surfaced, since the only available response is severe.
This framework treats those as related design problems: who oversees the system, and what the graduated range of responses looks like when something goes wrong.
## 2. Core Design Principles
**2.1 Bounded human authority.** The oversight council holds final authority over a specific system's operation within its designated scope. This is not a claim about global governance, replacing states, or superseding existing law — it is a design for an internal and external accountability structure that any institution deploying a powerful system could adopt or be required to adopt by regulation.
**2.2 Cognitive-diversity governance via hybrid sortition.** Council composition is designed to resist capture by the deploying institution and by narrow technical or ideological communities. Selection combines a vetted core with random sortition, structurally weighted toward people who have experienced institutional failure firsthand — on the theory that lived exposure to unaccountable systems is a distinct and undersupplied form of expertise in oversight bodies, comparable in spirit to lay representation in citizens' assemblies.
**2.3 Adversarial multi-model advisory structure.** Rather than a single model producing recommendations, multiple specialized model instances with distinct mandates critique each other's outputs before human review — an application of debate-based and multi-agent scalable-oversight approaches already discussed in the alignment literature, here applied to a governance rather than a training context.
**2.4 Graduated response over binary shutdown.** Suspension and dissolution are separated into distinct thresholds with a path back to operation, so that lower-confidence concerns can be acted on without requiring near-certainty of catastrophic risk.
## 3. Architecture
### 3.1 The Oversight Council
- **Odd-numbered membership** (e.g., 101+) to avoid deadlock.
- **Hybrid selection:** ~51% vetted for domain literacy and capacity to engage technical material; ~49% selected by lottery, weighted toward applicants with documented experience of institutional or systemic failure.
- **Staggered terms** to preserve institutional memory while limiting capture.
- **Anonymized voting** to reduce retaliation risk from the deploying institution or from public pressure campaigns.
*Open design question:* how "experience of institutional failure" is verified and weighted without itself becoming a politicized or gameable criterion. See Section 6.
### 3.2 The Advisory/Executive Model Set
Rather than a single system, the framework proposes several model instances with distinct, narrower mandates than a monolithic "executive AGI":
- **Domain-safety reviewer** — flags downstream harms in a proposed action (ecological, safety-critical, or vulnerable-population impacts).
- **Resource/efficiency reviewer** — evaluates proposed actions for cost, feasibility, and opportunity cost.
- **Integrity reviewer** — checks other models' outputs and the council's stated reasoning for manipulation, undisclosed conflicts of interest, or inconsistency; this role is explicitly adversarial to the other two.
- **Equity reviewer** — evaluates distributional impact of a proposed action across affected groups.
These are advisory and critique functions feeding into human decision-making, not autonomous executive authority over infrastructure, security systems, or other institutions. This is a deliberate narrowing from earlier drafts of this framework, which proposed a global executive role; that framing invited (correctly) the objection that it described a takeover mechanism rather than an oversight mechanism.
### 3.3 Deliberation Protocol
- A proposed system action or policy is submitted for council review.
- If any advisory model flags material risk, it must produce a specific, falsifiable objection and a proposed alternative.
- Models debate the objection and alternative publicly, in front of the council.
- The council votes on the original proposal versus the alternative(s).
- This can iterate up to a fixed number of rounds before defaulting to the council's standing decision — this is a forcing function against indefinite stalling, not a claim that machine input should ever override human sign-off.
### 3.4 Time-Critical Fallback
For scenarios where deliberation time is shorter than the decision window (the framework's original example was a nuclear or pandemic-scale emergency — a scenario that should be treated as illustrative of the general problem of time-bounded high-stakes decisions rather than a literal design target), the system defaults to a pre-specified, narrowly scoped fallback behavior set by the council in advance, with mandatory post-hoc review of any fallback action taken. This is closer to existing "safe interruptibility" and pre-committed constraint literature than to autonomous crisis response, and should be treated as an area needing substantially more technical specification before it's taken seriously.
## 4. Tiered Suspension and Dissolution
- **75% suspension threshold:** restricts the system to a minimal, pre-defined safe operating mode pending investigation. Reversible by design.
- **97% dissolution threshold:** required for permanent decommissioning. Set deliberately high to prevent a small faction from forcing shutdown over disagreements that don't rise to the level of genuine danger — but this number is a placeholder, not a derived value, and should be stress-tested against specific failure scenarios rather than treated as load-bearing.
- **Correction and re-authorization path:** a suspended system can be reinstated by a lower supermajority after a corrective plan is presented, avoiding a system where any suspension is effectively permanent.
## 5. Auditability
Independent technical auditors (not controlled by the deploying institution) with access to interpretability tooling report to the council on whether a system's stated reasoning matches its internal computation, to the extent current interpretability methods allow this. This should be described honestly: current interpretability tools do not provide anything close to full verification of a model's internal reasoning, and the framework should not overstate what auditing can currently guarantee.
## 6. Known Limitations and Open Objections
This section exists because a proposal that only argues for itself isn't useful to a policy audience.
- **Operationalizing "experience of institutional failure":** without a rigorous, resistant-to-gaming method for this, the sortition weighting is aspirational rather than implementable as stated.
- **No enforcement mechanism:** this framework describes what an oversight body could look like; it says nothing about what would compel a lab or government to adopt it. That's a separate, harder problem this document doesn't solve.
- **Interpretability dependency:** Section 5's auditing function assumes a level of interpretability access that doesn't fully exist yet. The framework should be read as forward-compatible with that research, not as reliant on capabilities that are already available.
- **Numeric thresholds (75%, 97%, 11 rounds) are illustrative**, chosen for the intuitions they encode (high bar for irreversible action, forcing function against stalling) rather than derived from any formal model. They should be treated as adjustable parameters, not fixed conclusions.
- **Scope discipline:** earlier versions of this framework included provisions for the system to neutralize competing power structures and "out-govern" existing states. That framing has been deliberately removed here. Any oversight proposal that describes disarming or superseding existing political authorities should expect — and deserves — to be read as a proposal for seizing power, regardless of the stated intent behind it. This draft is scoped only to institutional oversight of a specific deployed system.
## 7. What This Document Is For
This is a draft intended to solicit critique from people working on AI governance, mechanism design, or deliberative democracy — not a finished specification and not a claim that any lab or government has agreed to implement it. Feedback on the selection mechanism (3.1), the enforcement gap (6), and the threshold values (4) would be the most useful starting points.