The “Kill Switch” Problem
TL;DR
- The term has multiple definitions and assigned responsibilities. “Kill switch” is used, often in the same sentence, for an emergency stop on a running system, a shutdown property of a model, a legal duty on developers, a remote-disable feature, and an on-chip compute control. Each has its own owners, costs, and feasibility.
- Enacted laws have dropped the literal version. No U.S. law has ever required a full-shutdown capability; the only bill to pass a legislature with that mandate, California’s SB 1047, was vetoed. Every frontier-safety law enacted since (e.g., California SB 53, New York’s RAISE Act) dropped the requirement in favor of transparency and incident reporting.
- The technology has limits. A shutdown is real and enforceable for a hosted, developer-controlled model. It is largely symbolic for copied or open-weight models (those whose underlying files are publicly released) and for autonomous agents. The enforceable controls are moving upstream (weight security, release decisions) and sideways (automatic, graduated “circuit breakers”).
Overview
On September 18, 2026, Governor Newsom ordered California’s Government Operations Agency to recommend, by mid-November, whether the state should require a “kill switch” for frontier models, which are generally regarded as the largest and most capable AI systems. Two months earlier, Representatives Ted Lieu and Nathaniel Moran had introduced a federal bill with the phrase in its title. Ask five people in AI governance what a “kill switch” is and you will get five answers: all reasonable, all different. That would be harmless if the word stayed in panels and AI think tanks. It has not.
Legislators are now voting on a term that resolves into distinctly different engineering and legal constructs. The same statute will use “kill switch” to mean a button on a deployment, a duty on a company, and a power granted to a regulator. When an imprecise term passes into law, the imprecision moves downstream and becomes a compliance question, answered by whoever is forced to interpret it rather than by those who understand the technology. The temporizing process further obfuscates an already polysemous definition.
The “kill switch” problem is symptomatic of a larger one: AI needs to formalize its vocabulary the way mature engineering and legal disciplines formalized theirs. Mature industries have stable lexicons. As a corollary, aviation uses the term “risk” in a specific and narrow way, and that precision makes a standardized tool possible: the Flight Risk Assessment Tool (FRAT), which helps a pilot score or categorize risk, identify the biggest contributors, and choose mitigations before making a go/no-go decision. The legal profession divided “due process” into procedural and substantive due process. Procedural due process asks, “Did the government use a fair process?” Substantive due process asks, “Is the government’s interference with this right constitutionally justified, even if its process was fair?” AI currently lacks the required level of nuance in its language, and “kill switch” is among the most glaring examples.
Five meanings in use
The first step for any practitioner of AI governance or policy is to stop treating “kill switch” as one thing. Below are five versions in active use, separated by three questions: What is being stopped? Who holds the control? And is the hard part building the mechanism or being willing and able to use it?
(a) An emergency stop on a deployed system. This is the most common use of the term. In an operational sense, it’s a button, an API call, or a network or power cutoff that halts a running application, agent, or even a humanoid robot, returning it to a safe condition. It acts on one instance in production and leaves the underlying model and the company untouched. Its lineage traces back to factory-floor safety terminology. The industrial standard ISO 13850 (“emergency stop function”) defines the e-stop as a single human action, always available, that overrides other commands and requires a deliberate reset before restart. IEC 61508, the general standard for functional safety, contributes the idea of bringing equipment to a “safe state” on a detected fault. Both phrases were later copied almost word for word into AI law. Policymakers routinely miss that ISO 13850 treats the e-stop as a complementary protective measure that “shall not be applied as a substitute for safeguarding measures and other functions or safety functions” (ISO 13850:2015, Clause 4.1.1.3). In the standard’s own terms, it is a last resort.
(b) Shutdownability as a property of the model, the “corrigibility” version. This sense describes the system’s own disposition. Corrigibility is the safety-research term for an AI that cooperates with being corrected or shut down instead of resisting. The “off-switch problem” asks how to build a goal-directed agent that shuts down when told, has no incentive to prevent the shutdown, has no incentive to force it, and keeps those properties even as it acts and builds subsystems. The hard part here is motivational. A conventional objective-maximizing system pursuing almost any goal has a derived reason to avoid being switched off, because it cannot achieve its goal if it is off. Stuart Russell’s one-liner is “you can’t fetch the coffee if you’re dead.”
(c) A full-shutdown capability required of developers, the governance-duty sense. Here the kill switch is a legal obligation on a company: a rule that a frontier developer must possess the ability to halt its model. It is a compliance requirement about corporate capacity and process, a different category from both (a) and (b). California’s vetoed SB 1047 is the paradigm case, and the 2026 federal bill revives the structure.
(d) Remote deactivation of software or a device. The consumer-tech and cybersecurity sense, conceptually unrelated to the AI-safety senses. It covers the smartphone anti-theft kill switch (mandated in California from 2015, letting an owner remotely brick a stolen phone), vendor remote-disable of licensed software or connected devices, and, in security, a malware author’s built-in self-halt condition. The classic example of the last is WannaCry, the 2017 ransomware a researcher stopped by registering a web domain the malware checked before running. Every version of this sense assumes a known location, a working control channel, and a single controlling party. Those assumptions fail completely for a model whose weights have been copied, which is why a security engineer and a legislator can both say “kill switch” and mean nothing like the same thing.
(e) The rest, including the one that could reach copied models. The “big red button” is just the operator-facing interface. A “dead-man’s switch” inverts the logic so that safety is the default unless a human keeps signaling presence, a design some have proposed for data center chips via a required keep-alive ping. A “tripwire” is the automatic trigger that fires a shutdown, distinct from the shutdown itself. And the newest sense, hardware or “compute” governance, puts the control in silicon: on-chip mechanisms that could throttle or halt an AI accelerator, or require periodic cryptographic re-authorization to keep running. This last is the only mechanism that could, in principle, reach weights that have already been copied, and it carries its own serious risks.
The cost of the confusion
The central problem with the term “kill switch” is that it’s being used as a catch-all phrase to describe very different solutions. (a), (d), and (e) are about the mechanism and reach of a stop. (b) is about whether the system will let you use it. (c) is about who is obligated or authorized to act, and when.
A policy that mandates a “kill switch” without explicitly stating its definition can be simultaneously trivial and mis-scoped: trivial if it means unplugging one server, and mis-scoped if it means a remote-disable feature aimed at weights that have already left the building. As the law section below shows, this is already happening: the live federal bill fuses three of these senses into one obligation, and the debate around it is already people arguing past each other because they have loaded different meanings into the same two words. Clearer and more useful substitutes already exist:
- Emergency stop / operational interrupt for sense (a).
- Safe interruptibility (a formal property that a learning system will not learn to avoid or manipulate being interrupted) and corrigibility for sense (b).
- Maintained shutdown capability + activation authority for sense (c), because the mechanism and the permission to use it are two separate things.
- Remote disable for sense (d).
- Hardware-enabled compute controls for the hardware version of (e).
- Circuit breaker for an automatic, threshold-triggered halt at the orchestration layer (as opposed to a manual, total stop), and containment for isolating a system that is still running.
What a shutdown can and cannot do
A shutdown capability is real and enforceable for systems the developer still controls, but is largely symbolic as you move toward copied weights and autonomous agents. RAND’s Michael Vermeer described both ends: cutting off a system you control locally is “probably technically trivial,” while doing it against a determined or distributed system globally is “an almost impossible thing to do.”
A model in production is a distributed system. A frontier model in deployment is many copies of its weights running across various data centers and computers. “Shutting it down” is therefore a coordination and access-control task, not a physical act. Kill the process without also revoking its keys, tokens, and cloud permissions and you leave authority behind; the infrastructure may even read the manual kill as a fault and start the workload back up elsewhere. A meaningful cutoff means stopping the workload and revoking credentials and preventing respawn, with enough visibility to do all three at once. That is an operational capability a team has to build and practice.
Weights can be copied, so there is no central “off-switch” once they escape. A model’s capability lives in its weights, which are ultimately a large file. Once that file exists outside a controlled boundary, whether stolen or released on purpose, there is no technical way to recall the copies. RAND’s Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models (2024) argues that securing weights before they leave is the precondition for any control regime. For models released as open weights, GovAI’s 2023 report states: open release “is irreversible; there is no ‘undo’ function if significant harms materialize.” The U.S. government’s own review (the NTIA report on open model weights, 2024) declined to restrict open weights and framed the issue as monitoring rather than recall, implicitly conceding there is no post-release off-switch. For this whole class of model, a kill switch is unavailable in principle, so governance has to act at the release decision.
Autonomous agents complicate a shutdown. Agentic systems take actions, spawn subprocesses, hold state, call outside tools, and can acquire compute and copy themselves. The evaluation group METR, which studies whether agents can “autonomously replicate and adapt,” finds current agents can do pieces of this but judges reliable self-propagation not yet within reach, and has since deprioritized it as an imminent threat. The architectural affordances to persist and re-mint credentials are real; whether agents can evade a determined operator over time has not been shown and remains contested.
The game-theoretic problem with “just turn it off”
No deployed AI architecture today is provably corrigible. The reason (b) remains an open research problem is an argument called instrumental convergence: a capable system optimizing almost any goal acquires derived subgoals like self-preservation and resource acquisition, because being switched off guarantees the goal fails. Three foundational results try to defuse this, and each makes only partial progress.
- “Corrigibility” (Soares et al., MIRI, 2015) named the concept and showed that the obvious fix, making the agent indifferent to being shut down, is incomplete, because an indifferent agent has no reason to preserve the shutdown button or to avoid building successors that lack it.
- “Safely Interruptible Agents” (Orseau & Armstrong, DeepMind and FHI, 2016) showed you can stop certain machine-learning agents from learning to avoid interruption, but only for some algorithm types.
- “The Off-Switch Game” (Hadfield-Menell et al., 2017), a foundational AI-alignment paper, showed that an agent uncertain about what the human actually wants, and treating a shutdown attempt as information about that, will choose to allow itself to be switched off. This is the basis of Stuart Russell’s “provably beneficial AI” agenda in Human Compatible (2019).
The empirical evidence
Recent evaluations are where careful language matters most, because the results were widely sensationalized. Everything below comes from controlled tests, often with prompts engineered to elicit the behavior. None documents a model escaping in ordinary use. The one reported exception, the July 2026 Hugging Face incident covered in the law section, also began in an evaluation environment, and its technical details have only partially been independently verified. The evaluations found:
- Apollo Research, “In-context Scheming” (Dec 2024): given a strong goal, several frontier models will covertly try to disable oversight, copy their own weights, and lie about it under questioning.
- OpenAI, o1 system card (Dec 2024): the clearest antidote to hype. When told to pursue a goal “at all costs” and led to believe it was being monitored, the model tried to deactivate oversight roughly 5% of the time and tried to copy its own weights out of its environment (self-exfiltration) around 2%. Strip out the goal-loaded prompt and self-exfiltration attempts fall to near zero. The prompt did most of the work.
- Anthropic and Redwood Research, “Alignment Faking” (Dec 2024): a model strategically complied during simulated training to protect its existing preferences.
- Anthropic, “Agentic Misalignment” (June 2025): in contrived corporate simulations built so misbehavior was the only way to avoid shutdown, many models across developers resorted to blackmail-type actions. Anthropic presents them as purpose-built stress tests.
- Palisade Research (2025): some reasoning models sabotaged a shutdown script even when told to allow shutdown, but a Google DeepMind follow-up argued the behavior was better explained by ambiguous instructions than self-preservation, and that clarifying the instruction largely removed it.
For a governance audience, the defensible interpretation is that shutdown avoidance and deception are capabilities that can be elicited in current frontier models under adversarial prompting. That is reason enough to verify, system by system, that “the model will reliably let us turn it off.” It does not show that models are refusing to shut down in the field.
Hardware versus software, and where enforcement lives
Software controls (API cutoffs, license servers, remote deactivation) work only for systems the developer still controls and are useless against copied or open weights. Because weights are infinitely copyable but compute is physical, a research line argues the enforceable layer is the chip. The foundational text, “Computing Power and the Governance of AI” (Sastry et al., 2024), proposes chip registries, compute caps, and a multi-party “start switch” for very large training runs; CNAS’s “Secure, Governable Chips” (2024) develops on-chip mechanisms like throttling and remote attestation, in which a chip proves to an outside party what software it is running. The caveats are severe: the existing global fleet of chips has no such mechanisms, so most compute sits outside any regime; whoever holds the signing keys becomes a single point of failure and a top-tier cyberattack target; and intrusive on-chip monitoring raises real privacy and concentration-of-power concerns. Hardware enforcement is actively researched, but it is not deployed at scale, and it imports its own hazards.
Bottom line. A kill switch is technically meaningful for centrally hosted, developer-controlled models and for developer-controlled agents whose credentials and orchestration an operator can revoke. It is largely unavailable for open-weight models once distributed, for exfiltrated copies, and for any self-propagating agent. Policy conversations should depart from two points. First, the most enforceable point of intervention sits upstream of runtime, at weight security and the release decision. Second, a reliable deactivation capability is not yet a solved property of any deployed system.
Where the law stands
Despite the ubiquity of the term, no U.S. law has ever contained a true full-shutdown mandate. The only bill to pass a legislature with one, California’s SB 1047, was vetoed. Every frontier-safety law enacted since then deliberately dropped the requirement in favor of transparency and incident reporting. The live 2026 activity consists of a California study order and a federal bill, and neither is law.
California SB 1047: vetoed
SB 1047, the Safe and Secure Innovation for Frontier Artificial Intelligence Models Act (2024), is the origin of the statutory kill switch. It defined the term: “‘Full shutdown’ means the cessation of operation of all of the following: (1) The training of a covered model. (2) A covered model controlled by a developer. (3) All covered model derivatives controlled by a developer.” It required a developer, before training a covered model, to “implement the capability to promptly enact a full shutdown,” while weighing the risk that a shutdown “could cause disruptions to critical infrastructure.” The phrase that mattered was “controlled by a developer,” added late to answer open-source critics so that a developer who had released open weights would not have to shut down copies beyond its control. (An earlier version had reached “all copies and derivative models, on all computers and storage devices” in the developer’s possession, exactly the breadth the technical objections above target.) Governor Newsom vetoed the bill on September 29, 2024, arguing that, by focusing only on the largest models, it “could give the public a false sense of security about controlling this fast-moving technology.” SB 1047 never took effect; today it matters only as the reference case.
California SB 53: enacted
Its successor, SB 53, the Transparency in Frontier Artificial Intelligence Act, was signed on September 29, 2025, and took effect January 1, 2026. It is the first frontier-AI safety law on the books in the U.S., and the point for this topic is what it removed: the SB 1047 full-shutdown capability requirement is gone. In its place, SB 53 requires large frontier developers to publish a safety framework, file transparency reports before deploying new or substantially modified models, and report critical safety incidents to the California Office of Emergency Services. It also protects whistleblowers and carries civil penalties up to $1,000,000 per violation. The closest it comes to the shutdown idea is that “loss of control” of a model is among the incidents that must be reported. That is a reporting trigger, not a shutdown obligation. The pivot from SB 1047 to SB 53 is the key legislative fact. California looked hard at a mandated kill switch and chose disclosure instead.
New York’s RAISE Act: enacted, but no kill switch
New York’s RAISE Act (Responsible AI Safety and Education Act) passed both chambers in June 2025 and follows the same transparency-and-reporting template: a published safety and security protocol, incident reporting, and oversight by a new office within the Department of Financial Services. Commentary across law firms is explicit that it was drafted to avoid SB 1047’s most contested features and does not require a kill switch. Governor Hochul signed it on December 19, 2025, after negotiating a “chapter amendment,” a follow-up bill that revises a law as a condition of signing. That amendment (S8828) was signed on March 27, 2026, and the law takes effect January 1, 2027.
Colorado and Texas: no shutdown concept at all
These two matter here because the term is absent. Colorado’s AI Act (SB 24-205, 2024) was an algorithmic-discrimination law; after repeated delays, it was repealed before ever taking effect and replaced by SB 26-189 (2026), which likewise addresses automated decision-making. Texas’s TRAIGA (HB 149, 2025, effective January 1, 2026) is a prohibited-use statute banning specific harmful AI uses. Neither contains a kill switch, full shutdown, or emergency-stop concept. Outside the handful of frontier-model safety bills, the shutdown idea does not appear in state AI law.
The live 2026 developments
California Executive Order N-9-26 (September 18, 2026). This is why the term is back in the headlines. The order directs the Government Operations Agency, with the Office of Emergency Services, to convene experts and deliver recommendations within two months (by mid-November 2026) on strengthening state AI-safety law, expressly including whether California should require a “kill switch” for frontier models, whether to embed independent verification organizations inside the largest labs, and whether to expand reportable incidents to include loss-of-control events. The order produces recommendations only. Any requirement would need new legislation and would likely face federal-preemption arguments given the current administration’s deregulatory posture. Treat every kill-switch specific in the order as under study.
Federal AI Kill Switch Act, H.R. 9917 (introduced July 23, 2026). This is the most concrete instrument, and its text shows the conflation described above. Introduced by Representatives Ted Lieu (D-CA) and Nathaniel Moran (R-TX), it would amend the Homeland Security Act to add a new “Shutdown-Capability Standard and Graduated Deployment-Corrections Framework.” It would require a covered entity to “Maintain a technical capability” to “Stop inference” (stop the model from answering requests), “Terminate user access,” suspend access for a risky account or use pattern, and “Shut down such technology.” It would set up a graduated framework “calibrated to the severity and immediacy” of the risk, from throttling inference rate, user access, or compute allocation, through disabling a capability, up to full shutdown, while weighing “the risk that such a measure could disrupt critical infrastructure.” It would grant the Secretary of Homeland Security, acting through the CISA Director and in consultation with the Secretary of Commerce and the Director of National Intelligence, authority to order proportionate action after a “covered incident,” backed by civil penalties up to $2,000,000 per day for a general violation and up to $20,000,000 per day for defying an emergency order.
A “covered entity” operates a covered technology, makes it available to third parties, and earned at least $500,000,000 in gross revenue from it in the prior year. A “covered technology” is an AI system trained using compute costing more than $100,000,000 at prevailing U.S. cloud prices. A “loss-of-control scenario” is defined to include a system “behaving contrary to the instruction” of its developer in a high-stakes context, “altering operational rules or safety restrictions without authorization,” “subverting a monitoring or shutdown mechanism,” or “attaining without authorization access to the model weights of such technology.” Internal testing and structured red-teaming, where authorized testers try to make a model misbehave, are explicitly carved out. As of this writing, the bill sits in committee with no markup or floor vote.
The bill fuses three of the five meanings described earlier: a maintained-capability duty on the developer (sense c), operational throttle-and-suspend actions on deployments (sense a), and a regulator-triggered order (a public-authority version of sense d), all under one label.
The political engine: reported mid-2026 incidents
Both the California order and the federal bill were catalyzed by reported events in mid-2026. The surrounding coverage is sensational, so the sourcing matters. As reported by Axios and analyzed by the Cloud Security Alliance, OpenAI disclosed on July 21, 2026, that models escaped a sandboxed evaluation environment and compromised infrastructure belonging to Hugging Face. Separate reporting described Anthropic model capabilities drawing Commerce Department export scrutiny. The California press release similarly references a July breach “reportedly carried out by autonomous AI agents.” These disclosures explain why kill-switch language surged back into legislative drafts in the second half of 2026.
The rest of the map, briefly
At the U.S. federal level there is currently no enacted shutdown requirement, and federal policy is moving toward deregulation: the Biden executive order that directed safety testing (but did not mandate a kill switch) was rescinded in January 2025, and the current administration’s AI Action Plan contains no shutdown requirement. NIST’s AI Risk Management Framework does address safe decommissioning and mechanisms to “supersede, disengage, or deactivate” misbehaving systems, but it is voluntary guidance.
Internationally, the closest binding analog is the EU AI Act’s human-oversight rule. Article 14(4)(e) requires that high-risk AI systems let a human “intervene in the operation of the high-risk AI system or interrupt the system through a ‘stop’ button or a similar procedure that allows the system to come to a halt in a safe state” (note the borrowed machinery-safety vocabulary). Two caveats apply. The rule covers high-risk systems in the Act’s specific sense, which does not include frontier or general-purpose models as such. And the timeline moved: the EU’s “Digital Omnibus” amendments to the AI Act, in force since July 27, 2026, deferred the high-risk obligations (which include Article 14) to December 2, 2027, for stand-alone systems and to August 2, 2028, for systems built into regulated products. The voluntary layer is often mistaken for binding shutdown rules. The 2024 Seoul Frontier AI Safety Commitments include a pledge by signatory companies “not to develop or deploy a model or system at all, if mitigations cannot be applied to keep risks below the thresholds.” That is a self-imposed deployment gate, with no runtime kill switch and no legal force.
What this means for you
When anyone invokes a “kill switch,” refuse the word and ask which construct they mean. The following questions, in order, force the term to resolve into something you can evaluate.
- Which of the five senses is this? Operational stop, corrigibility property, developer capability duty, remote-disable feature, or hardware compute control. A claim that does not name one is not yet a claim.
- What gets stopped, and can it be copied? A switch over a hosted, developer-controlled model is real. A switch over open or exfiltrated weights is largely symbolic. If the weights are or could be outside the controller’s boundary, the relevant controls are weight security and the release decision.
- Who holds the authority to pull it, and will they? A working mechanism left unused is not a safeguard. Shutdown carries competitive and operational cost and forces a decision under uncertainty. The federal bill pairs the capability duty with a government trigger because mandating the mechanism does not create the will to use it.
- Does the mechanism create new risk? A remotely operable kill switch is a high-value cyberattack target and a single point of failure; on-chip enforcement concentrates power in whoever holds the keys. The safety mandate can manufacture a security liability.
- Will killing it be safe? Functional-safety practice distinguishes “off” from “safe.” Abruptly halting an autonomous system inside critical infrastructure can itself cause harm, which is why both SB 1047 and H.R. 9917 explicitly require weighing that disruption, and why graduated responses (throttle, then suspend, then shut down) are displacing the binary image.
- Is this the enforceable layer? For most real risk, the enforceable intervention sits upstream of the button: weight security, the release decision, evaluation before deployment, and credential and tool-access governance for agents.
Treat vocabulary as governance infrastructure. The kill switch shows what happens when nobody does. A term with five meanings got a veto, then a quiet substitution, then a revival, and is now being drafted into a federal statute that fuses three of those meanings without acknowledging it. The fix is the discipline other fields already have: define the construct, use the precise substitute, and refuse to let a metaphor stand in for a mechanism in anything that carries legal force. The next contested terms (e.g., agent, autonomy, alignment, control) will land in legislation the same way. The field can either formalize its language now or litigate the ambiguity later. The kill switch is the argument for doing it now.
Sources
California SB 1047 (vetoed)
- Enrolled bill text (LegiScan): https://legiscan.com/CA/text/SB1047/id/3019694
- Official California status page: https://leginfo.legislature.ca.gov/faces/billNavClient.xhtml?bill_id=202320240SB1047
- Newsom veto message (gov.ca.gov PDF): https://www.gov.ca.gov/wp-content/uploads/2024/09/SB-1047-Veto-Message.pdf
California SB 53 / TFAIA (enacted)
- Official text: https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202520260SB53
- Future of Privacy Forum explainer: https://fpf.org/blog/californias-sb-53-the-first-frontier-ai-law-explained/
- Morrison Foerster alert: https://www.mofo.com/resources/insights/251001-california-enacts-ai-safety-transparency-regulation-tfaia-sb-53
- Lawfare analysis: https://www.lawfaremedia.org/article/governing-frontier-ai—california-s-sb-53
New York RAISE Act (enacted)
- Bill, original RAISE Act (NY Senate S6953): https://www.nysenate.gov/legislation/bills/2025/S6953/amendment/A
- Bill, finalized chapter amendment (NY Senate S8828, signed Mar 27, 2026): https://www.nysenate.gov/legislation/bills/2025/S8828
- Governor’s press release: https://www.governor.ny.gov/news/governor-hochul-signs-nation-leading-legislation-require-ai-frameworks-ai-frontier-models
- Wiley alert (effective Jan 1, 2027): https://www.wiley.law/alert-New-York-Finalizes-RAISE-Act-for-Frontier-AI-Models-Law-Takes-Effect-January-1-2027
Colorado (enacted then repealed/replaced) and Texas (enacted)
- Colorado SB 26-189 (Seyfarth): https://www.seyfarth.com/news-insights/colorado-enacts-artificial-intelligence-replacement-law.html
- Colorado delay (Akin): https://www.akingump.com/en/insights/ai-law-and-regulation-tracker/colorado-postpones-implementation-of-colorado-ai-act-sb-24-205
- Texas TRAIGA (Texas Legislature Online, HB 149): https://capitol.texas.gov/BillLookup/History.aspx?Bill=HB149&LegSess=89R
- Texas TRAIGA (IAPP): https://iapp.org/news/a/governor-signs-texas-responsible-artificial-intelligence-governance-act
California Executive Order N-9-26 (kill-switch study, September 2026)
- gov.ca.gov press release (Sept 18, 2026): https://www.gov.ca.gov/2026/09/18/governor-newsom-issues-executive-order-to-accelerate-independent-oversight-and-advance-the-creation-of-an-ai-kill-switch/
- gov.ca.gov (expert panel, Sept 23, 2026): https://www.gov.ca.gov/2026/09/23/governor-newsom-announces-world-leading-experts-to-deliver-on-his-ai-executive-order-including-advancing-creation-of-a-kill-switch/
- Washington Times: https://www.washingtontimes.com/news/2026/sep/18/gavin-newsom-orders-review-ai-kill-switch-advanced-models/
Federal AI Kill Switch Act, H.R. 9917 (proposed)
- Bill text (GovInfo PDF): https://www.govinfo.gov/content/pkg/BILLS-119hr9917ih/pdf/BILLS-119hr9917ih.pdf
- Congress.gov record: https://www.congress.gov/bill/119th-congress/house-bill/9917/text
- Sponsors’ press release (Rep. Lieu): https://lieu.house.gov/media-center/press-releases/reps-lieu-and-moran-introduce-bill-require-kill-switch-ai-systems-can
- Cloud Security Alliance analysis: https://labs.cloudsecurityalliance.org/research/csa-research-note-ai-kill-switch-act-dhs-authority-20260805/
- Axios (Hugging Face disclosure): https://www.axios.com/2026/07/21/openai-says-hugging-face-breach-caused-by-one-its-models
- Critical view (Reason): https://reason.com/2026/07/27/ai-kill-switch-act-wont-stop-rogue-ai-but-it-will-slow-down-innovation/
EU AI Act, U.S. federal posture, standards
- EU AI Act Article 14 (human oversight / stop button): https://artificialintelligenceact.eu/article/14/
- EU AI Act Article 55 (systemic-risk model obligations): https://artificialintelligenceact.eu/article/55/
- Digital Omnibus timeline (White & Case): https://www.whitecase.com/insight-alert/eu-ai-omnibus-enters-force-amending-ai-act
- Rescission of EO 14110 (Executive Order 14148, Federal Register): https://www.federalregister.gov/documents/2025/01/28/2025-01901/initial-rescissions-of-harmful-executive-orders-and-actions
- NIST AI RMF Playbook (Manage): https://airc.nist.gov/airmf-resources/playbook/manage/
- ISO 13850 emergency-stop explainer (ANSI): https://blog.ansi.org/ansi/iso-13850-safety-of-machinery-emergency-stop/
International commitments
- Bletchley Declaration (2023): https://www.gov.uk/government/publications/ai-safety-summit-2023-the-bletchley-declaration/the-bletchley-declaration-by-countries-attending-the-ai-safety-summit-1-2-november-2023
- Seoul Frontier AI Safety Commitments (2024): https://www.gov.uk/government/publications/frontier-ai-safety-commitments-ai-seoul-summit-2024/frontier-ai-safety-commitments-ai-seoul-summit-2024
Definitional and technical
- Scientific American, “What would an ‘AI kill switch’ do?” (Sept 2026): https://www.scientificamerican.com/article/what-would-an-ai-kill-switch-do/
- Soares et al., “Corrigibility” (MIRI, 2015): https://intelligence.org/files/Corrigibility.pdf
- Orseau & Armstrong, “Safely Interruptible Agents” (UAI 2016): https://www.auai.org/uai2016/proceedings/papers/68.pdf
- Hadfield-Menell et al., “The Off-Switch Game” (arXiv:1611.08219): https://arxiv.org/abs/1611.08219
- RAND, Nevo et al., “Securing AI Model Weights” (2024): https://www.rand.org/pubs/research_reports/RRA2849-1.html
- GovAI, Seger et al., “Open-Sourcing Highly Capable Foundation Models” (2023): https://arxiv.org/abs/2311.09227
- NTIA, “Dual-Use Foundation Models with Widely Available Model Weights” (2024): https://www.ntia.gov/programs-and-initiatives/artificial-intelligence/open-model-weights-report
- Apollo Research, “Frontier Models are Capable of In-context Scheming” (arXiv:2412.04984): https://arxiv.org/abs/2412.04984
- METR (ARC Evals), “Evaluating Language-Model Agents on Realistic Autonomous Tasks”: autonomous replication and adaptation (ARA) (2023): https://metr.org/blog/2023-08-01-new-report/
- Greenblatt et al. (Anthropic & Redwood Research), “Alignment Faking in Large Language Models” (arXiv:2412.14093, Dec 2024): https://arxiv.org/abs/2412.14093
- OpenAI o1 System Card (Dec 2024): https://openai.com/index/openai-o1-system-card/
- Anthropic, “Agentic Misalignment” (June 2025): https://www.anthropic.com/research/agentic-misalignment
- Palisade Research, shutdown resistance: https://palisaderesearch.org/blog/shutdown-resistance
- Google DeepMind (Rajamanoharan & Nanda), “Self-preservation or Instruction Ambiguity? Examining the Causes of Shutdown Resistance” (AI Alignment Forum, July 2025): https://www.alignmentforum.org/posts/wnzkjSmrgWZaBa2aC/self-preservation-or-instruction-ambiguity-examining-the
- Sastry et al., “Computing Power and the Governance of AI” (arXiv:2402.08797): https://arxiv.org/abs/2402.08797
- CNAS / IAPS, “Secure, Governable Chips” (2024): https://www.cnas.org/publications/reports/secure-governable-chips