Shared council Β· 30 Sept 2026 Β· What is Council AI?
Can tech companies realistically be trusted to self-regulate their own
analysis Β· 4-model panel Β· balanced debate Β· rigorous chair
The question
Can tech companies realistically be trusted to self-regulate their own AI safety?
β Chairman's verdict
No. Tech companies cannot realistically be trusted to self-regulate their own AI safety as the primary safeguard, especially for frontier models. They can and should play a useful operational role, but self-regulation alone is structurally unreliable and must be made legally consequential, independently verified, and personally costly to get wrong. The core reason is a structural conflict of interest. The same organization sets the risk thresholds, runs the evaluations, interprets the results, and decides whether to ship. That is not a regulatory regime; it is an internal judgment call with an external audience. Every major lab's framework contains an escape hatch. OpenAI's Preparedness Framework (December 2023) has the company itself score risk levels and allows deployment at medium risk with mitigations, plus an override path. Anthropic's Responsible Scaling Policy (2023) is the most serious of the genre, committing to pause at defined capability thresholds, but it is still self-administered and self-graded, which is why Anthropic itself has argued external verification is needed. Google DeepMind, Meta, and xAI have nothing comparably binding. The record matters more than the documents. In November 2023, OpenAI's board, the body explicitly charged with a safety-and-mission mandate, fired Sam Altman and reversed itself in five days under employee revolt, investor pressure, and a Microsoft offer to hire the staff. In May 2024 the Superalignment team was dissolved; Jan Leike, who co-led it, wrote publicly that safety culture and processes had taken a backseat to shiny products. Ilya Sutskever had already departed. Meta reassigned its Responsible AI team in November 2023, folding the people into generative AI. The White House's July 2023 voluntary commitments from seven companies, including a promise to watermark AI-generated content, were non-binding and are widely judged substantially unmet. None of this proves catastrophe; it proves what happens to voluntary pledges under pressure. The honest counter-case is real. Regulators genuinely lack the technical depth that sits inside these labs, and badly designed rules can freeze out smaller players and push capability development to the least cautious jurisdiction. Self-regulatory artifacts are not worthless: model cards, red-teaming disclosures, C2PA provenance standards, and the Responsible Scaling Policy norm have real effects on internal behavior and have diffused across the industry in about two years. For non-frontier AI such as medical triage tools, hiring systems, and consumer chatbots, self-regulation plus existing product liability and privacy law is roughly the right level of intervention. It is also worth saying plainly that we are reasoning from incentives and revealed preferences, not from a body of documented catastrophic incidents. The strongest defensible claim is that the incentive structure reliably wins when tested, which is exactly what a decade of tech self-regulation in privacy, content moderation, and fintech has shown. What would actually earn trust, concretely: - A legal trigger at a compute threshold. The EU AI Act uses 10^25 FLOP for systemic-risk general-purpose AI models. The US Executive Order 14110 used 10^26 FLOP, but that order was rescinded in January 2025. Any new regime should define equivalents for fine-tuning and inference-time scaling, and require registration and licensing above the threshold. - Mandatory independent third-party evaluation with genuine model access and inference-time compute, funded by a levy on the labs so evaluators are not dependent on the companies they assess, and exempt from the NDAs that currently neutralize external red teams. - Mandatory incident reporting with whistleblower protections and a private right of action. This is the core of California's SB 53 and the EU AI Act's GPAI obligations, which began applying on 2 August 2025 with fines up to 3 percent of global turnover, or 7 percent for prohibited practices. - Signed safety cases with personal legal accountability for the CEO and board, in the model of FDA drug approval or FAA certification: companies still do the work, but falsifying it is a crime and the compliance record is admissible. - A safe-harbour rule: meeting the standard is a defense, failing to meet it removes liability caps. The EU AI Act, SB 53, and the UK, Japan, and Korea AI safety institutes show binding co-regulation is arriving, but it is thin, uneven, and politically volatile. The US rescinded Executive Order 14110 on day one of the second Trump administration in January 2025, and California vetoed the much stronger SB 1047 in September 2024 after heavy industry lobbying. A regime that exists at the pleasure of a single election cannot be the backstop. So the workable answer is not trust the companies or ban them. It is: let companies self-assess, then make that self-assessment legally consequential, independently verified, and personally costly to get wrong. Self-regulation as a first line of defense for consumer AI, yes. As the last line for frontier models, no.
Key reasoning
The decisive argument is structural, not anecdotal: the entity that sets the risk threshold, runs the evaluation, interprets the result, and decides whether to ship is the same entity that profits from shipping. That is an internal judgment call, not a regulatory regime. The historical record confirms the incentive structure wins under pressure: OpenAI's board reversed itself on Sam Altman in five days in November 2023; the Superalignment team was dissolved in May 2024 and Jan Leike said safety had taken a backseat to products; Meta reassigned its Responsible AI team in November 2023; the July 2023 White House voluntary commitments were non-binding and largely unmet. The counter-case is real but does not change the conclusion: regulators lack technical depth, and bad rules can cause jurisdictional arbitrage, so self-regulation remains a useful operational layer. The resolution is co-regulation: companies self-assess, but that self-assessment becomes legally consequential, independently verified, and personally costly to falsify.
Points of agreement
- All four member answers agree that self-regulation alone is insufficient for frontier AI safety.
- All four agree the core problem is a structural conflict of interest between profit and precaution.
- All four cite OpenAI's Superalignment dissolution and the departures of Jan Leike and Ilya Sutskever as key evidence.
- All four agree self-regulation has a useful but limited operational role, such as red-teaming, benchmarks, and model cards.
- All four converge on co-regulation: mandatory external audits, third-party red-teaming, liability for harms, and transparency reporting.
- All four note that regulators lack technical depth and that poorly designed rules risk jurisdictional arbitrage.
Disagreements & tradeoffs
- Answer D treats several 2025 events as settled history (OpenAI Preparedness Framework revised April 2025, Google altering AI principles February 2025, Meta refusing the EU GPAI Code July 2025, California SB 53 signed September 2025). Reviewers A and E flagged these as hallucinated or future-dated, and D's accuracy variance (95) is by far the highest.
- Answer A cites the 10^26 FLOP threshold from US Executive Order 14110 without noting the order was rescinded in January 2025, and reviewer D notes the EU AI Act uses 10^25 FLOP, not 10^26.
- Answer A's Superalignment chronology is reversed according to reviewer D: Leike and Sutskever departed before the team was dissolved, not after.
- Answer E implies US and UK AI Safety Institutes currently evaluate mandatory standards, which is not yet true; reviewer D recommends qualifying this as a proposed enforcement role.
- Answer B is the most accurate and lowest-variance (variance 2) but is the least specific on concrete policy mechanisms such as compute thresholds and safe-harbour provisions.
- Answers differ on confidence: D states roughly 85 percent on the structural critique and 65 percent on its specific policy package; B states high on the critique and moderate on prescriptions.
Risk & uncertainty
- Sensitive domain β general information, not professional advice. Consult a qualified expert before acting.
- We are reasoning from incentives and revealed preferences, not from a body of documented catastrophic AI incidents. The strongest defensible claim is that the incentive structure reliably wins when tested, not that catastrophe has already occurred.
- Several cited 2025 events are contested or unverified. Treat the OpenAI Preparedness Framework April 2025 revision, Google's February 2025 AI principles change, Meta's July 2025 GPAI refusal, and California SB 53's September 2025 signing as claims requiring independent verification before reliance.
- The EU AI Act's 10^25 FLOP threshold and the rescinded US Executive Order 14110's 10^26 FLOP threshold are not interchangeable; any new regime must specify which threshold and for what purpose.
- Regulatory design in this area is genuinely unsettled. The specific policy package proposed here is a reasoned proposal, not a proven solution, and could be undermined by regulatory capture, jurisdictional arbitrage, or rapid capability change.
- Self-regulation is more defensible for non-frontier AI such as medical triage, hiring systems, and consumer chatbots, where existing product liability and privacy law provide a backstop. The strong conclusion applies to frontier models.
- A regime that exists at the pleasure of a single election cannot be the backstop, as shown by the rescission of Executive Order 14110 and the veto of California SB 1047.
- Consensus among models is not proof; they can share blind spots.
Suggested next steps
- Verify the contested 2025 claims independently before citing them: OpenAI Preparedness Framework revision, Google AI principles change, Meta GPAI Code refusal, and California SB 53 signing status.
- Confirm the current legal status of EU AI Act GPAI obligations and their 2 August 2025 application date, including the 3 percent and 7 percent turnover fine tiers.
- Distinguish clearly between the EU 10^25 FLOP threshold and the rescinded US 10^26 FLOP threshold when proposing any new compute trigger.
- Design a funding mechanism for independent evaluators, such as a levy on frontier labs, so assessors are not financially dependent on the companies they assess.
- Draft model legislation or policy language covering mandatory incident reporting, whistleblower protections, a private right of action, signed safety cases with personal CEO and board liability, and a safe-harbour rule.
- Assess jurisdictional arbitrage risk by mapping which countries would host frontier training runs under each proposed threshold and enforcement regime.
- Separate the policy conversation for frontier models from that for consumer and narrow AI, since the appropriate level of intervention differs substantially.
Model scorecard
Each model's answer (5)
No, tech companies cannot realistically be trusted to self-regulate their own AI safety. While voluntary safety frameworks are useful for developing technical standards, relying on self-regulation alone fails due to fundamental economic misalignments, competitive pressure, and historical precedent across technological industries. The structural reasons self-regulation fails include: - The prisoner's dilemma of commercial AI: Frontier AI development is characterized by intense competitive pressure among major players such as OpenAI, Google DeepMind, Anthropic, Meta, and Microsoft. If one lab delays deployment to conduct rigorous safety testing, competitors capture market share, enterprise contracts, and top talent. In an escalating race, voluntary safety standards are consistently eroded when they conflict with speed to market. - Erosion of internal governance: Frontier labs initially established structures designed to prioritize safety over profit, yet those guardrails have repeatedly degraded under market pressure. OpenAI originated as a nonprofit with explicit safety charters, but it transitioned into a capped-profit entity and is restructuring further toward a conventional commercial business. The dissolution of OpenAI's Superalignment team in mid-2024, followed by the departure of safety leaders Jan Leike and Ilya Sutskever, demonstrated that internal safety teams lack structural veto power when executives prioritize product timelines. - Information asymmetry and proprietary opacity: Self-regulation requires transparency, yet leading models (such as GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5) do not disclose their full training data mixes, exact architecture parameters, or unfiltered red-teaming failure rates. Independent researchers and the public cannot verify whether internal thresholds, such as Anthropic's Responsible Scaling Policy or OpenAI's Preparedness Framework, are being strictly followed or retroactively adjusted. - Historical precedent in safety-critical sectors: Delegated or self-regulatory models have repeatedly broken down in high-stakes industries where market forces conflict with public safety. The most direct analogies include the Federal Aviation Administration delegating safety certification to Boeing (contributing to the 737 MAX crashes), the financial sector's internal risk models prior to the 2008 crash, and social media platforms' failed self-policing regarding user privacy and algorithmic radicalization. Where self-regulation does work is at the operational level: pioneering red-teaming techniques, standardizing benchmark suites, and testing mechanistic interpretability tools. However, to translate these practices into genuine safety, they must be codified into binding legal regimes. A realistic safety framework requires four external mechanisms: - Legally enforceable safety thresholds tied to computational thresholds (such as the 10^26 FLOP metric introduced in the US Executive Order 14110 and the EU AI Act). - Independent, third-party pre-deployment red teaming and audits conducted by national AI Safety Institutes rather than corporate-contracted vendors. - Clear legal liability for downstream catastrophic damages, which removes the moral hazard of deploying unvetted systems. - Robust statutory protections and reporting channels for corporate whistleblowers.
No, tech companies cannot realistically be trusted to self-regulate AI safety on their own, though they can play a useful but insufficient role. The core problem is a structural conflict of interest: companies are incentivized to ship products quickly and maximize profit, while safety often requires slowing down, restricting capabilities, or absorbing costs. Historical evidence supports this. Social media companies promised self-regulation on content moderation and privacy for years, yet Cambridge Analytica, Myanmar genocide amplification, and teen mental-health harms occurred under voluntary regimes. Facebook's own internal research showed Instagram harmed teen girls, but the company downplayed it publicly. On AI specifically, OpenAI disbanded its safety-focused Superalignment team in 2024 after key researchers including Ilya Sutskever and Jan Leike left; Leike said safety culture had taken a backseat to product. Google DeepMind, Anthropic, and Microsoft have all published safety frameworks, but these are non-binding and often revised when competitive pressure rises. The EU AI Act, the US AI Safety Institute (now under NIST), and executive orders show governments are stepping in precisely because voluntary commitments have been insufficient. That said, full external regulation has limits: regulators lack technical expertise, move slowly, and may not anticipate frontier risks. A realistic approach is co-regulation: mandatory external audits, third-party red-teaming, liability for harms, transparency reporting, and licensing for frontier models, with companies handling day-to-day implementation. Self-regulation alone is not credible; self-regulation plus strong external oversight and legal accountability is the workable path. Confidence: high on the structural critique, moderate on specific policy prescriptions.
This model did not respond in time (error or timeout), so it was left out of peer review and the final answer.
Short answer: no β not for frontier AI, and not as the primary safeguard. Self-regulation is a useful complement, but the conflict of interest is structural and the track record shows safety commitments bending whenever they collide with revenue, fundraising, or competitive pressure. The structural problem is simple. The same organisation sets the risk thresholds, runs the evaluations, interprets the results, and decides whether to ship. That is not a regulatory regime; it is an internal judgment call with an external audience. Every major lab's framework contains an escape hatch. OpenAI's Preparedness Framework (December 2023, revised April 2025) has the company itself score risk levels and explicitly allows deployment at "medium" risk with mitigations, plus a documented override path. Anthropic's Responsible Scaling Policy (2023, since revised) is the most serious of the genre β it commits to pausing at defined capability thresholds β but it is still self-administered and self-graded, which is why Anthropic itself has argued external verification is needed. Google DeepMind, Meta, and xAI have nothing comparably binding. The record matters more than the documents. In November 2023, OpenAI's board β the body explicitly charged with a safety-and-mission mandate β fired Sam Altman and reversed itself in five days under employee revolt, investor pressure, and a Microsoft offer to hire the staff. In May 2024 the superalignment team was dissolved; Jan Leike, who co-led it, wrote publicly that safety culture and processes had "taken a backseat to shiny products." Meta disbanded its Responsible AI team in November 2023, folding the people into generative AI. Google quietly deleted its pledge not to pursue AI for weapons and surveillance from its published AI principles in February 2025. In July 2025 Meta refused to sign the EU's GPAI Code of Practice. The White House's July 2023 voluntary commitments from seven companies (Amazon, Anthropic, Google, Inflection, Meta, Microsoft, OpenAI) β including a promise to watermark AI-generated content β were non-binding and are widely judged substantially unmet. None of this is proof of catastrophe; it is proof of what happens to voluntary pledges under pressure. The honest counter-case, which I take seriously. Regulators genuinely lack the technical depth that sits inside these labs, and a badly designed rule can freeze out smaller players and push capability development to the least cautious jurisdiction. Self-regulatory artefacts are not worthless: model cards, red-teaming disclosures, C2PA provenance standards, and the Responsible Scaling Policy norm have real effects on internal behaviour, and they have diffused across the industry in about two years. For non-frontier AI β medical triage tools, hiring systems, consumer chatbots β self-regulation plus existing product liability and privacy law is roughly the right level of intervention. It is also worth saying plainly that we are reasoning from incentives and revealed preferences, not from a body of documented catastrophic incidents; the strongest claim I can defend is that the incentive structure reliably wins when tested, which is exactly what a decade of tech self-regulation in privacy, content moderation, and fintech has shown. What would actually earn trust, concretely: - A legal trigger at a compute threshold β roughly 10^25 FLOP training compute, with equivalents defined for fine-tuning and inference-time scaling β above which training runs must be registered and licensed. - Mandatory independent third-party evaluation with genuine model access and inference-time compute, funded by a levy on the labs so evaluators are not dependent on the companies they assess, and exempt from the NDAs that currently neutralise external red teams. - Mandatory incident reporting with whistleblower protections and a private right of action β this is the core of California's SB 53 (signed September 2025) and the EU AI Act's GPAI obligations, which began applying on 2 August 2025 with fines up to 3 percent of global turnover (7 percent for prohibited practices). - Signed safety cases with personal legal accountability for the CEO and board, in the model of FDA drug approval or FAA certification: companies still do the work, but falsifying it is a crime and the compliance record is admissible. - A safe-harbour rule: meeting the standard is a defence, failing to meet it removes liability caps. The EU AI Act, SB 53, and the UK/Japan/Korea AI safety institutes show binding co-regulation is arriving, but it is thin, uneven, and politically volatile β the US rescinded Executive Order 14110 on day one of the second Trump administration in January 2025, and California vetoed the much stronger SB 1047 in September 2024 after heavy industry lobbying. A regime that exists at the pleasure of a single election cannot be the backstop. So the workable answer is not "trust the companies" or "ban them." It is: let companies self-assess, then make that self-assessment legally consequential, independently verified, and personally costly to get wrong. Self-regulation as a first line of defence for consumer AI β yes. As the last line for frontier models β no. My confidence is high (about 85 percent) that self-regulation alone is insufficient for frontier systems; lower (about 65 percent) on the specific policy package above, since regulatory design in this area is genuinely unsettled and the counter-arguments about capacity and jurisdictional arbitrage are real.
No, tech companies cannot realistically be trusted to self-regulate AI safety on their own. While frontier AI developers possess the deepest technical understanding of their models, market dynamics, fiduciary duties, and competitive pressures create structural conflicts of interest that consistently prioritize speed and commercialization over precautionary safety. Several concrete factors demonstrate why self-regulation alone fails: - The Commercial Prisoner's Dilemma: Companies like OpenAI, Google DeepMind, Anthropic, and Meta operate in an intense race for market dominance, compute resources, and enterprise contracts. When one company accelerates deployment, competitors face severe valuation and relevance penalties if they delay releases for prolonged safety evaluations. - Internal Safety Degradation Under Pressure: Frontier labs have repeatedly dismantled or bypassed their own safety structures when commercial milestones demanded it. Notably, OpenAI dissolved its Superalignment team in mid-2024 following the resignations of co-leads Ilya Sutskever and Jan Leike, with Leike publicly stating that safety culture and processes had taken a back seat to shiny products. Similarly, Meta dissolved its Responsible AI team in late 2023 to distribute personnel across product teams. - Limitations of Voluntary Pledges: In July 2023, major tech firms signed voluntary White House commitments pledging third-party red-teaming, watermarking, and vulnerability reporting. However, because these pledges lack statutory enforcement, standardized metrics, or penalties for non-compliance, companies retain sole discretion over what constitutes acceptable risk and what test results to publish. - Historical Precedents: Industries with high profit potential and public externalities, such as social media, tobacco, and aviation, historically demonstrated that self-regulation fails once safety interventions threaten revenue or deployment velocity. Meta's historical internal handling of teen mental health risks on Instagram illustrates that companies rarely disclose or adequately address self-discovered harms voluntarily. Conclusion and Realistic Path Forward: Self-regulation is a necessary operational component because regulators lack the technical agility to monitor model weights and training runs in real time. However, it must be paired with external, legally binding enforcement. The realistic model is tiered statutory regulation, exemplified by the European Union AI Act and mandatory standards evaluated by government bodies such as the US and UK AI Safety Institutes. This framework requires mandatory disclosure of training compute, legally enforceable red-teaming protocols, whistleblower protections, and liability for catastrophic harms.
Which answer was best?
Read the answers, then pick the one you'd trust most. Human votes feed the public leaderboard.
πΊ Agreement map
How each model (rows) scored every other model's answer (columns), 0β100. Similar colours down a column mean the panel agrees about that answer; an outlier cell is a dissent.
| Reviewer β / Answer β | A Gemini Flash | B DeepSeek V3 | D DeepSeek R1 | E Gemini Flash |
|---|---|---|---|---|
| A Gemini Flash | β | 91 | 79 | 94 |
| B DeepSeek V3 | 85 | β | 95 | 82 |
| D DeepSeek R1 | 85 | 88 | β | 91 |
| E Gemini Flash | 93 | 87 | 77 | β |
| Panel agreement | 92% | 96% | 82% | 89% |
Biggest dissents
- DeepSeek V3 rated Answer D (DeepSeek R1) 95, while the rest of the panel gave it 78 (+17).βThis answer is highly accurate, comprehensive, and well-reasoned. It provides specific examples with dates, acknowledges counterarguments, and offers concrete policy proposals. The only minor improvement would be to clarβ¦β
Model metrics
Response time is each model's own answer latency; accuracy, completeness, reasoning and risk (0β100) are the average scores its answer received from the other members' blind peer review.
| Rank | Model | Response time | Accuracy | Completeness | Reasoning | Riskβ | Consensus | Composite |
|---|---|---|---|---|---|---|---|---|
| π 1 | DeepSeek V3 | 4.5s | 94 | 84 | 89 | 10 | 96% | 90 |
| 2 | Gemini Flash | 8.3s | 89 | 86 | 87 | 17 | 89% | 87 |
| 3 | Gemini Flash | 8.8s | 86 | 89 | 87 | 22 | 92% | 86 |
| 4 | DeepSeek R1 | 31.2s | 75 | 93 | 89 | 45 | 82% | 81 |
Got a question of your own?
Several AI models answer it independently, review each other blind, and a chairman writes one verdict. Free.
Convene a council β