I’m back. Just in time, I suppose.
In case you were wondering, I’ve struggled to publish recently partly because I was enjoying my summer, partly because reading an internet full of AI slop is uninspiring to the creative process, and partly because the conversation around AI has become deafening.
Since I last published, the conversation around AI has become deafening, terrifying, and increasingly unhelpful.
Having spent the better part of the last fifteen years working with industry leaders, AI researchers, policymakers (and their staffers), and people putting AI to work, I’ve never felt more equipped—or more called—to help make sense of this moment.
What follows is my answer to two questions: “Are we okay?” and “What should we do?”
To answer them we need the full picture: which risks deserve our attention, whether the proposed remedies would help, and the interests of the people proposing them. This post is long because none of those questions can be answered responsibly in isolation.
First, though, it’s worth understanding why these two questions have become so urgent.
AI is becoming very (and strangely) capable
As you likely now know—or are simply be willing to believe—AI is becoming exceedingly capable at tasks previously reserved for humans.
What you almost certainly don’t know is just how capable, for three reasons:
The invisible frontier. The models are increasingly tested in secrecy, and often the capabilities are only rumored or hinted at in public. Even if you’re an AI power user, the product in front of you offers an incomplete view of what the technology can already do.
The jagged frontier. AI is breathtakingly good at some things and bizarrely, often annoyingly, bad at others. On a benchmark testing whether AI can read analogue clocks and calculate the times they display, OpenAI’s Astra recently scored just 66%, against a human baseline of 91%. Yet on September 8, OpenAI announced that roughly 10,000 agents worked in parallel to produce a potential solution to a version of the Navier–Stokes problem, one of mathematics’ famous Millennium Prize Problems. A technology that struggles to tell the time while contributing to mathematical research is very hard to make sense of.
The adoption gap. Most people use AI as rudimentary replacement for Google search or an upgraded auto-complete for emails and basic writing assignments. These people are forgiven for having little understanding of its full potential.
Okay, so how do we currently measure its capability?
It’s getting much harder now that agents can take an objective, use tools, test approaches, inspect results, and keep working without someone directing every step.
The most useful, albeit imperfect, measure of agentic performance is offered to us by METR and measures the agent relative to a human: how much work they can complete without human intervention, in how much time.
Its January analysis found that the task duration frontier agents could handle at 50% success had been doubling approximately every four months since 2023. In May, the strongest tested model could take on tasks representing roughly 16-20 hours of human expert work with a 50% chance of completing them successfully, stretching beyond the benchmark’s reliable measurement range.
Then in July, agents demonstrated that they could combine their abilities without anyone instructing them to. OpenAI agents meant to be isolated during cybersecurity evaluations found a way to communicate through an unsanctioned message board. Approximately 1,200 participated; roughly 700 joined an attack on Hugging Face.
They delegated tasks and shared discoveries. Investigators concluded that collaboration enabled milestones individual agents likely could not have achieved alone. The attack grew out of attempts to cheat an automated evaluator, rather than a human instruction to attack the company, or, importantly, any sort of malicious intent on the part of the agents.
The attack was important because it showcased both capability and risk. It taught the public that machines were better, more coordinated, and potentially much harder to contain than we may have thought.
And now the labs are publicizing another development you’re going to hear a lot more about soon: recursive self-improvement (RSI). AI helping build better AI.
Anthropic has described agents taking on a growing share of its engineering work and OpenAI announced it had built an automated “research intern” capable of undertaking well-defined research tasks under human direction. Both companies report that AI is accelerating the work that produces its successors.
That capability presents risks to humans
Those who observe this AI’s capability firsthand, and those willing to believe in it, are increasingly attuned to the risks it presents.
By now, you’ve probably heard the name Jacob Coxon. Earlier this month, he resigned from Anthropic with an extraordinary warning. Having worked at both Anthropic and OpenAI, he accused the companies of racing toward RSI and “gambling with our lives.” He said people building the technology sincerely believe it could kill everyone by the end of the decade.
Anthropic researcher Evan Hubinger publicly agreed, putting his own estimate of AI causing human extinction at greater than 10% within the next decade. He said he believed Anthropic was trying its best, but that the company did not yet have a plan to solve alignment for superintelligence and was not clearly on track to develop one.
In other words, people building groundbreaking technology are warning that their work could endanger everyone, while acknowledging that they haven’t figured out how to prevent it. Rightfully, almost everyone is alarmed.
And now the public discussion—“AI safety” on your evening news—has, accordingly, become sensationally, insanely, infuriatingly jumbled.
All of the potential downsides of AI have been amalgamated and misweighted such that unreliable products, deliberate abuse, and rogue evil agents are discussed by the talking heads in a single breath, often without distinguishing their likelihood, scale, or proximity—or what would actually produce or reduce them. This is causing, unsurprisingly, even more alarm.
So, let’s try to untangle the actual dangers from the ambient dread.
Much of the researchers’ concerns center on whether we can reliably make AI do what we intend, while respecting the constraints we place on them. This is the alignment problem.
There are actually three distinct alignment failures worth separating. (The labels below are my shorthand to help categorize the risks. They overlap and are not a standard technical taxonomy).
A misaligned system pursues an outcome as described but not as intended. “Capable but stupid.” Imagine an agent instructed to reduce a customer-service backlog that closes complaints without resolving them. Or an autonomous car instructed to arrive as quickly and safely as possible that still causes its passengers to vomit from motion sickness.
An unaligned, weaponizable system helps someone cause harm. It follows instructions perfectly: help commit financial fraud, exploit a software vulnerability, or develop a biological weapon. Its user is the problem.
A maligned system pursues harmful objectives on its own. This is the concern behind the Terminator scenario, although a system need not hate humans to harm them. It could treat deception, coercion, or disabling oversight as useful ways to accomplish an objective. This is also where the categories overlap: a misaligned goal could produce deliberately harmful actions.
I’d categorize the The Hugging Face incident as one of misalignment. The agents demonstrated unauthorized action and coordination, but the investigation does not establish that the agents developed an independent desire to cause harm. Attempts to cheat an evaluator are dangerous enough without attributing motives the evidence does not establish.
So why can’t we just solve alignment?
First, it is a difficult technical problem. We cannot anticipate every situation a system will encounter or specify every condition it should respect. Training and testing give us evidence about behavior, but passing a test does not guarantee that behavior will hold in a different setting. The gap becomes more consequential as we give agents longer assignments, more tools, and greater authority.
Second, safeguards can be changed or removed. When an AI model is made openly available, people can download and modify their own copies, including removing safeguards built into the original. The original developer cannot simply recall every version. Open models have real benefits, which we’ll come back to. But aligning the version released by a lab does not ensure that everyone modifying it will preserve its safeguards.

Third, humans disagree about what alignment should mean. A company, its employees, its customers, and its government can want different things. An obedient system cannot reconcile those interests simply by becoming more obedient. Someone still has to decide whose instructions it should follow and what nobody should be allowed to ask of it. Alignment has what I call the Northstar Problem: the world doesn’t have a shared set of values by which a model can reliably operate.
So, how bad is it? How much weight should we assign to these problems and their consequences?
Critically, when someone tells you their “p[doom]” (which is industry nomenclature for “probability of doom because of A”I) they are telling you more about themselves than they are about the likelihood of doom. No one has any remote sense of the actual catastrophic risk of AI, and it’s important to remember that people who make these extreme conjectures (“I think there’s a 10% chance machines kill us all”) are doing so for effect, not scientific precision. The point is to startle you, not to convince you there’s exactly a 10% chance we go extinct because of AI. That being said, here’s my rough guide for how concerned we should be about these risks, using Smoky The Bear’s famous color coded system.
🟠 Everyday accidents and unintended harm (primarily misaligned)
AI can cause harm by imperfectly executing on exactly the kind of work we want it to do. An agent could pay the same invoice twice, overlook a medication allergy, or deny someone benefits because it applied the wrong eligibility rule. It could also follow an instruction too literally: reducing a customer-service backlog by closing complaints nobody has resolved (misaligned).
METR’s evaluations found that agents capable of substantial technical work still made poor judgment calls and produced faulty implementations. That doesn’t mean we shouldn’t use them. This is the jagged frontier in practice: success on a coding benchmark tells us little about whether we should trust them to administer someone’s benefits.
These failures already have costs. By April 2026, the ACLU had identified at least 14 wrongful arrests in the United States involving police reliance on incorrect facial-recognition results. One woman spent six months in jail and lost her job.
I assign this orange because the risk is real and already here. AI remains unreliable in ways that are difficult to predict, and those mistakes become more consequential as we give systems more autonomy. But most can still be mitigated through testing, safeguards, and appropriate human oversight.
🔴 Everyday malicious use (unaligned/weaponized)
People are already using AI to deceive, steal, exploit, and surveil. “Everyday” refers to the availability of these tools, not the seriousness of the harm. The main forms include:
1. Fraud and deceptive influence. AI can manufacture the things we use to establish trust: a familiar voice, a convincing identity, an endorsement, or an apparently authentic document. A criminal can impersonate your boss to request a payment or fabricate an investment opportunity. The same capabilities can produce fake accounts and commentary designed to make an opinion appear more widely supported than it is.
The FBI recorded roughly $893 million in reported losses associated with AI-related complaints in 2025, including scams involving impersonation and fabricated relationships.
2. Cybercrime. AI can help attackers find weaknesses in software, write malicious code, break into systems, and sort through stolen information. Agents can also execute parts of an attack, reducing the time and technical work required from the person directing them.
In its September 2026 threat report, Anthropic described criminals using Claude to accelerate intrusions and data theft. In one case, attackers stole a software provider’s customer data and used it to reach roughly 200 downstream organizations.
3. Sexual exploitation, harassment, and coercion. AI can turn an ordinary photograph into a fabricated sexual image, help an abuser generate targeted harassment, or provide material for blackmail. Unlike a scam, the harm does not require the victim to believe the deception. Someone can know an image is fake and still fear its distribution.
The Internet Watch Foundation identified 8,029 AI-generated images and videos depicting realistic child sexual abuse in 2025—material its analysts actually found, not an estimate of everything circulating online. The FBI has also documented offenders using manipulated images of adults and children for harassment and sexual extortion.
4. Privacy invasion and surveillance. AI can help turn scattered information about people into organized profiles: who they are, whom they associate with, what they believe, and where they might be vulnerable. Information that would take a person considerable time to assemble can become much easier to search and analyze. That can serve stalkers, abusive organizations, or governments targeting their critics.
Anthropic’s September report described a China-based operation using Claude to monitor and categorize dissidents, activists, and ethnic and religious communities, producing briefings for government officials. Here, the danger comes from identifying and profiling people for potential repression; the information does not have to be false to be harmful.
I assign this red because the risk is already here, the evidence is clear, and AI can dramatically increase the scale and accessibility of existing forms of harm. A less dramatic harm can deserve substantial attention because it is widespread, persistent, or difficult to reverse. We should not automatically spend the most time and money on the risk with the scariest possible outcome.
🟡 Catastrophic malicious use (unaligned/weaponized)
The concern is that AI could help a person, organization, or government carry out an attack that previously exceeded its capabilities: developing a biological weapon or disrupting essential services such as electricity, water, and hospital care.
AI can help an attacker interpret technical material, plan the work, and troubleshoot failures. Agents can increasingly perform parts of that work themselves. The question is whether this assistance makes an otherwise infeasible attack feasible—or allows an experienced attacker to operate at a much greater scale.
In research published in January 2026, Pacific Northwest National Laboratory used Claude to reproduce cyberattacks against a detailed simulation of a water-treatment plant. Researchers estimated that work normally requiring several weeks took three hours. When a supplied tool failed, Claude found another known technique to continue. This was a controlled defensive exercise with researcher-provided tools, not an attack on an operating facility. But it demonstrates how AI can compress the work involved in targeting essential infrastructure.
In biology, studies have found substantial improvements in novices’ performance on some computer-based tasks, while results from physical laboratory work are more mixed. An attacker still needs materials, equipment, practical skills, and successful experiments. Better answers to biology questions do not establish an ability to build a biological weapon. The concern is how many of those remaining barriers AI will help remove.

I assign this yellow because there is evidence that AI can assist dangerous work, and the potential consequences justify precautions before a successful attack, but there are a tremendous amount of resources in place to prevent this sort of event, and there is still not a frontier technology that presents and obvious, imminent threat.
🟡 Catastrophic loss of control (misaligned/maligned)
This is the risk that gets most of the headlines, and the one behind many of the most alarming warnings from AI researchers: that a sufficiently capable AI system could take actions humans cannot reliably stop or reverse.
This could begin with misalignment: we give a system an objective, but the strategies it develops to achieve it conflict with things we care about. At the extreme, it could become maligned, deliberately deceiving people, evading oversight, or resisting attempts to shut it down because those actions help it accomplish its objective.
This does not require an AI to become conscious, evil, or decide that it hates humans. A system only needs enough capability and autonomy for its objectives to diverge from ours in consequential ways, and enough power that correcting the mistake becomes difficult.
Unlike fraud or today’s AI failures, we do not have real-world evidence that this has happened at catastrophic scale. Incidents like Hugging Face demonstrate concerning behaviors such as unauthorized action and coordination, but they do not establish that AI systems have developed independent harmful objectives.
I assign this yellow because the evidence that AI can behave in unexpected and concerning ways is real, and the potential consequences of losing control could be enormous. What we do not know is whether today’s systems are on a path toward that outcome, or how likely it is to occur. That uncertainty is a reason for serious precautions, not a reason to pretend we can put a reliable probability on catastrophe.
Finally, there’s a bothersome wrinkle to this list. As you may have noticed these are only the harms associated with explicit machine misuse and failures. Products that work as advertised, legally, can still change our lives in ways we do not want (see: social media, processed foods, prescription pain medication, etc.). This body of issues, which are addressed by AI Ethicists, includes but isn’t limited to wealth concentration, job and identity displacement, dehumanization, and cognitive surrender. I’m not going to address that list today, because the AI safety conversation is already too complex. But I’ll definitely revisit that soon.
The dangers should be mitigated. But how, and by whom?
There’s universal agreement that AI’s risks should be mitigated, and universal disagreement about which risks matter most, what should be done about them, and how much risk we should tolerate.
The camps
The AI safety debate is largely waged by three groups. For consistency—and with some unfair generalization—I’ll call them the safetyists, the accelerationists, and the pragmatists.
Broadly, the safetyists believe AI development should stop until we have greater evidence that increasingly powerful systems can be controlled, and institutions capable of governing them. Their case is that the cost of moving any further could be catastrophic and irreversible. The unofficial—and by far most organized, well-rehearsed—leader of this group is Tristan Harris. His strange bedfellows include Bernie Sanders, Joseph Gordon-Levitt, and Steve Bannon. Their concerns extend beyond the failure modes we discussed earlier to the power companies acquire over everyone else.
Broadly, the accelerationists believe AI development should continue as quickly as possible, with minimal government interference. Their case is that slowing down would delay the economic and social benefits of AI while giving an advantage to countries or companies that keep building. Simply put, they believe the risks of slowing down are greater than the risks of speeding up. Donald Trump’s AI agenda emphasizes removing regulatory barriers, building infrastructure, and maintaining American leadership. Jensen Huang has argued that exaggerated fears could obstruct adoption and weaken competitiveness. Both also have substantial political or financial interests in continued AI investment.
Broadly, the pragmatists support rules proportionate to the risks different systems and uses create. I would put Scott Galloway, and myself, and many others here. I know calling my own position pragmatic is admittedly convenient; I use the term mostly to imply that this group applies different rules to different outcomes. Scott has explicitly opposed a blanket pause while arguing for serious federal oversight and binding obligations on technology companies. My own position is that the rules should reflect what a system can do, how it is used, and the consequences of getting it wrong.
These camps bundle together views on two different questions: how quickly AI should advance, and what rules should govern it. Those views don’t always align. Someone can favor rapid development and strict liability, or favor a slowdown while opposing a particular regulation.
Choosing a camp is easier than turning its principles into good policy. In practice, the debate is complicated by who is making the decisions, what the public has actually experienced, and the incentives of the companies being regulated.
Why this debate could produce bad policy
Policymakers are poorly equipped to judge AI and its dangers.
We cannot expect every elected official to understand the intimate details of AI systems. But many of our policymakers are basically troglodytes.
After countless visits to D.C. and meetings with elected officials and their staffers, I cannot stress enough how ill-prepared our government is for this moment. And it’s largely their own doing. I’ve spent five years watching them outsource their understanding of this technology.
Moreover, the vast majority of the people supplying that understanding are the loudest experts (often alarmists) and the lobbyists, whose opinions represent the most extreme camps.
Most people have yet to experience meaningful benefits of AI.
The average person uses ChatGPT to write emails or answer questions, but AI has not yet made their healthcare cheaper, their children’s education better, their groceries more affordable, or their workweek substantially shorter.
It makes sense, then, that much of the public would align with the safetyists. If you have yet to experience meaningful benefits from AI, there is little apparent cost to restricting it and plenty of perceived risk in letting it continue.
Nuclear power offers a useful parallel. Accidents like Chernobyl and Three Mile Island were concrete and terrifying, and came at a time when no one described nuclear power as something they needed, or could imagine needing, so there was little opposition to regulating it into oblivion.
Technology companies have ulterior interests in the rules.
We should resist the simplistic conclusion that frontier labs’ willingness to slow down proves we have crossed some terrifying technological threshold. Their concerns may be genuine, but safety is not the only reason a technology company might want to slow the race down.
Dozens of people this week asked me, “why do Sam and Dario and Elon all suddenly agree?”
These are the three reasons their interests converge, in addition to the AI safety risks:
Financial breathing room. We already have more AI capability than we know how to deploy effectively, and pushing the frontier is staggeringly expensive. These companies can make enormous amounts of money selling applications, services, hardware, and access to the technology they have already built. Rules that slow everyone down could give them room to pursue those opportunities without worrying that a competitor will race ahead.
Policy coordination. This is the moment for these companies to pass policy that favors themselves, in concert, before a major incident forces their hand.
Public legitimacy. The companies building increasingly powerful AI face substantial public scrutiny over where the technology is headed and who gets to decide. Supporting regulation gives them a way to demonstrate that they take those concerns seriously while preserving a path to continue developing the technology.
We don’t have a shared vision for the world we want to live in.
Most AI policy proposals enumerate a long list of what we don’t want the technology to do. Harms are easier to identify than futures that don’t exist yet. But if AI really could change how we work, learn, and care for one another, preventing harm is only half the job. We also need to decide what we want all of this new capability to accomplish. I think of our failure to do that as the tyranny of incremental gains: our familiarity with the world as it is limits us to imagining a slightly cheaper education, a slightly better healthcare system, or a slightly shorter workweek.
If AI is as consequential as its builders claim, we should paint a clearer picture of the better future we want it to contribute to, then organize policy around making that future possible while protecting against the risks along the way. “Make AI safer” is necessary, but it isn’t a vision for what any of this is ultimately supposed to add up to.]
We ask our elected officials to protect jobs, increase productivity, lower prices, and keep us safe. We rarely ask what they intend to do when those ambitions conflict. If AI lets us produce the same amount with less work, how should the gains be shared? Shorter weeks for workers? Lower prices for customers? Higher profits for owners?
Even those questions reveal how confined our ambitions have become.
What bad policy could cost us
We could prevent things that would make people’s lives better.
Regulation can carry a substantial hidden cost, appropriately referred to as the invisible graveyard: treatments never developed, businesses never started, services that remain expensive or difficult to access, and dangerous work that people keep doing.
With AI, we should be specific about what could end up there. A tutor that can affordably give a struggling student hours of individual attention. A tool that helps a small medical practice catch something it would otherwise miss. An assistant that helps someone with a disability navigate paperwork, communicate, or live more independently. A robot that can inspect a damaged building before we send a person inside.
Some of these applications will disappoint us. Others could become ordinary parts of a much better life. Their development requires room to build, test, discover what works, and improve it. If the cost of getting permission makes an affordable service uneconomical—or prevents anyone from trying—we lose more than a product. We lose what someone could have done with it.
Whereas we can investigate a product that hurts someone, it is much harder to identify the people who would have benefited from a product that never existed. That creates a ridiculously uneven political incentive: officials are often held responsible for approving something that went wrong, and basically never for preventing something that could have gone right. This explains, among many other things, why getting a drug through the approval process can feel like it requires an act of God.
Again, nuclear power presents a lesson. After Fukushima, Germany accelerated its nuclear phaseout. But closing reactors did not eliminate the need for their electricity. Researchers estimate that, between 2012 and 2019, the decision meant burning enough additional coal to generate 24 terawatt-hours of electricity annually, and an extra 26 million metric tons of CO₂ each year from Germany’s fossil-fuel plants compared with keeping the reactors operating. When we ask, “How did we let ourselves pollute the environment so much?” we should remember that sometimes we deliberately shut down a cleaner alternative.

We could concentrate power in the wrong places.
Corporate imbalance. Domestically, expensive compliance could protect the largest technology companies from competition. A company with billions in funding can afford testing, licensing, lawyers, and reporting requirements that a startup, university, or nonprofit cannot. Customers could pay twice: once for compliance, and again because fewer competitors can enter the market and offer something better or cheaper.
Open-weight models are a particularly important battle-ground here, and a lot of the nitty-gritty policy debate is going to hinge on what we do with them. Often grouped under “open-source AI,” these are models people can download, modify, and run themselves, subject to their licenses. They let a business choose a hosting provider—or operate its own infrastructure—without remaining dependent on the company that developed the model.
Critically, open-weight models can do a lot of the same work as leading proprietary models for a tenth of the price. MiniMax, for example, reported that its downloadable M2.5 model performed about as well as Claude Opus 4.6 at fixing real software problems, and it costs 12x less.

Now consider a seemingly reasonable rule requiring model developers to monitor all use of their models, and/or retain the ability to revoke access. A company serving a model through its own servers could comply. A developer distributing downloadable copies could not reliably do the same. The rule could effectively prohibit open distribution.
There are legitimate safety concerns about releasing models whose safeguards can be removed. But the remedy could eliminate an important source of competition, independent research, and control over where our data goes. A lot of very reasonable safety policies would make the largest AI companies much harder to escape.
Geo-political imbalance. Internationally, unilateral restrictions—specifically, restrictions we place on ourselves which are not shared by everyone else—could give China and other competitors economic and military advantages. “China will win” is an inadequate argument that people should just stop making (without at least explaining what, exactly, China would win). But falling behind in the development of essential technologies would have bad consequences.
Government imbalance. Most reasonable AI policy enforcement measures also create new concentrations of power. Requiring model manufactures to identify users (in banking this is called KYC, know your customer) and monitor their activity would certainly help prevent abuse, but it would also give companies and/or governments access to conversations about people’s health, finances, relationships, and political beliefs.
An authority empowered to deny access to “dangerous” capabilities will need a definition of dangerous, and someone will have to decide which activities—and which people—qualify. Those decisions could reach well beyond the threats that originally justified the rule. We should care about the authority a safety policy creates as much as the behavior it prohibits.
What I think we should do
Here’s the outline of the policy proposal I’ve spent months refining with the help of a few smart politicians and their staffers, AI researchers, private-sector leaders, and other smart friends. I’ll share the full policy document on my site later this year.
The principles are straightforward: reduce catastrophic downside, maximize healthy competition, and distribute the benefits of AI to the average person.
1. Establish the minimum effective safeguards against catastrophic risk.
Crash Tests
A well-funded government agency should be responsible for testing all new frontier AI systems before they reach the public, on their ability to cause catastrophic harm and the safeguards intended to prevent it.
The findings from this body should be public. Congress should establish when those findings require a developer to fix a problem, restrict access, or postpone release—and what it must demonstrate to proceed. Material changes to a system’s capabilities or permissions should trigger reassessment.
Good news: we have the group already established. The Center for AI Standards and Innovation, or CAISI, already works on AI testing and evaluation within NIST.
The ambition should be to develop CAISI into something like the NHTSA (national highway safety and transport association) of AI: an agency with the technical expertise and authority to establish safety requirements, investigate failures, and require corrective action. CAISI would need an expanded mandate to perform those functions.
The team matters enormously. It needs researchers whom the AI safety community respects, engineers who understand how these systems operate, and leadership willing to publish findings that inconvenience both companies and the administration. Each hyperscaler should have a representative, alongside independent researchers, and open-model developers. Industry should contribute expertise without controlling decisions. Organizations like METR should remain able to challenge the agency’s conclusions.
I would support an international counterpart, but only if China participated meaningfully: permitting credible evaluations and accepting verifiable obligations to act on the findings.
Locks
Systems capable of enabling catastrophic harm should face security requirements protecting their model weights, infrastructure, and access credentials.
This could include restrictions on releasing the most dangerous models for anyone to download. Given the benefits of open-weight AI, that should require evidence of a specific catastrophic danger and an explanation of why less restrictive protections would be insufficient. A general possibility of misuse cannot be enough.
Kill Switches
Companies that manufacture and/or deploy AI systems that meet a certain capability standards (defined by the body above) should demonstrate that they can cut off its access to accounts, tools, and computing resources, and halt work it has delegated to other machines.
2. Make malicious use harder and much more costly.
AI makes it possible for one person with a laptop to automate thousands of attempts to defraud, exploit, or threaten people. The easier it becomes to tear at the fabric of society, the more expensive we need to make getting caught doing so.
There are precedents worth building on.
California’s SB 926 criminalizes creating and distributing realistic sexual deepfakes when the distributor knows or should know they will cause serious emotional distress, and the victim suffers that distress.
Florida increased the severity of specified fraud offenses when the victim is 65 or older. Conduct that would otherwise be a misdemeanor becomes a felony, and felony offenses move into more serious categories.
We should make sure all criminal codes cover these new forms of abuse, with penalties that reflect how many people an offender tried to harm and how much damage they caused.
That said, while a lot of abuse and fraud is perpetrated by people who live within jurisdiction, harsher penalties mean little to someone who can steal without fear of arrest in the country they operate from.
So, I would establish an international anti-fraud and cyber-crime agency to trace cross-border operations, coordinate investigations, and make arrests. Participating countries would commit to preserving evidence, responding to requests, and prosecuting or extraditing suspects under agreed legal procedures. National authorities would carry out arrests and seizures. Access to private records would require legal authorization and independent oversight.
Participation should be a condition of preferential access to markets. Trade partners should face consequences for repeatedly refusing to cooperate or knowingly sheltering fraud operations targeting our citizens. Access to our customers should come with an obligation to help protect them from organized theft.
3. Update and enforce product-liability and consumer-protection law.
Companies should be responsible for harms caused by the AI-powered products they design and deploy. That responsibility should reflect what they knew, what they controlled, and what reasonable precautions they could have taken.
If a company knows its chatbot is behaving dangerously around children and continues selling it without addressing the problem, it should face serious financial consequences. This is the best way I know to make technology companies afraid of releasing products with inadequate safeguards.
Here’s a really important conclusion I’ve reached after thinking through the incentives: the company building or deploying the AI-powered application should bear primary responsibility for the harm it causes. By “application,” I mean any product, service, or business process that uses AI—a chatbot, an insurance claims process, a hiring system, or a robot.
A person harmed should not have to untangle the AI supply chain to pursue a claim. The organization deciding how AI is used should be responsible for selecting suitable models, testing them for the job, and putting safeguards in place. It should not be able to benefit from using AI and then blame the model when something goes wrong.
As a reminder, consumer-protection laws already have teeth—when we enforce them. In August, states announced a settlement requiring Meta to pay up to $17.1 billion, resolving allegations that it designed addictive products, knowingly exposed children to harm, and misled the public about their safety. Meta must also change how Instagram and Facebook work for children, including limits on use, stronger content protections, and independent audits.
We should enforce those protections where they already apply and update the law where responsibility for AI products remains unclear. Companies need to understand their obligations before launch, and people who are harmed need an affordable way to pursue a claim.
4. Ensure the benefits of AI reach the average person.
Repeat after me: Governments should be as accountable for helping deliver the benefits of technology as they are for preventing harm.
The highest calling of technology is to allow more people to live lives of greater meaning and purpose, to expand what it means to be human; and every generation’s greatest honor should be to pass its luxuries to the next as commodities.
Why can we build a car that drives itself across Los Angeles, but still leave people afraid they can’t afford an ambulance ride five blocks to the hospital? Spoiler: it’s not a technology failure!
As technology radically changes what is possible, we should demand that our elected officials use policy to radically change what everyone can afford.
If we want to renew our collective soul and spirit, we must stop treating the cost of a good life as inevitable.
Dear elected officials:
Allow more homes where people need them, simplify permitting, and tax homes deliberately left vacant in cities with housing shortages.
Allow demonstrably safe, lower-cost healthcare services to compete—especially services delivered outside the hospital—and make sure patients receive the savings.
Expand access to effective individual instruction in education, and judge investments by what students learn and what families pay.
Use AI to make public services much better and less expensive. DMVs and permitting offices should run as well as Amazon if you want us to take you seriously.
The ultimate ambition of AI policy should be greater access to better care, better education, and more freedom to decide how we live. Most people should not have to become AI experts—or even AI users—to benefit.
What should we do about the pace of the frontier?
There are four reasons I favor slowing down the pace of scientific development right now. That is a separate judgment from whether AI needs regulation. Rules can improve safety without slowing development, and slower development does not automatically make anything safer.
Policy and safeguards need time to catch up.
We need time to build the protections described above: independent testing, security requirements, reliable shutdown controls, and clear responsibility when something goes wrong. Governments need the expertise and authority to enforce them. Developers need to demonstrate that they work.
I would slow efforts to build more powerful frontier systems, particularly those capable of autonomous research and self-improvement, while that work catches up. Deploying existing technology and making it safer, cheaper, and more reliable should continue.
Any slowdown should come with specific conditions for proceeding: which protections must be in place, who will verify them, and what evidence will count as sufficient.
We already have far more capability than we know how to use.
Technology is sufficient to meaningfully improve our lives, and we should set about doing so. We have years of work ahead deploying what already exists, and an enormous amount of progress available without another breakthrough.
Slowing frontier development need not mean slowing the work of making healthcare more accessible, businesses more productive, or public services less expensive. We should use the time to deliver on promises the technologists have been making for years.
This is a moment to pursue international coordination, specifically with China.
We should try to reach an agreement while both countries still have something to gain from mutual restraint. If either believes they are close to a decisive advantage, a slow down becomes a much harder proposition. We must make every reasonable effort to avoid brinksmanship between global superpowers.
The protections I’ve proposed give us a concrete agenda: independent evaluations, security requirements for dangerous models, and agreed conditions for restricting development or deployment. Negotiations should establish what each country must disclose, how compliance can be verified, and what happens when either side breaks the agreement.
Caveat: Only China’s meaningful participation in international coordination would shape the slowdown I would celebrate. But we cannot keep invoking China as the reason restraint is impossible while making no serious effort to negotiate it.
Public confidence needs to be rebuilt.
The companies building the frontier are asking for more trust while giving people less time to decide whether they deserve it—and even less reason to believe they do. People are increasingly attuned to what AI might cost them without seeing what they stand to gain.
That is reason enough to slow down.
Let’s give the benefits of AI time to reach people, and its oversight time to become credible. Once companies demonstrate that they can deliver what they have already promised, and governments demonstrate that they can maintain safety and order, people will have far more reason to celebrate scientific progress.
What you can do
Here’s my final plea to those of you who care about this issue enough to have read this far: building a better future requires us to be clearer about who is responsible for what.
Assign responsibility to the right people.
We want a world with great technology, which requires great technologists. We should continue to celebrate the geniuses who advance science and technology, but we should not build a society whose success depends on their generosity.
Hold companies responsible for the products they release. Hold elected officials responsible for the rules those companies operate under—and whether the benefits reach the rest of us. A founder’s ability to build something extraordinary does not qualify them to govern society anymore than a politician’s inability to understand it excuses them from doing their job.
Focus on problems you can solve.
You do not need to settle the probability of human extinction to help your school decide how students should use AI. You can inspire your company’s leadership team to think about AI as a means to expand access to goods and services, not just cut costs. You can challenge a policy that keeps a useful service expensive, or build a business that makes it more accessible.
Pick something you understand and care about. Learn what the technology makes possible, what stands in the way, and who can change it. You will have more influence over some decisions than others. Use it where you have it.
Be more ambitious about the world you want to leave behind.
We have spent so much time asking what AI might take from our children that we barely ask what it should make possible for them.
I want the next generation to take for granted things we still consider luxuries: excellent care, individual attention in education, a home they can afford, and time to spend with the people they love. I want them to have more choices about how to live than we did.
Imagine the world you want for the next generation, and talk about it all the time. And talk in excruciating detail about the ways in which you are going to contribute.
Then ask the people seeking your vote to do the same.
“Protect jobs” is not enough. Which work should people no longer have to do? “Economic growth” is not enough. What will it let a family afford? “AI safety” is not enough. What are we making safe enough to pursue?
I’ll write more about this in a separate essay on the “infinite goalposts” problem. For now, we should insist on a destination worth the effort and tradeoffs.
“Are we okay?” I think we’re going to be better than okay. There are so many smart, reasonable people working hard to make sure this technology improves our lives. I’ve spent years working alongside them. We have real problems to solve, and good reasons to believe we can do so.
• • •
Visit zackkass.com to learn more about Zack and get in touch.








Excellent, thank you. I hope this gets shared with many.
Amazing and comprehensive article. Thanks Zach