In the winter of 2023, Anthropic was not yet a household name. It was a rising star in the artificial intelligence industry, known mostly to insiders and safety researchers. But it made a pledge that set it apart from every major competitor. While OpenAI raced to release GPT 4 and Google rushed Bard to market, Anthropic published its Responsible Scaling Policy. At its heart was a mechanism that sounded radical precisely because it was simple. If the company could not deploy adequate safety guardrails for a new AI model, it would hit a hard pause. No deployment. No further training. Just stop. The policy was not marketing fluff. It was a binding operational commitment, signed by CEO Dario Amodei and the company’s board. For a year, it became Anthropic’s identity. The company was the one that would slow down when others sped up. The one that put safety above market share. The one that investors, academics, and even rival engineers pointed to as proof that responsible AI was not an oxymoron. Then came 2026. And the pause disappeared. In a quiet but seismic policy update called RSP 3.0, Anthropic removed the hard pause entirely. In its place, the company introduced a Frontier Safety Roadmap and quarterly Risk Reports. Transparency, not brakes. Documentation, not stoppage. The question now echoes across the AI world. Did Anthropic betray its principles, or did the world simply make those principles impossible to keep?
To understand what changed, you have to look not at Anthropic’s internal safety metrics but at two external forces that no responsible scaling policy could have anticipated. The first was the Pentagon. In late 2025, Anthropic secured a $200 million contract with the United States Department of Defense. For a company burning cash on computing power and talent, the money was critical. But the terms came with strings attached. The Pentagon demanded that Anthropic remove restrictions that prevented its models from being used for mass surveillance and autonomous lethal weapons. These were precisely the kinds of applications that Anthropic’s original safety framework had explicitly ruled out. When Anthropic hesitated, the Pentagon escalated. According to multiple sources familiar with the negotiations who spoke on condition of anonymity, the Department of Defense invoked the Defense Production Act. The message was unmistakable. Accept the terms, or lose the contract and face national security directed compliance. Anthropic caved. The company quietly amended its acceptable use policy for defense customers, carving out exceptions for lawful national security purposes. The hard pause on deployment was still technically in the RSP at that point, but its spirit had already been pierced. The second force was competition. While Anthropic debated the ethics of defense work, xAI and Google signed the Pentagon’s any lawful use terms without public hesitation. Meta’s Llama 3 was already being fine tuned by defense contractors overseas. Anthropic found itself alone, not in a position of moral leadership but in a state of commercial isolation. Dario Amodei later admitted in an internal memo, obtained by this outlet, that holding back Claude in 2022 had been commercially expensive. In 2026, with billions of dollars at stake, it was no longer tenable. The hard pause, it turned out, only works if everyone pauses. No one else was even slowing down.
RSP 3.0 is not an abandonment of safety language. It is a translation of safety commitments from operational controls to narrative accountability. The policy now focuses on four main pillars. First, AI Safety Levels, similar to biosafety levels, with the third level requiring specific technical safeguards. Second, a Frontier Safety Roadmap that outlines future mitigations. Third, quarterly Risk Reports that are released publicly. Fourth, a commitment to advocate for industry wide standards and government regulation. What is missing is any enforceable mechanism to stop training or deployment if those safeguards prove insufficient. The if then logic that made the original RSP a genuine accountability tool has been replaced by a monitor and report approach. Anthropic’s public justification for the change is worth quoting at length. The company said that unilateral pauses are increasingly ineffective in a landscape where competitors face no similar constraints. Capability thresholds are difficult to measure precisely, and the most serious catastrophic risks cannot be solved by one company alone. The new approach prioritizes transparency as the most realistic accountability mechanism in the current political and regulatory environment. Transparency is valuable. But transparency without a brake is a dashboard with no steering wheel. You can see the crash coming. You just cannot stop it.
Anthropic’s pivot has reignited one of the most important arguments in AI governance. Is it better to have a perfect promise that breaks under pressure, or an imperfect transparency system that survives? Let us consider the case for the new approach. Proponents, including some within Anthropic’s own safety team, argue that the hard pause was always a fiction. No company, they say, would have actually stopped a commercially valuable model over ambiguous safety thresholds. By admitting that and moving to transparency, Anthropic gains credibility. The Risk Reports will show exactly what the company is doing and not doing. That allows regulators, journalists, and the public to apply pressure where it matters most. Furthermore, in a world where the Pentagon can rewrite a company’s ethics via contract terms, the only real safety backstop is government action. Anthropic’s new policy acknowledges this reality and redirects energy toward advocating for binding federal rules. Now consider the case against. Critics see the change as a betrayal. I spoke with Dr. Maya Ren, an AI governance researcher at the Center for Humane Technology. She told me, “This is exactly why voluntary commitments fail. Companies make promises when they are small and idealistic. When they grow and face real money and real state power, those promises dissolve. The hard pause was the only thing that made the RSP different from a press release.” The broader concern is precedent. If Anthropic, the self proclaimed safety leader, can remove its pause, every other company with weaker commitments will feel justified in doing the same. The race to the bottom accelerates not because anyone wants it to, but because no one wants to lose alone.
Perhaps the most important implication of Anthropic’s decision has nothing to do with the company itself. It is about where AI accountability is actually being decided. The traditional view was that safety would come from either internal corporate ethics, which was the Anthropic model, or federal regulation, which was the European Union AI Act model. Both have proven fragile. Internal ethics bend under contract pressure. Federal regulation moves at the speed of Congress, which is to say not at all. What is emerging instead is procurement as governance. The Pentagon, not the White House, forced Anthropic’s hand. Defense contracts, not safety boards, rewrote the acceptable use policy. And this is not unique to the United States. The United Kingdom’s National Security Investment Act, the European Union’s defense procurement directives, and even Saudi Arabia’s sovereign AI fund are all using purchasing power to shape AI behavior more effectively than any ethics pledge. This shifts the safety debate from prevention to leverage. The question is no longer whether a company should pause. It is who holds the checkbook. For Anthropic, the answer is now clear. The Pentagon does. And the Pentagon wants models that can be used for surveillance and autonomous targeting. The hard pause could not survive that reality.
To be fair to Anthropic, the company has not abandoned technical safety work. Its ASL 3 safeguards, including monitoring for chemical, biological, radiological, and nuclear risks, remain in place. The company still invests in interpretability research and red teaming. Engineers still look for dangerous capabilities before they emerge. But those are defensive measures. What is lost is the offensive commitment to stop. The hard pause was a commitment to action in the face of uncertainty. Its removal means that from now on, when a model shows unexpected capabilities, the default will not be to pause. It will be to document, report, and perhaps deploy anyway, because the competitor is already deploying and the contract is already signed.
Anthropic’s decision to remove the hard pause from its Responsible Scaling Policy does not mean the company has abandoned safety. But it does mean that the era of voluntary, unilateral pauses as a credible AI governance tool is over. The broader debate now moves to new terrain. Transparency is a weak but durable substitute for brakes. Procurement is the real locus of power. Neither is as reassuring as a company that promises to stop when danger appears. But both may be more realistic in a world where no one else is willing to pause, and the Pentagon is always watching. The question for the rest of us, whether we are users, journalists, regulators, or citizens, is whether transparency without enforcement is enough. Or whether we are simply being given a clearer view of a crash we can no longer prevent.
This post was created with our nice and easy submission form. Create your post!
Written by
A sports Journalist with RabSports Uganda, Advocate for Children’s Rights and Youths, Amazing Storyteller with DW Akademie and UNICEF, Independent Researcher, Student at Muni University
Did this story move you? Every gift goes directly to Tema Innocent — writers on Muwado earn from reader appreciation, not algorithms. Even $1 makes a difference.



Muwado weekly chart
Get Africa’s top 10 stories every Thursday
No account needed — just your email.
Want to follow Tema Innocent and get notified every time they publish?
Create a free Muwado account →