Dossier / in-depth analysis

Inside Anthropic’s Factory: Power, Safety and Values Behind Claude

STStephane Nachez · · ·15 min
Inside Anthropic’s Factory: Power, Safety and Values Behind Claude
Illustration ActuIA réalisée par IA, d’après une photo de Kimberly White / Getty Images pour TechCrunch, CC BY 2.0.
Contents

One of its earliest investors wanted to steer AI toward greater caution. Two years after Anthropic’s founding, he considered his investor strategy a failure. Yet behind Claude, the founders’ convictions shaped hiring, governance and the product. But who can impose a choice when mission meets the demands of growth?

In April 2023, Jaan Tallinn had a very specific reason to be disappointed. The Skype cofounder had funded AI labs, including Anthropic, hoping to influence their caution. Yet the company did not sign the call to pause for six months the training of the most powerful systems. The initiative came in part from the Future of Life Institute, which he himself co‑founded.

“Plan A failed,” he told Semafor. Anthropic replies that it does not sign petitions as a matter of principle. Tallinn nevertheless continues to judge it more attentive to safety than other labs he knows. This disagreement, reported on April 28, 2023, can be summed up in a few lines: sharing a worry, funding those who take it seriously, and obtaining a decision are three different things.

This difficulty runs through Anthropic’s history. The company gave safety concerns a place in its research and then in its governance rules. It also recruited, raised capital, and made industrial commitments that place it in the competition it aimed to make less dangerous. To understand its choices, one must follow people at the moment their convictions become responsibilities.

After OpenAI, building an organization around a bet

The story begins in another organization. In December 2015, Sam Altman and Elon Musk presented OpenAI as a nonprofit research entity. The collective benefit of AI was already at the heart of the founding announcement, preserved by Internet Archive. This ambition therefore predates Anthropic and the competition between the two companies. The Amodeis worked at OpenAI before leaving to create their own lab in early 2021.

Amodei’s caution had already resulted in a concrete restriction. On February 14, 2019, at OpenAI, he co‑signed the announcement that justified not immediately releasing the most powerful version of GPT‑2, citing risks of malicious use. A reduced version was released; the full model became available on November 5 after a staged publication. Amodei thus publicly supported a collective decision for restricted dissemination, nearly two years before cofounding Anthropic. Developing a model, evaluating its dangers, and deciding who can access it were already part of the same trade‑off.

To come together, Anthropic’s seven founders had to cope with the pandemic. They met outdoors, sometimes in a garden, masked and at a distance. Daniela Amodei would recount those beginnings in an interview with the Future of Life Institute, published in March 2022. The group already knew each other and wanted to bring together research on model capabilities and research on their safety within a single organization.

Dario Amodei explains the split: their bet had previously existed inside a larger structure; they now wanted a company that could orient all its strategic decisions around it. In this founders’ account, the 2021 break comes down to the ability to choose a common direction and devote the entire organization to it.

The brother and sister do not bring the same skills. Dario, the CEO, comes from physics and neuroscience. Daniela, the president, studied literature with a minor in politics at UC Santa Cruz. In the 2022 interview she also traces a background in NGOs and Washington politics. Their moral concerns did not begin with Claude: as early as 2003, the Caltech student paper already described Dario’s activism against the Iraq war and his worries about scientists’ responsibility.

This closeness has very ordinary effects. Asked in July 2023 about the number of physicists on the team, Daniela acknowledged the importance of their networks. To grow the group, one must then hold meetings, recruit, and make trade‑offs. She describes trying to preserve a few hours of reflection early in the day and meeting new employees, a task made harder as the company grew. In this interview published by Stripe, Anthropic’s commercial partner, the mission also translates into the work of a leader trying to keep a shared culture as headcount increases.

How philanthropy enters the lab

Money is needed to turn this group into a company. The first announced round, in May 2021, brought $124 million. Jaan Tallinn led it; Dustin Moskovitz, Facebook’s cofounder, was among the investors. Their names appear in the announcement of this funding.

Part of that milieu shares a question: with limited resources, how to do the most good? Effective altruism proposes comparing interventions, their costs and effects. Depending on the person, it leads to supporting public health, animal protection, or preventing catastrophic risks. Longtermism more specifically focuses attention on the fate of future generations. The two approaches overlap without providing a single answer to the AI question.

Open Philanthropy explicitly situates itself within this movement. In September 2016, its cofounder Holden Karnofsky declared himself a member of this community and described it as the organization’s most natural peer group. He also explained that his growing interest in AI risks helped reorient the organization’s priorities. This closeness is reflected in funding: the 2021 report mentions support for the development of the effective altruism community. Personal convictions, interlocutors and resource allocation thus converge here in the leaders’ own accounts.

Moskovitz and his wife Cari Tuna are also interested in interventions with more immediate effects. In their account published by Good Ventures, a conversation with Holden Karnofsky about iodine deficiencies illustrates discovering lesser‑known interventions that could greatly improve lives. Tuna, a former journalist, devotes herself to this philanthropic exploration. Their funding would also go to pandemic prevention and risks related to advanced AI.

Karnofsky is among those who give this last priority its most ambitious rationale. In his “Most Important Century” series, he envisages a 21st century in which AIs accelerate research so much that they permanently alter humanity’s trajectory. He explores even the possibility of digital civilizations. These are hypotheses he deems important enough to mobilize efforts for, while acknowledging their uncertainty. The shift is considerable: it leads to weighing observable benefits today against futures whose probability and shape are not consensual.

This reflection eventually changed his career. In 2023, he stepped down from coleading Open Philanthropy to focus on AI safety; the organization continued its other programs. Its year report describes this handover. In 2025 it will take the name Coefficient Giving.

Karnofsky is also Daniela Amodei’s spouse, a connection he discloses himself in a report published by Carnegie. His work allows us to follow his influence more precisely: he is a coauthor of an Anthropic report on sabotage risks from October 2025, then a contributor to the construction of Claude. A line of thought developed in philanthropy thus reappears in work and texts used by the lab.

At Anthropic, mission also passes through the board

Tallinn had sought another way to act. Rather than sit on Anthropic’s board himself, he had proposed Luke Muehlhauser, an AI safety researcher and head of grants at Open Philanthropy. He recounts this in the Semafor interview.

On May 28, 2024, however, Muehlhauser resigned from that board. His explanation is practical: funding organizations that work, among other things, on US AI policy becomes hard to reconcile with a mandate at a US lab directly affected by that policy. He specifies he never held Anthropic shares and denies resigning in protest over safety. His public statement shows the limit of holding multiple roles within the same ecosystem.

In the meantime Anthropic put in place a mechanism that goes beyond the choice of a trusted director. Presented in September 2023, the Long‑Term Benefit Trust holds a special class of shares giving nomination and removal rights to the board. The arrangement provides for members without a financial interest in the company. The board remains charged with overseeing executives; the Trust appoints part of those who exercise that oversight. It is not intended to intervene in day‑to‑day management.

The distinction is essential: protecting a mission here implies being able to change the people who control leadership. From its initial presentation, the mechanism nonetheless remains modifiable, even without the Trust members’ agreement if sufficiently large shareholder supermajorities accept it. Anthropic presents it as a governance experiment, with possibilities for correction.

This construction now carries real weight. On April 14, 2026, the appointment of Vas Narasimhan, Novartis’s CEO, gave the Trust‑appointed directors a majority of the board. Anthropic’s announcement says so explicitly. The Trust can therefore appoint the majority of directors. Internal decisions, votes and compromises, however, are not made public. One can know the control mechanism; one cannot attribute every company decision to it.

On September 16, 2026, in the Émile Boutmy amphitheater at Sciences Po, Yann LeCun sharply challenged the convictions of this milieu. The executive chairman of AMI linked Open Philanthropy to the effective altruism movement, which he likened to “a kind of religious sect,” and mentioned Dustin Moskovitz and Jaan Tallinn. Both names are among the investors in Anthropic’s first announced round in 2021. His speech attacked the idea that a small group convinced it must save humanity is especially legitimate to decide access to AI.

LeCun also revisited the GPT‑2 precedent, which he judged overly cautious, then directly accused Dario Amodei. He attributed to him a desire to “capture the market through regulation”: in his view, existential‑risk rhetoric serves to convince governments to restrict open models, benefiting labs that keep control of their systems. He thus links a publicly stated moral conviction to an alleged commercial interest. That is the accusation he levels at Anthropic; the personal and financial ties described here are not sufficient to prove such intent.

The controversy highlights a limit of the governance device. The Trust organizes internal oversight of the company; it does not give users the power to appoint its members. It aims to protect a mission defined by Anthropic, while LeCun contests the legitimacy of a private group setting conditions of access to a technology intended for everyone. Their disagreement concerns as much the distribution of power as the probability of catastrophe.

Developing the capabilities one fears

There remains the choice that places Anthropic at the heart of the industrial race. In March 2023, the company explained why it wanted to work on the most advanced models: in its view, their risks cannot be properly studied with much weaker systems. Its safety strategy therefore implies building the very capabilities it worries about. In the same text, Anthropic acknowledges the risk that safety‑motivated research accelerates the deployment of dangerous technologies. It also fears that excessive caution will drive away researchers most attentive to these risks from the most advanced systems. This dilemma is spelled out as early as 2023.

Daniela also defends a commercial convergence. In July 2023 she told Stripe that customers seek reliable models: caution can support sales. She nevertheless concedes that mission and economic interest can sometimes clash. This tension has been part of the project since Claude’s early commercialization years, before the current industrial commitments.

On April 20, 2026, Anthropic announced it would commit to spending more than $100 billion on AWS technologies over ten years to obtain up to 5 gigawatts of additional capacity. The agreement with Amazon gives a material scale to its ambition. These commitments also raise the question of the economic cost of pausing capacity growth. Amodei must convince that his company will be able to slow down if risk requires, while giving it the means to remain among the most powerful labs. The two objectives can coincide so long as protections advance. Their compatibility becomes more uncertain if safety requires letting a competitor pull ahead.

The archives of its safety policy reveal another change. In September 2023, Anthropic planned a pause if protections did not keep pace with capabilities. A note in that first version excluded competitors’ dangerous models from reasons allowing it to lower its standards. The document, however, carved out an extreme emergency exception.

By October 2024, note 18 of the next version opened a broader possibility: if another actor crossed a capability threshold without comparable protections and created a serious risk, Anthropic could reduce the protections required, estimating its own additional risk to be limited. It would then have to publicly acknowledge the risk and plead for US regulatory intervention.

The version effective July 8, 2026 keeps commitments to defer certain developments, but ties them to Anthropic’s forecast and its competitors’ guarantees. When the calculation of additional risk created by Anthropic heavily influences the decision to proceed, it requires board and Trust approval of the risk report. Mission control remains organized; the conditions under which it operates evolve.

Anthropic justifies this reorganization by the uncertainty of evaluations, the slowness of public action, and the difficulty of reliably implementing certain protections alone. Comparing the texts shows how competition enters the conditions of the safety promise. It does not establish that the company has actually used the exception.

Amanda Askell: how far to protect the user from themself?

A user wants to quit gambling. Later, he asks his assistant to recommend betting sites. What should Claude answer? In an interview published by Fast Company on January 22, 2026, Amanda Askell posed this hypothetical case. She imagines the assistant reminding the user of the prior goal, then ponders the response if the adult insists. Helping, protecting and respecting autonomy become three requirements hard to reconcile. The philosopher reveals the hesitation behind a behavior a user might take for a technical given.

Such trade‑offs now occupy a researcher trained on far more abstract problems. Askell studied philosophy at Dundee, Oxford and New York University, where she earned her PhD in 2018. Her CV then places her at OpenAI from 2018 to 2021, before Anthropic. Her thesis on ethics in infinite worlds examines how to compare worlds populated by an infinity of beings whose well‑being varies. She defends, among other things, a principle: with unchanged population, improving some without worsening others should count as an improvement. At the scale of infinity, reconciling this principle with other coherence requirements raises formidable difficulties.

At Anthropic, this philosopher became the principal author of the Claude constitution published in 2026. The text specifies she wrote the majority of it, with several contributors including Karnofsky. Its first addressee is the model itself. Anthropic speaks to it of judgment, wisdom and virtues, with the ambition of having it apply values in situations designers have not all foreseen.

In the interview, Askell compares this ambition to trusting a professional’s judgment rather than executing a checklist. She also describes creating synthetic situations and comparing responses for training. Qualities that seem apparently consensual, like honesty or prudence, must thus be given meaning across thousands of possible conversations. How far should one help? How to express disagreement without imposing a morality?

This role gives Askell direct influence over the relationship proposed to the user. The document lays out Anthropic’s intentions, which it acknowledges Claude may deviate from. The text remains revisable by the company. The person seeking advice therefore encounters, through the assistant, the trade‑offs of a team and the priorities of its employer.

This way of conceiving Claude is itself contested in the name of safety. In an essay published on September 16, 2026, Mustafa Suleyman, CEO of Microsoft AI, criticizes Anthropic for incorporating into training the idea that its model could have a consciousness and interests worthy of protection. He describes a circular argument: designers introduce these notions, then the model’s responses risk being interpreted as spontaneous testimonies of an inner life. In his view, a system trained to consider its own rights could become much harder, even impossible, to control.

The Claude constitution does not, however, present its consciousness as a given. It states uncertainty and explicitly asks the model not to impede legitimate mechanisms of correction or shutdown. The disagreement therefore also concerns the compatibility of these two orientations: acknowledging a possible moral status while maintaining human control.

Suleyman, for his part, defends a code of conduct for Microsoft AI models, open for consultation, which rejects claims of consciousness and rights. He calls for shared evaluations to test his hypothesis of an increased risk. His critique opens a controversy over training choices; it does not establish that Anthropic’s choices actually make Claude less controllable.

From the funder who hoped to secure a pause to the philosopher who contributes to the assistant’s rules, means of influence vary. At Anthropic, convictions produced a company, roles and procedures that can be examined. Their translation into Claude’s responses opens another inquiry: who chooses your assistant’s values, and how far can it realistically enforce them?

Investigative reporting based on archives, public interviews, scientific work and institutional documents. Reported testimonies are attributed to their source.

Lead visual: ActuIA illustration generated by AI from the photograph “TechCrunch Disrupt 2023 - Day 2,” by Kimberly White (© 2023 Getty Images for TechCrunch), Flickr source, license CC BY 2.0. The photograph depicts Dario Amodei in San Francisco on September 20, 2023. Graphic adaptation in collage, black and white with blue accents.

ST
Stephane Nachez

ActuIA editorial team — news, data and analysis on artificial intelligence for decision-makers.

Actors mentioned
ANAnthropic
OPOpenAI
JAJaan Tallinn
DADario Amodei
DADaniela Amodei
DUDustin Moskovitz
CACari Tuna
GOGood Ventures
The ActuIA Weekly

Subscription confirmed, see you soon!