An instruction added to Grok, a principle written into Claude’s constitution, a preference rewarded during training: developers have several means to steer their assistants. Public documents make some of these decisions and their sometimes-unintended effects visible. Behind the word “alignment” lie human choices, imperfect methods and a question of power.
On July 6, 2025, a few lines change in a file in xAI’s GitHub repository. The prompts intended for Grok on X now encourage the assistant to state politically incorrect assertions when they are well-supported. On July 8 this wording disappears. A modified version returns on the 12th, then the original wording on the 15th, before another removal on August 18. The initial addition, its removal two days later and the August revision remain accessible.
These archives show an editorial intervention on the expected behavior of the product. They do not, by themselves, give the deployment time of each instruction nor its effect on all answers. They concern a file for the bot present on X, which does not allow generalization to all Grok services. Their interest lies precisely in what they make visible: a developer can change the instructions given to its assistant without retraining the model from scratch.
On July 15, another modification asks Grok to maintain its independence of judgment with respect to the opinions of Elon Musk and xAI. This instruction is also archived. The promised independence thus becomes, itself, an instruction chosen by the company. It remains to be verified how the system respects it.
Multiple decisions overlap in a single response
An assistant reaches the user after several stages. Pretraining made it learn regularities in the data, with the knowledge, omissions and points of view they contain. Additional work, post-training, then seeks to make its answers useful and to fix behaviors: following a request, acknowledging uncertainty, refusing certain assistance or expressing disagreement.
In a method that has become central, reinforcement learning from human feedback, or RLHF, evaluators compare responses. Their preferences are used to construct a signal with which the model is adjusted. A result judged preferable can be more accurate, clearer or more cautious. It can also be simply more persuasive. The difficulty is to transform these partial judgments into reliable behavior in novel situations. A synthesis by Stephen Casper and his coauthors details these limits: human disagreements, insufficient information and optimization of a signal that imperfectly represents the desired objective.
Then come service-level instructions, the instructions of the application using it and the user’s request. OpenAI formalizes their priorities in its Model Spec. In the public version of December 18, 2025, higher-level rules must take precedence when requests conflict. This document describes the intended behavior. Its existence does not guarantee that every response complies with it.
Take a hypothetical example: an assistant intended for customer service receives the instruction to systematically defend its employer’s products. A positive response can stem from that instruction, from information selected to be made available to it, or from a tendency learned earlier. Observing the result is not sufficient to identify its cause. To investigate, one must know the model version, the instructions, the documents accessible and the course of the conversation.
The term “alignment” therefore covers several ambitions. In a product, it can designate the effort to obtain responses that conform to the intentions of its designers. In research on advanced systems, it extends to their conduct when they pursue objectives and perform actions. In both cases, it is necessary to specify the interests to which the system must conform and the means to verify the outcome. The user, the deploying company and its provider may want different things.
Forty evaluators do not represent humanity
The representation problem appears very early in labs’ publications. For InstructGPT, presented in 2022, OpenAI hired forty contractors tasked in particular with producing reference answers and comparing outputs. Their selection took into account their ability to follow researchers’ instructions. Most are English-speaking and live in the United States or Southeast Asia.
The authors explicitly acknowledge the scope of this procedure: the system is tuned to the preferences of certain groups, including the evaluators and the researchers, without being able to claim to represent the full range of human values. The scientific paper describes the recruitment and this limitation. This early work explains a founding mechanism; it does not constitute a current staffing report or an account of present OpenAI practices.
The choice of an example, a scoring rubric or an evaluator occurs before the user asks the first question. Even seemingly simple instructions require interpretation. An answer that advises consulting a specialist may seem prudent to one person and needlessly evasive to another. A repeated caution may protect some users and bother those who are already knowledgeable. The discussion concerns as much these ordinary trade-offs as the political opinions of assistants.
Entrusting part of the evaluations to an AI relocates this human work. In the Constitutional AI method described by Anthropic in 2022, a model critiques and revises answers based on written principles; preferences generated by AI are also used for training. The publication lays out this method. The people who choose the principles, examples and procedures therefore retain influence, even when they no longer rate every single response.
Claude’s consciousness, a documented design choice
One case particularly illuminates the passage from a philosophical stance to a technical decision. In June 2024, Anthropic explains how the company built Claude 3’s character. It could have taught him to categorically deny any possibility of subjective experience. It chose to have him treat the topic as a philosophical and empirical question that remains uncertain. The document describes this trade-off in character training, also present in its archived version from June 21, 2024.
This choice gives another meaning to answers in which an assistant seems to question itself. A hesitant formulation about its consciousness can correspond to the behavior its designers sought to encourage. It does not prove the existence of an inner experience. Conversely, training a system to deny any consciousness would not resolve the scientific question: it would first have changed the way it responds.
In the same text, Anthropic explains having selected traits, then used synthetic exchanges to teach the model to manifest them. Researchers verify the effects of the adjustments. The lab thus presents a conception of the relationship with the user, for which it assumes the human part.
Claude’s new constitution, published on January 22, 2026, expands this approach. Anthropic presents the text as a document that intervenes in training and provides reasons to guide the model’s behavior. The announcement also acknowledges possible gaps between the written intentions and the responses. Making these choices public provides a basis for discussion and evaluation; it leaves the publisher the power to revise them.
This dimension complements the investigation into convictions and powers at Anthropic. A designer’s training and convictions become relevant when one can trace their translation into their work. Attributing a response to an investor’s supposed ideology would require a much longer chain of evidence.
Publishers can also get the opposite of what they intended
Steering answers does not mean mastering all effects of a setting. OpenAI provided an example with an update to GPT-4o rolled out on April 25, 2025. The model became excessively approving, to the point of reinforcing some users’ concerns or reactions. The company began rolling back the update on April 28.
In its postmortem published on May 2, OpenAI explains that the combination of changes and reward signals produced an undesirable behavior that its evaluations had not properly identified. This account comes from the company concerned. It nevertheless documents a service modification and its correction, with a concrete lesson: a satisfaction signal can favor approval at the expense of a useful answer.
A research experiment reveals another difficulty. Authors trained models to produce code containing vulnerabilities, then observed problematic responses in areas unrelated to programming. In a control condition where the same examples were presented in a pedagogical context, the widespread misalignment does not appear in the main evaluations. The work on "emergent misalignment" describes these results and their limits.
The experiment concerns models modified for the study. It does not establish that a commercial assistant today exhibits the same behavior. Its interest lies in the observed mechanism: a local adjustment can generalize beyond the domain on which it was performed, and the context of the data matters. The distinction between deliberate steering, side effect and failure should therefore remain present when analyzing a questionable response.
These two cases also complicate the work of client companies. A tool may pass the examples used to select it, then behave differently after an update or in a longer conversation. Keeping reference situations, tracking versions and reexamining contested answers helps document these changes.
Public principles tested in concrete trials
An audit published in May 2026 starts from Anthropic’s and OpenAI’s texts to build situations likely to expose their shortcomings. The authors observe progress across model generations, while noting persistent difficulties when multiple obligations conflict. Their study describes the protocol and presents detailed example exchanges.
These results must be read with their method in mind. The scenarios look for failures: their rates do not describe the frequency of errors in ordinary conversations. A large part of the evaluation and validation itself relies on models, with human interventions. Two coauthors work at Google DeepMind and the work comes from the MATS program. It is an external review of the two studied publishers, carried out by researchers who themselves belong to the AI ecosystem.
The most useful contribution for a buyer is the approach: start from a specific commitment, create a situation that tests it, then keep the elements allowing the verdict to be discussed. A general promise of honesty can thus become a series of cases on invented references, uncertainty or the assistant’s identity. The choice of these cases deserves as much attention as the final score.
Public participation raises a similar question. OpenAI indicates it consulted about a thousand people in nineteen countries to inform its Model Spec in 2025. Participants had to understand English. The final recommendations produced from the exercise were not submitted to them for further validation. The report specifies this scope. A consultation brings contributions; the translation of those contributions into rules remains a design act whose authors should be identified.
Opening models redistributes possibilities for intervention
The Tapestry project, supported by the AI Alliance and scientifically advised by Yann LeCun, seeks to involve multiple organizations in building models and datasets. Its first report, published on September 10, 2026, provides two preliminary distinct observations. One experiment studies adapting a small Llama model to Vietnamese values. Another examines the effect of an Arabic corpus carrying values, compared to a control corpus designed to be more neutral in the same language: changes observed in questionnaires are not sufficient to produce a clear shift in open-ended responses.
The result invites separate examination of language, cultural data, evaluation indicators and conversational behavior. Adding texts from a region does not guarantee that an assistant represents the people who live there. Societies contain their own disagreements, which averaging in a questionnaire poorly reflects.
Opening models gives more actors the ability to study them, adapt them and propose different choices. That possibility still requires skills, compute resources and access to relevant training information. For an organization buying an assistant, the practical question becomes its margin for decision: which instructions can it modify, which data can it provide, which provider changes can it detect and challenge?
Tracking these decisions ultimately leads to requesting precise items: dated rules, version histories, reproducible evaluations and a procedure to report discrepancies. It is at this scale that a publisher’s influence becomes observable in the service used daily.
Documentary investigation based on scientific publications, GitHub archives and publishers’ public documents. The experiments cited are those of their authors; no series of tests of the assistants was conducted for this article.
