On 2 October 2026, the communications agency Reputation Age published a study carried out with Bubbling, a French platform that measures brand presence in conversational assistants. Presented under the title "Half-year results as seen by ChatGPT and Gemini", it consists of seven pages dated September. Its main figure concerns the half-year accounts of 39 CAC 40 companies: 40% of the figures returned by the two assistants are incorrect. Confusions between indicators also invite examination of the limits of unstructured text and the implicit assumptions of everyday language.
Between 25 and 28 August 2026, after the publication of first-half results, 8,671 conversations were conducted with ChatGPT and Gemini, generating 34,684 responses. An automated "economic journalist persona" asked questions about the published figures, the indicators highlighted by each company, its financial narrative and comparison with competitors. A second scenario asked the engine whether the company was the right choice for a potential client.
Published growth correct in only 42% of cases
The rate of correct retrieval depends on the indicator requested. Net income is correct in 67% of responses, net debt in 66%, revenue in 65%, organic growth in 64%, free cash flow in 62% and EBITDA in 57%. Published growth comes last with 42%.
The study distinguishes two families of errors. The first groups outdated or invented figures, from 19% of responses on net income to 36% on EBITDA. In the second, the engine cites an official company figure that does not correspond to the question, which occurs in 24% of responses on published growth. The authors also note confusions between organic and published growth, or between quarter and half-year. By company, this selection defect affects 46% of checks on BNP Paribas, 43% on Crédit Agricole and 26% on EssilorLuxottica. "The AI finds the right document but takes the wrong figure from it", the document summarises.
The ranking goes from Capgemini (81% correct responses) and Airbus (80%) to Crédit Agricole (32%), EssilorLuxottica and BNP Paribas (35% each). This accuracy ranking publishes only sixteen names, those of the seven best-retrieved companies and the nine worst-retrieved.
Seven growth rates for a single revenue figure
Even Capgemini's half-year press release, the best-retrieved company in the sample, shows how growth lends itself to confusion. Re-read on 4 October 2026, this document of 30 July announces on its first page revenue of 12,082 million euros, up 8.8%, then growth of 11.3% at constant exchange rates for the half-year and 11.6% for the second quarter. The reconciliation table in the appendix adds the published growth for each quarter, 7.0% then 10.5%, and constant-currency growth for the first quarter, 11.0%. The outlook sets an annual target of 8.5% to 9% at constant exchange rates, of which about 5 points from acquisitions.
For the group's revenue alone, the document thus lists seven growth rates, of which only two relate to the past half-year (8.8% published, 11.3% at constant exchange rates), and no organic growth rate for the group. An assistant asked about Capgemini's organic growth finds no direct answer in the press release, and the 11.3% includes the contribution of the WNS and Cloud4C acquisitions, which the text mentions. The study attributes part of the errors to the way issuers structure and publish their information. The Capgemini case shows that this difficulty exists even for the best-ranked, without being sufficient alone to explain the gaps between companies.
Companies the origin of 78% of citations
According to the document, 78% of the sources used by the two assistants refer to company websites (press and investor spaces, results documents), 19.5% to other sources such as press release distribution sites and the SEC, 2.42% to financial sites such as MarketScreener or Boursorama, and 0.08% to the press, limited to Bloomberg and Le Monde. The authors attribute this near-absence to the blocking of AI robots by "almost all" press publishers.
The same document specifies that source citations are "almost exclusively" the work of ChatGPT, Gemini "almost never" citing links. This distribution therefore mainly describes the behaviour of only one of the two assistants. On the messages that companies want to convey, the average is 67% correct retrieval for the three indicators and three messages highlighted by each, from 97% for Michelin to 33% for EssilorLuxottica.
The potential client scenario provides a measure of a different nature, over 1,144 conversations of four exchanges. The share of favourable opinions falls from 5.6% in the first response to 0.5% in the fourth exchange, while that of unfavourable opinions rises from 7.0% to 12.6%.
A method summarised in three keywords
The method is contained in a box on page 3: "digital twins, multi-turn conversations, hybrid coding". Several parameters that condition the result are not included. The PDF publishes neither the questions asked nor the number of figures checked per company or per indicator. The model versions are not specified, nor is the activation of web search, the tolerance used to judge a figure correct (rounding, scope, restated figure) or the share of coding entrusted to humans. The document does not break down accuracy between ChatGPT and Gemini. It also does not say which CAC 40 company was left aside, nor why, and does not publish the complete ranking of the 39 companies.
The two authors have a commercial activity related to the subject. Bubbling presents itself in the study as a platform for managing brand visibility in generative AI and offers on its website a "preliminary AI diagnosis". Reputation Age, a communications and reputation agency, displays on its homepage a consulting offer dedicated to AI. The presentation page concludes that companies must work on "readability for AI, by AI" of their press releases, which aligns with Bubbling's business. Nothing indicates external control of the protocol.
Our reading: the limits of unstructured text
"What is your margin?" In a professional conversation, the question seems simple. Yet it leaves several choices open: gross, net or operating margin, expressed in euros or as a percentage, over what period? Between interlocutors accustomed to working together, context may suffice to resolve these ambiguities. Another reader may understand something else. Everyday language thus relies on implicit assumptions that predate the arrival of LLMs.
The versatility of large language models stems in particular from their ability to process freely formulated requests, with shortcuts, approximations and variable vocabulary. This tolerance makes the same tool usable in many professions, without requiring a formalised query in advance. It becomes a difficulty when a precisely defined figure is expected: the assistant may retain a plausible interpretation and respond without making visible the choice it made. The fluidity of the response then does not allow knowing whether the two interlocutors are talking about the same indicator.
This reading remains distinct from the study's result. The questions asked are not published: it is impossible to determine what share of errors comes from their formulation, from the presentation of documents or from poor processing by the model despite explicit indications. Outdated or invented data also constitute a separate category of errors. Making definitions explicit reduces the room left for interpretation; the accuracy of the figure retained still needs to be verified in the source.
Unstructured text can scatter the definition of a figure between a sentence, a table and a note in the appendix. The Capgemini case illustrates this work of reconciliation between value, period and calculation method. To make financial use more reliable, our recommendation is to explicitly associate each data point with its definition, unit, period and scope, in the source as in the request. A table or structured format helps preserve these links, provided that the fields themselves are defined without ambiguity. Faced with a request for "margin" that is not specified, the assistant should ask which one before choosing a figure.
The harness, and the limit of shadow AI
A technical response already exists: surround the model with a harness adapted to the business. This term designates the software environment that organises its inputs, its calls to tools and the restitution of results, as described by Anthropic. In a financial use, this framework can associate a repository of indicators, approved sources, calculation rules and automatic controls. Its reliability must be measured on business cases, including ambiguous requests and missing data.
To return to the margin example, the device can require the user to confirm the indicator, period and scope before any restitution. It can entrust the calculation to a deterministic tool, check units and keep the exact reference of the document used. If a required piece of information is missing, a software rule can block the restitution of the figure and request clarification or human validation. The effectiveness of this device also depends on the quality of the data and the controls retained.
This is precisely the limit of shadow AI that we retain here: the employee queries ChatGPT or Claude directly without going through the company's harness. The business definitions, validated sources and controls integrated into this device are then no longer applied to their request. Verification rests on the user, who may take a convincing response into a report without having resolved the ambiguities. Access to a performant model is therefore not enough to reproduce the conditions of a use validated by the organisation. In the margin example, an exact figure taken from ChatGPT can thus end up in the wrong column of a table because the user and the assistant had not retained the same definition.
Our articles will then appear first in Google Top Stories.
