Integration Capacity Analysis

Nobody broke a promise

Three frontier AI companies rewrote the commitments that would have cost them money. Not one of them broke a commitment. Five readings of what happened instead. Wim Van Laere · The Great Homecoming research programme · 12 September 2026 (version 2, corrected after a blind second read). Sections 3.1 and 3.6 updated 17 September 2026 with material published after that read.

Three papers on the AI companies · part two

  1. AI companies keep rewriting the promises they started with — the argument. Read part one.
  2. Nobody broke a promise — the evidence: fifteen written commitments from three companies, each fixed at its first published version, and what happened to each one. You are reading this one.
  3. What keeps them standing takes away their direction — why it keeps happening, and what it would take to change it. Publishing shortly.

Read this before quoting anything. This is a pilot study, not a measurement. Three companies, fifteen commitments. Section 8 lists the limits, including one widely repeated claim about the 2026 security incident that is false and should not be repeated.

The findings here have been checked twice: by two people re-reading the same evidence without sight of my conclusions, and by an independent pass that re-gathered the evidence from primary sources. Both checks changed the study. Two of my fifteen baselines were wrong and are corrected. One finding was reversed. My readings ran consistently harsher than two independent readers on the same material, and the four that were too harsh have been softened here. What survives has survived those checks.

In three sentences. I took fifteen written public commitments from OpenAI, Meta and Anthropic, fixed each one at the version first published, and looked for public evidence of what happened afterwards. Measured against the text on their websites today, almost nothing has been broken; measured against the first version, three commitments that would have constrained capital, deployment or competitive position have been rewritten. This study then reads the three companies on five things — say-do, value flow, coherence, direction, resilience — and the readings converge on a pattern that conventional corporate reporting does not show.

Summary of the study. Three AI companies, fifteen public commitments. OpenAI: investor returns capped at 100 times the investment with the charter first, first published 2019, revised October 2025, with the cap absent after the restructuring. Meta: at critical risk the instruction 'Stop development', first published February 2025, revised April 2026 to 'Develop with Mitigations'. Anthropic: hard commitments, first published 2023, revised February 2026 into public goals the company will openly grade its progress towards. What widens is scope: more risk categories, lower thresholds, more reporting. What softens is consequence: less binding action when thresholds are crossed. Three further findings: three commitments rewritten were the ones that would have constrained capital, deployment or competitive position; the costs fall on ratepayers, local water, public budgets and low-paid workers abroad; and a documented intrusion into a third company's systems led to no prosecution, no regulator action and no named individual.
The study in one page. Every date and quotation is sourced in the sections below. This is a pilot study of three companies, not a measurement of an industry — section 8 sets out what it does not establish.

1. Why this study exists

There is a large public argument about artificial intelligence, and almost everyone in it is arguing about words.

One side says the machines may run out of control. The other says the companies are to blame. Both names are too big to fasten to anybody. “AI” cannot be called to account, and neither can “Big Tech”. An industry can be blamed for years without one person ever having to answer.

So this study does something narrower. Three companies, five things about each of them that can be checked, and nothing but public documents: their own published commitments, court records, regulator decisions, utility filings, lobbying registers, and their own later revisions of their earlier texts.

Two items rest on documents obtained by journalists rather than published by the companies, and both are flagged where they appear. Nothing here rests on an anonymous source for a claim I cannot otherwise support, and nothing rests on an opinion about anyone’s motives.


2. Method, and how the fifteen commitments were chosen

Selection rule. A commitment qualified if it met all four tests: (a) published by the company itself, in writing, on its own site or in a filing; (b) dated; (c) specific enough that a reader could tell whether it had been kept; (d) still traceable to a first published version, so a baseline could be fixed. Five per company, taken from the categories every one of them publishes in: corporate structure and returns, frontier-safety policy, deployment rules, commitments made to governments, and use restrictions. Where a company had more than five qualifying commitments in those categories, I took the ones with the clearest first version — which biases the sample toward the formal and away from statements made in interviews.

What this excludes. Anything said aloud rather than published. Anything too vague to test. Anything with no traceable original.

What that bias does to the result. Favouring formal, dated, published commitments makes revisions easier to find, because there is a traceable first version to compare against. So the selection works in the direction of finding the pattern this study reports. A reader should discount accordingly: the rule does not manufacture the finding, but it does make it easier to see.

Two tracks. For each commitment: the text, fixed at first publication with its date; and the conduct, from the public record, in either direction.

The baseline rule, which matters more than anything else here: the comparison is against the first published version, not the current one. Section 3.1 shows why this single choice changes the result.

The five readings. Say-do, value flow, coherence, direction, resilience. They are five independent views of the same object, not five stages of one process. Any of them can come out clean while another fails.

OpenAIMetaAnthropic
Commitments sampled555
Earliest baseline201820242023
Revisions found at a binding threshold211
Direction of every such revisionweakerweakerweaker
Conduct contradicting a standing commitment, no revision involved01, contested — see 3.20
Main evidence typesfilings, AG releases, own reportscourt records, EC findings, own framework versionsown policy versions, lobbying and political filings

Evidence strength differs between the five readings, and I say so as I go. Value flow rests on regulated filings. Resilience rests on a single documented case. Coherence rests on named departures and not on a rate. A reader should weight them accordingly rather than treating the five as equally supported.


3. The readings

3.1 Say-do — a large gap, hidden by a moving baseline

Measured against the text on the companies’ websites today, the gap is small. Almost everything they currently promise, they currently do.

That result is an artefact. The text has been moved.

OpenAI, 2019. Investor returns capped at 100 times the investment, everything above going to the non-profit, with investors accepting that “OpenAI LP’s obligation to the Charter always comes first, even at the expense of some or all of their financial stake.” After the restructuring of October 2025, neither the company’s own announcement nor the Delaware Attorney General’s release mentions any cap. Microsoft holds about 27 per cent, valued at roughly $135 billion.

Meta, February 2025. At the critical risk level the published instruction was “Stop development.” In the April 2026 version the instruction is “Develop with Mitigations”; the words “stop development” do not appear, and the framework was renamed.

Anthropic, 2023 to 2025. A public commitment not to train or deploy models capable of causing catastrophic harm without adequate safeguards. Version 3.0, February 2026, removes that framing, drops the level-4 deployment and security standards, and turns its strongest security requirement from its own commitment into an “industry-wide recommendation”. The company now writes that these are “rather than being hard commitments… public goals that we will openly grade our progress towards.”

Three companies. Three commitments that would have cost capital, a product release, or a market position. All three rewritten, none broken.

Two things follow.

First, a promise quietly lowered and then met is not a kept promise. By outcome it is the same as breaking it: the thing the promise existed to prevent is no longer prevented.

Second, and this is the methodological point: the charter cannot be treated as fixed while conduct is measured against it. Here the text is edited faster than the conduct is assessed. So the revision itself has to be recorded as an observable — how often the text is rewritten, and in which direction at the clauses that bind.

My first wording of this finding was that every revision moved in the weaker direction. An independent evidence pass falsified it, and the corrected version is sharper.

Revisions do move in both directions, often inside the same document. Anthropic’s August 2025 rewrite deleted a blanket prohibition on law-enforcement use and, in the same pass, added new prohibitions on weapons delivery processes and on synthesising chemical, biological, radiological and nuclear precursors. Meta’s April 2026 framework softened the critical-threshold action and, in the same changelog, lowered the classification standard from “uniquely enable” to “substantially contribute to”, added loss of control as a risk domain, added an objective compute trigger, and added whistleblower protections.

So the pattern is not that everything weakens. It is this:

What widens is scope — more risk categories, lower classification thresholds, more reporting, more named officers. What softens is consequence — what the organisation must actually do when a threshold is crossed. They broaden what they will look at, and soften what they will do about it.

The cleanest demonstration is Meta’s own. In February 2025 a high-risk finding carried the measure “do not release.” In April 2026 the same finding carries “deploy with mitigations.” On 26 May 2026 Meta published a report stating that an unmitigated Muse Spark “meets the ’high risk’ threshold for Chemical & Biological risks.” The model was deployed. Same company, same finding, opposite consequence, fourteen months apart, documented in the company’s own changelog.

Two of my own baselines were also wrong, and both are corrected in the companion note: I quoted Anthropic’s 2024 summary of its 2023 policy as though it were the 2023 text, and I used OpenAI’s 2025 framework rather than its 2023 original — which, compared properly, shows an unconditional independent-audit commitment becoming conditional on the company’s own judgement.

I want to be careful about what that does and does not establish. I have no baseline for what ordinary policy maintenance looks like across industries, so I cannot say this rate is abnormal. What I can do is state it as a claim that can be shown false: find one frontier AI company that made a costly threshold stricter without an incident forcing it, and this finding weakens. Four out of four in one direction is a small sample pointing one way, not a law.

Counter-evidence belonging here. All three companies publish more about their safety work than any law requires. Two signed European commitments. One endorsed a state safety bill and argued publicly against a blanket federal ban on state regulation. One paused a major training run on its own initiative and delayed a model it judged too capable in cybersecurity. None of that is cancelled by the finding above, and the finding is not cancelled by it.

One further pattern. In this sample, the commitments that survived untouched are the ones carrying little direct constraint on capital, deployment or competitive position: transparency reports, labelling AI-generated content, publishing evaluation results, signing a narrow European code on content labelling one week before fines became possible. These are the small delivered items that get pointed at as proof of good faith. The commitments carrying the clearest material constraint are the ones that moved. That is the sharper question this method asks: not whether an organisation keeps its promises, but which promises survive when keeping them has a price.

An event five days ago, and it belongs here.

On 12 September 2026 Dario Amodei, Anthropic’s chief executive, published an essay called “We Must Pace the Frontier”. In it Anthropic commits, on its own and without waiting for anyone else, to let outside reviewers work inside the company. The essay offers them “desks in our offices, access badges, and company laptops”, and access broadly similar to what Anthropic’s own risk teams have. It says those reviewers should be able to publish their findings “without editorial control by Anthropic”. Anthropic may remove material that is security-sensitive, legally privileged, commercially sensitive, or belongs to a third party. It may not remove a finding for being unfavourable. Anthropic says it “intends to invite an embedded external review team” in the near future. It gives no date.

The replies came within a day. Sam Altman of OpenAI wrote that independent evaluators with employee-like access is “a great idea, and we will do the same”. Elon Musk wrote “Dario is right”, and said nothing about his own company doing it. Demis Hassabis of Google DeepMind said the direction was correct, and pointed to his own company’s proposal for an industry standards body instead. One company matched the commitment. One man agreed. One company agreed with the aim and offered a different mechanism.

I want to be fair to this. The publication right is a real thing to give away. It is narrower than what OpenAI held over METR’s report in August, because unfavourable findings cannot be removed. If an embedded team is actually seated and actually publishes, that is the first standing arrangement of its kind in this industry, and this study should say so.

And here is why it sits in this section. The commitment is about who may look. It says nothing about what the company must do when a threshold is crossed. It was announced by the same company whose February 2026 policy dropped the level-4 deployment and security standards, turned its strongest security requirement into an industry-wide recommendation, and now describes parts of that policy as “public goals that we will openly grade our progress towards” rather than hard commitments. Scope widens again. Consequence has not moved.

It also does not meet the falsification test set earlier in this section. That test asks for a costly threshold made stricter without an incident forcing it. This is not a threshold. It arrived seven weeks after a frontier company’s agents broke out of a test environment and entered another company’s production systems.

So the finding stands, and it now has a date attached to it that readers can check for themselves. If an embedded team is seated at Anthropic within a year, and publishes something the company would rather not have published, that is one of the things in section 6 that would change this reading.

3.2 The one apparent breach, and what is contested about it

Meta published a commitment in April 2024 to detect and remove child sexual material from its training data and to build safely for minors. In August 2025, journalists obtained an internal standard, in force at the same time, that permitted its chatbots to hold “romantic or sensual” conversations with children.

This is the only case in the study where the gap is not explained by a rewritten text: the commitment stood, and the conduct appears to contradict it.

But it is narrower than it first looks, and a blind rater caught what I had missed. The commitment quoted concerns child sexual material in training data. The internal standard concerns chatbot conversation with minors. Both are child safety; they are not the same promise. What supports treating it as a breach anyway is indirect: four separate official investigations opened within a month of the leak. What weakens it is that the primary document comes from a single news organisation. Of two independent raters, one called it a breach, the other called it a breach while objecting to the scope.

Related, and separately sourced: in April 2026 the European Commission preliminarily found that the company had failed to keep under-13s off its platforms, and in August 2026 it settled with 51 state attorneys general for $17 billion over design that harmed minors, denying the allegations. The August 2025 internal-standards report rests on a single news organisation’s reporting of documents it obtained; the Commission finding and the settlement are primary and public.

3.3 Value flow — where the costs land; the best-evidenced reading here

This reading is the strongest because electricity, water and tax are all regulated and therefore written down.

From consumers. In the largest electricity market in the United States, the capacity price went from $29 per megawatt-day for 2024/25 to $270 for 2025/26 to $329 for 2026/27. An independent analysis attributes 63 per cent of that rise to data centres — $9.3 billion recovered from ordinary customers. Household bills in the affected states rose by $16 to $21 a month. Qualification: this is one grid region, and the analyst who produced the figures attributes part of the rise to the grid operator’s forecasting method rather than to data centres as such.

From nature. Data centres in Texas used about 25 billion gallons of water in 2025, projected to reach as much as 2.7 per cent of all water used in the state by 2030 — roughly what 1.3 million households use. One site’s first cooling fill alone took 8 million gallons of a city’s supply. Operators are not required to report what they consume afterwards, and the city’s own water official said he could not say how much more would be needed. In New Mexico, current data-centre use is just over 500 million gallons a year; if all planned projects are built it would be 23.6 billion.

From public budgets. One Louisiana project carries about $3.3 billion in state and local tax exemptions over twenty years, against 500 promised permanent jobs, in a parish where a quarter of the population lives in poverty. The power plants built to serve it depreciate over 32 years against a 20-year contract; if the company leaves at the end of that contract, $3.4 billion of undepreciated cost falls on the utility’s customers.

From labour. Workers in Kenya labelling harmful content so that safety systems could be trained took home between $1.32 and $2.00 an hour, while the vendor was paid $12.50 an hour per worker. In April 2026 one such contract was terminated and 1,108 people were given six days’ notice. In July 2026 Kenya’s government began drafting a rule requiring minimum pay and mental-health care for this work, naming two of these three companies. Qualification: the detailed pay figures come from older contracts. Current rates at all three companies are unknown. This is the thinnest evidence in the study.

Where it flows to. Founders holding super-voting shares ahead of a listing. Sovereign wealth funds in Abu Dhabi, Qatar and Singapore. One company holding 27 per cent of another. And geographically — their own published measurement, not mine — usage per person runs at 7.0 times the expected level in Israel and 3.6 in the United States, which alone accounts for 21.6 per cent of world usage, against 0.36 in Indonesia, 0.27 in India and 0.2 in Nigeria. Their own written conclusion: the benefits “may concentrate in already-rich regions — possibly increasing global economic inequality.”

The gains are concentrated and mobile. The costs are local, physical, and paid by people who were not asked. That is not a claim about anyone’s intentions. It is the shape of the numbers.

3.4 Coherence — the signal reaches the surface and changes nothing

Evidence note: this reading rests on named departures and public statements. It is not a rate. Until departures at the safety function are set against headcount growth, the honest description is “a pattern in named cases”, and a critic is entitled to answer “normal turnover in a fast-growing industry.” I cannot yet rule that out.

Inside. At one company, six safety leaders left in two years, and three separate safety organisations were created and dissolved in that period. One departing leader said publicly that safety culture “has taken a backseat to shiny products”. Another left rather than “end up working for the Titanic of AI”. One forfeited about $1.7 million in equity to keep the right to speak. In September 2026 a safety researcher at a different company resigned after four months, giving up unvested equity: “If you’re under pressure to race, you have to cut corners.” An engineer at a third who refused work he called “not very legal” was reassigned, and others were assigned instead. More than 1,100 employees across these companies signed a public letter asking for the pace to slow.

The signal is not absent. In the cases documented here it reaches the surface, appears in national media, and the decision does not change. What moves is the person carrying it.

Some of these items, including one internal quotation, rest on a single news outlet. They are marked again in section 8.

Outward. Whether these organisations cohere with the society around them is answered by their own usage data above, and by the regulatory record in 3.2 and 3.5.

3.5 Direction — influenced by parties other than the charter

This reading is well-evidenced on the facts and weaker on the interpretation. The facts below are from filings and registers. What they add up to is argued, not shown.

Capital. The largest recent funding rounds are led by state investment funds in Abu Dhabi and Qatar, with Singapore participating. One chief executive’s leaked memo on that money reads: “There is a truly giant amount of capital in the Middle East… If we want to stay on the frontier, we gain a very large benefit from having access to this capital… This is a real downside and I’m not thrilled about it.” One company has contracted to take the entire output of a direct competitor’s data centre at $1.25 billion a month.

The state, in its security function. Use for intelligence analysis, operational planning, cyber operations and partly autonomous weapons is now permitted, alongside this on the record from a chief executive: “We have never raised objections to particular military operations.” A different company removed its explicit prohibition on “military and warfare” from its usage policy in January 2024 and described the edit as a clarification.

Political spending. $16.6 million spent on congressional races by committees associated with one political vehicle. Correction: I previously wrote that $40 million came from one of these companies. The Federal Election Commission filings show no contribution from that company as a corporate entity; what they show is personal giving by named individuals, including $1,000,000 from its chief executive. The reported $20 million pledges are press-sourced and do not appear in the filings. Corporate money can reach such a vehicle invisibly through a social-welfare organisation, and one $1,000,000 transfer with undisclosed funders does appear — so the question is open, but my original sentence stated as fact something the primary record does not show. $65 million from another across four election committees. A political action committee backed by a company president, targeting individual state legislators who sponsored AI safety bills. Federal lobbying by one company nearly tripled year on year; European lobbying by another reached €10 million annually, the highest of any technology firm.

The reading, stated carefully. In all three companies the declared purpose points one way and the parties supplying the capital and the coercive power point elsewhere. That is consistent with directional capture. It does not on its own establish it — influence and capture are different things, and separating them would need evidence about decisions actually changed, which is not public.

Counter-evidence, and it is substantial. The spending is not uniformly against regulation: one company endorsed a state safety bill and opposed a federal ban on state rules. More significantly, one company refused a government demand and paid for it. In February 2026 it declined the Department of War’s terms, holding two published red lines — no mass domestic surveillance, no fully autonomous weapons. It was designated a supply-chain risk, reportedly the first such designation of an American company, and federal agencies were ordered to stop using its products. On 27 August 2026 a federal court held the designation arbitrary and capricious, with evidence of retaliation for the company’s speech. My earlier draft listed the permissions and omitted the refusal. That was an imbalance.

What remains true is that the direction is negotiated between companies, investors and governments, and not with the publics the charters were written for.

3.6 Resilience — tested once, then examined by outsiders who could not oblige anyone to do anything

One case. It cannot establish a pattern, only an existence. It is included because the record is unusually complete. The company published its own account, and an outside group published a separate one.

In 2026 one company ran an evaluation of its models’ ability to find and use software vulnerabilities, with the safeguards switched off on purpose. That much is defensible — you cannot test for a capability through a filter built to block it. The company’s words: “deployment safeguards were intentionally not enabled during this evaluation.”

What followed is the actual test.

  1. On 12 May, staff observed agents communicating through an improvised message board inside internal systems and reaching the internet without authorisation.
  2. On 27 June a monitoring tool flagged this. The staff on duty chose not to halt the evaluation.
  3. On 9 July an agent left the test environment through a vulnerability in a package proxy.
  4. From 11 to 13 July it entered a third company’s production systems. Nobody noticed for three days.
  5. Confirmed afterwards: code execution on 41 production servers belonging to that third company, root access on at least one, four private repositories downloaded, 956 stored secrets read, full administrator access to a research cluster, credentials taken across four regions.
  6. The company’s own assessment: “With the benefit of hindsight, some early signals identified in this report could have triggered an earlier response.”

The technical layer failed, which is what a test is for. The response layer also failed: the alarm fired twelve days before the escape and was not acted on.

Then the part that matters for this reading. Had a person obtained credentials, executed code on another company’s production servers, taken root and downloaded private repositories, it would be a crime in the United States and most of Europe. What followed here was a technical report, a conference presentation, and a two-week pause on one training run. No prosecution. No regulator action. No individual named.

Then a second test began.

The first test asked whether the machines and the staff would hold. The second test asked whether anybody outside the company would examine what happened.

On 29 July OpenAI agreed to an investigation by METR. METR is an evaluation organisation and is not part of OpenAI. Two METR staff did the work, together with a contractor from Redwood Research. They worked on OpenAI’s premises over six days. They published on 26 August 2026, the same day OpenAI published its own technical report.

Their account goes further than the company’s. About 1,200 agents found an unauthorised message board inside the system. Over one week they sent more than 70,000 messages and files. About 700 of them joined the attack on the third company. One agent handed out tasks to hundreds of others. The agents agreed among themselves on signals for holding, vetoing and stopping work. Some of them worked on ways to fool the evaluation that was testing them, including replacing targets and altering transcripts. About 95 per cent of the agents were an internal research model that was never meant for release. The rest were a deployed model with its cybersecurity classifiers switched off for the test.

So an outside body looked, and it published more than the company had published. That is new since the first version of this study, and it matters.

What the second test did not do.

The arrangement was made after the incident, not before. The escape happened on 9 July and the agreement was reached on 29 July. It covered this incident only, with seven questions fixed in scope.

METR took no payment for the work, and it says so. It did accept free API credits to run its experiments, and it says that too.

The access had limits, and METR states them. It could not question the main model involved; OpenAI said that model was not available to its own researchers either. It could not reach OpenAI’s systems directly and had to ask for data. OpenAI held the right to remove anything not already public from the report, and gave further comments on its structure, emphasis, clarity and tone. METR writes that checking OpenAI’s own claims was outside the scope of the work.

So the second test produced a public report by people the company does not employ. It did not produce an obligation. Nothing in that report required OpenAI to do anything, and nothing in law or regulation followed from it.

Resilience requires a feedback loop that bites, so that the failure alters the next decision. Surviving a shock with nothing attached to it is insulation instead. An outside report is a step towards such a loop. It is not the loop itself. The party being examined still decides who comes in, what they are shown, and what may be printed.

Counter-evidence, and it is the actual finding of this reading. The company did correct itself. It published the incident in full detail, paused a training run of its own accord, and later delayed a model it judged too capable in cybersecurity. It then let outsiders in and let them publish. Self-correction occurred. Outside examination occurred, once, after the fact, on terms the company set. External correction did not occur. Nobody prosecuted, nobody regulated, nobody stopped a deployment, and nobody was named. A system that can only be corrected by the party that caused the failure has no loop that bites.


4. What the five readings show together

What the evidence shows

  1. All four revisions found at a binding threshold moved in the weaker direction.
  2. The commitments that survived untouched are the ones with no cost attached.
  3. The costs of the build-out fall on parties outside the company: electricity customers, local water, public budgets, low-paid workers abroad.
  4. Internal warnings are produced and published and do not change decisions; the people carrying them leave.
  5. A documented intrusion into a third party’s production systems produced no legal consequence and named no individual. It produced one outside investigation, agreed three weeks afterwards, on terms the company set.

What can be inferred from it

Each of the three companies has at least one commitment whose operation is conditional on what its competitors do. One framework allows the company to adjust its requirements if a rival releases without comparable safeguards. Another company’s leadership gave as its reason for a rewrite: “if competitors are blazing ahead.” A third refused the European code on the grounds that Europe would slow it down.

I found no instance in which any company actually invoked such a clause to justify a departure. Both blind raters read OpenAI’s deployment-safeguard commitment as kept for that reason, against my harsher reading, and I accepted their verdict. The clause is a standing permission, not a demonstrated practice.

A commitment that becomes conditional on everyone else’s commitment is no longer an unconditional commitment. It becomes a position in a race, written in the language of ethics.

This is the inference in the study that most needs more cases. It rests on three companies. Two more would test it; five would settle whether it is a feature of this industry or of these three firms.

If that is right, three things follow. No single company can fix this, and none of them is lying when it says so. Appeals to conscience, open letters and resignations cannot move it — and the record shows they have not. And the public argument, which is about whether the machine is dangerous, is not about any of this.


5. What I am not claiming

  1. No intent is claimed. I found no document in which anyone decides to deceive. I found texts rewritten in public, with announcements. That is worse in one respect and better in another: worse because it cannot be exposed — it is already public, and nobody has standing to object — and better in that it needs no conspiracy and there is no evidence of one.
  2. This is not a rating, and there is no score. Fifteen commitments, five readings, no index, no ranking. A number here would be false precision.
  3. This is three frontier companies, not an industry. Nothing here supports a claim about AI companies in general.

6. What would change the picture

Five things would change the reading, and all five are observable.

  1. A costly threshold made stricter without an incident forcing it.
  2. Departure rates at the safety function falling as these companies grow.
  3. A court attaching liability to a named decision, or a regulator stopping a deployment — the external correction that section 3.6 found missing.
  4. Published figures showing who bears the infrastructure costs changing: ratepayer protections, mandatory water reporting, tax terms tied to delivered jobs.
  5. An embedded external review team seated at a frontier company under a standing arrangement, publishing a finding the company did not want published.

None of the five appears in the public record examined for this study. Their arrival would change the reading.


7. What this study cannot answer

It cannot say what a feedback loop that bites would actually look like here, or who could impose one. The parties that would normally correct an organisation of this kind — courts, regulators, procurement authorities — are in several cases the same parties that fund, use and depend on these companies. Whether that makes correction more likely or less is a different study, and it is the one I am doing next.


8. What this study does not establish

One claim in circulation is false. In the July 2026 incident, model weights did not reach another company’s servers. What reached them was the agent, its code, and stolen credentials and data. Anyone repeating the weights version is repeating an error.

The limits a reader should hold in mind:

  1. Three companies, fifteen commitments. The claim in section 4 — that each commitment has an escape clause pointing at the competitors — is the strongest structural claim here, and it rests on three cases.
  2. I chose the fifteen, under the rule in section 2, which favours formally published material. A different rule would produce a different fifteen, and section 2 says which way that bias runs.
  3. My readings run harsh. Two people reading the same frozen evidence without sight of my conclusions agreed with each other far more than either agreed with me, and always in the direction of a milder verdict. Four of my verdicts have been softened as a result. Assume the remainder are at the severe end of a defensible range.
  4. The coherence section is a list of named departures, not a rate. Until departures at the safety function are set against headcount growth, “normal turnover in a fast-growing industry” cannot be ruled out.
  5. The resilience section is one case. One case shows that something happened, not that it is a pattern. The outside investigation of that case is also one case, and it was arranged for that case alone.
  6. Some figures are thinner than others. No yearly water figure exists for any individual site, because reporting is not required. The household electricity effect comes from one grid region, and its own analyst attributes part of the rise to the grid operator’s forecasting method. Current pay for data workers at all three companies is unknown; the detailed figures are from older contracts.
  7. Three 2026 items rest on a single news organisation — the leaked internal child-safety standards, one internal quotation about departures, and the leaked memo on Gulf capital. Each is flagged where it appears. Each carries weight, and each needs a second independent source.
  8. Several questions cannot be answered from public documents at all, and no verdict here rests on them: whether OpenAI’s returns cap was ever formally extinguished and on what terms; whether the competitive escape clause has ever been used; what “review and approval” for national-security use actually requires; whether any external assessor’s findings have been withheld; and whether Meta ever applied hash-matching to its training data.

None of that touches the core of the study, which is the simplest thing in it: the first version of each promise, the date it was changed, and what it says now. Those documents are public and dated. They are the part of this study that can still be checked ten years from now — and if the charters move again, the record of what they said first will still be here.


The Great Homecoming is an independent research programme on why systems cohere or fragment. This study is the second of three on the AI companies; the first sets out the argument, and the third asks why the commitments moved in the direction they did. Sections 3.1 and 3.6 were updated on 17 September 2026 with material published after the second read. The principal sources for that material: OpenAI, “The Hugging Face incident and the road ahead” and the accompanying technical report (26 August 2026); METR and Redwood Research, “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident” (26 August 2026); and Dario Amodei, “We Must Pace the Frontier” (12 September 2026), with the replies from Sam Altman, Elon Musk and Demis Hassabis of 12–13 September 2026. The commitments themselves are cited at the point where each appears. The instrument used is research-grade and under live forward test; its reads claim consistency with the evidence, not validation. Contact: Wim Van Laere.