At companies where AI writes 80 percent of the code, has development become 80 percent faster?

日本語

Contents of this article

Summary

I hold a hypothesis: what decides the value of software from here on may be the autonomous continuity in which software makes software. To verify that hypothesis, I ran two passes of research through the public primary sources of the major companies that build AI models. The research asked how far each company officially says it uses its own AI models in its own model development, and how far the companies and outside researchers have been able to verify with actual measurement the claims of acceleration, of how much faster development has become as a result.

I read four kinds of sources: each company’s announcements, the documents that gather a provider’s evaluations of performance and safety, published when it releases an AI model (system cards), research papers, and an economic estimate in which independent researchers rework the figures each company published within a calculation framework of their own. I read all of them down to the body text, not just summaries. From the research I found both large numbers that support the hypothesis and conclusions from the AI companies themselves that weaken it.

In this article I go through what the primary documents of the three companies actually say, with the definitions of the numbers attached. I too operate a chain of the same structure, in which software makes software, together with AI at the scale of an individual. I add the measurements from running it, and explain how far this hypothesis can be claimed now and where it stops being claimable.

Let me lay out the structure of the body up front. First, I confirm the exact definition of the figure that says AI writes 80 percent of the code inside a company. Next, I restate the hypothesis as a structure in which the gap between models keeps widening the gap between the next models. Third, I recount how the methods I had settled on earlier for the convenience of my own operation matched the current recommendations of these companies, which I read afterward. Last, I read the judgment in which the parties themselves conclude that they cannot attribute the acceleration of their own progress to AI.

What you can take away

This article is for people who are building a setup where they develop alongside AI agents. It is also for people who want to know from primary sources where the story of AI building models actually stands right now. In this article, I explain two insights. The first is how to read acceleration figures, the numbers for how much faster AI has made development. Each company describes the same acceleration as large in a promotional setting and as small in a safety evaluation. So you must not line up numbers with different definitions and compare them as they are. I show this way of reading with quotations from the original text. The second is a lesson that applies directly to individual development. What I settled on is this: the lower the cost of verification in a domain, the larger the effect the chain shows there. I also explain the conclusion that what has the highest value to hand over to AI is the record of judgment, of why a person decided the way they did.

The problem and what came of it

To check my hypothesis against reality, I read the public primary documents of the major companies and an independent economic estimate down to the body text. After the first pass, I went back over the premises to make my conclusion more exact. My material was those public documents, together with measurements from the running of this platform itself.

The hypothesis

What AI can write moves closer to what anyone can build. If the value of building software itself keeps thinning out under competition, value should remain somewhere else. I think that place is the autonomous continuity in which software, once built, builds the next software. I build this membership platform for a technical blog together with AI: the articles, the search, and the conversation feature alike. Including the chain that records the discussions held here, turns them into content, and returns them to the knowledge of that AI itself, this platform is also a live experiment testing this hypothesis.

Still, a hypothesis is a hypothesis. To check it against reality, I went and read public primary documents down to the body text, to see whether this is really happening at the leading edge and how far the companies and outside researchers have been able to verify it.

What each company officially says

First, I will quote only the official words of the parties involved. In an announcement dated 2026-07-09, OpenAI wrote that over the preceding six months, the share of its research compute allocated to running AI that writes code internally had grown a hundredfold. In the same paragraph, though, it notes on its own that this is a measure of usage and not a number that measures research progress itself. In an official explainer article, Anthropic published that, as of May 2026, its own models wrote more than 80 percent of the code integrated into its codebase. That 80 percent has a footnote as well, which gives it the conservative definition of the share of lines integrated into production that can be attributed to a model, and goes as far as writing that line count is a measure of quantity and not of quality. Google’s chief executive said in a talk that 75 percent of new code inside the company is AI generated and has passed engineer approval. This one is defined as the share of generation that passed approval, so a few lines emitted by a completion feature can enter the numerator. The two figures, 80 and 75, look close, but their definitions differ, so they cannot be compared side by side.

In the same explainer article, Anthropic has also published an internal evaluation in which the rate at which a model beats human judgment on choosing the next research move rose from 51 percent to 64 percent. As for OpenAI, reports carry an announcement that a higher-tier AI model autonomously performed post-training, the additional training done after a model is built, on a lower-tier model. An OpenAI employee added a note to this, though: it did not build the training procedure from scratch but adapted its own post-training configuration for a small model, and for humans that would be about two weeks of work for two researchers. OpenAI’s system card, which gathers its evaluations of performance and safety, also states plainly that reliably designing and executing a complete post-training procedure across diverse models is not yet possible. Even so, one of Anthropic’s co-founders goes so far as writing, in an essay, that it may be one or two years until the point where the current generation of AI autonomously builds the next.

Read only the disclosures up to this point, and the chain in which software makes software looks like it is already functioning in full.

The structure where the gap between models widens the next gap

The competitive picture is not decided by whether a company has the chain in which software makes software. Because every company has begun to depend on the AI models that carry the chain, today’s gap between models becomes the performance gap of the next models, and that gap keeps widening, so latecomers find it harder and harder to catch up. More than whether a company has the chain at all, what decides the competition is that the gap in the results the chain produces goes on widening with every turn. This is my hypothesis restated more exactly. It is not a verified fact, and I write it as my own thinking.

I confirmed this reading most strongly in a passage of an Anthropic system card. It says that, on the grounds that recent models have the ability to accelerate their own development, they implemented an intervention that, for requests aimed at developing frontier large language models, lowers the model’s effectiveness in a way invisible to the user. In other words, Anthropic throttles the power its own model holds to build the next model whenever the request comes from someone else. The party writes this much in its own document, so I read it as support for that power sitting at the center of the competition.

On the other hand, when I read the primary sources, I found that you cannot draw a line among the leading companies between those that have the chain and those that don’t. The first to publish the most concrete instance of a closed chain was Google. AlphaEvolve is a coding agent that runs on Gemini. A function it found keeps recovering 0.7 percent of the compute in the company’s data centers. There was also an improvement that made the whole of the core computation of training 23 percent faster on average, which cut Gemini’s training time by 1 percent. Google adopted a circuit design that AlphaEvolve proposed for the next generation of TPU. On the circuit, though, the authors of the paper themselves add that an existing synthesis tool had independently found the same improvement. The paper states plainly that this is a new instance of Gemini optimizing its own training process through AlphaEvolve. Google published this about a year before the other two companies began disclosures of the same kind. And these are measured numbers from improvements deployed in production, not a self-reported productivity survey. Google’s disclosure was modest, but its example of the chain was the most concrete.

The chain I run at an individual scale

I operate a chain of the same structure at the scale of one person. Let me start with the layer of rules documents. The documents that set down the discipline of this operation come to 54, counting the one at the top. Every time I point something out, the AI adds that lesson to the rules documents on the spot. On revision, rather than stacking additions, the AI rewrites the whole body so that only the current rule stands there. Inside the documents, 248 places state their origin as a dated remark or comment of mine. That is a count of occurrences, so the same lesson gets counted more than once across documents. Even so, the default I keep for this operation is a form in which every decision traces back to when it was made and on whose judgment. The other layer is the raw record. In the body of a commit, the AI writes, on top of what changed, why it changed, which comment and which judgment it followed, and what it verified. The commits piled up over these 20 days come to 849 when counted with merges and squashed changes excluded. Before and after the day I decided to write fuller commit messages, the median length of the body moved from 223 characters to 312. It is not a controlled comparison, so I cannot claim causation, but I have been able to measure the correlation: once the rule was written, the writing changed.

The match came out most clearly on the second pass, when I read what each company currently recommends. One company’s official guide recommended keeping one record per lesson, and writing down why it mattered as well. Official documents also carried guidance to version prompts and rules with git in your own repository rather than entrusting them to an external management feature. A research conclusion that raw records should be kept alongside summaries rather than replaced by them points the same way. So does the judgment, near-identical across the three companies, that acceleration is concentrated in execution and is not reaching judgment. The methods I had settled on earlier for the convenience of running the operation matched, one after another, the current recommendations I read afterward. This match between methods is the center of what the two passes of research told me.

That said, for my own chain too, the quantity that matters remains unmeasured. How much did output rise per unit of AI capability put in? In fact, the authors of the independent economics paper write, on their own, that no one has been able to measure this quantity, not even at the leading companies. They release almost none of their internal metrics, for competitive reasons. Then if I can measure the same quantity on a small chain at an individual scale and publish it, that becomes primary data of a kind those companies do not release.

The most modest judgment came from the parties themselves

If I do not write this part, this article becomes nothing but a summary of promotion. So I write the side that weakens the hypothesis just as fully.

The strongest counterevidence was Anthropic’s own judgment. Its April 2026 system card, while granting that the growth of its capability had turned steep, goes as far as writing that the growth it could identify is confidently attributable to human research and is not due to AI assistance, and that it confirmed this by interviewing the people involved. Even with employees self-reporting a fourfold output, combining that with an estimate of its impact on progress put the overall multiple below two. The judgment does not change in the latest system card, released on 2026-07-24, which says the acceleration is concentrated in engineering execution rather than research judgment. OpenAI’s system card also states plainly that none of its three new models reaches the threshold of High, the highest risk level, in the AI self-improvement evaluation, and an independent evaluation body judged that fully automated AI research and development will not become possible. In the independent economic estimate, the condition needed for self-sustaining acceleration is a 15 percent gain in productivity per unit of capability, and the measured figure, taking the parties’ self-reports at face value, was 9 percent. It does not reach the condition. The paper closes, though, by saying that it appears to be strengthening.

And one pattern common to the three companies comes into view. Each company reports the same acceleration at its maximum in a product announcement and at its minimum in a safety evaluation. One party put out a figure of eight times as much code integrated per day compared with 2024. That same party concluded from the same internal data that overall progress was less than double, while an outside evaluation body read more than double from the same data. Unless you check that definitions and measures agree before you check that numbers agree, you can build two opposite pictures of the same company.

The more primary sources I read, the more modest the wording of the claims becomes. The paper’s authors write their own reservations, the party behind the system card itself denies the attribution, and the party that put out the usage figure notes on its own that it is not a measure of results. The most modest telling of the story that AI has started building AI was in the primary documents of the parties, the ones who should have been claiming it most strongly.

Where the hypothesis stands, neither denied nor confirmed

The economics paper carries one reservation that matters for this hypothesis. It predicts that acceleration of narrow capability can run ahead of acceleration of broad capability, and that the narrow acceleration should concentrate in domains where verifying an answer is cheap, such as software development and mathematics. If what my hypothesis points at is this narrow acceleration, it has not been denied yet. And as long as the quantity at its core is unmeasured, it has not been confirmed either.

Let me narrow down what can be said now. If a low cost of verification is what decides the effect of the chain, then adding more checks before publication is the obvious move. If acceleration concentrates in execution and does not reach judgment, then the record of judgment is what has the highest value to hand over to AI. I design my own operation on the basis of these two expectations. I will keep measuring here whether those two expectations were right, and writing down the results.

Research and primary sources I referred to

The numbers and quotations in the body are all as the documents read when I checked them on 2026-07-26. This article is an analysis that organizes what public documents state, and it does not represent the views of any of the companies named.