# Intelligence Without Permission

Tarry Singh · Synaptic · Complete research draft · 9 September 2026

40,376 narrative words · 18 sections · 38 figures in the web edition, including 10 optional 3D studies. Captions, headings and sources are excluded from the narrative count.

This reading copy contains the complete prose and source notes. Interactive figures are available in the web edition; figure markers below identify their position and purpose. The article is a local draft, not a published announcement.

## The room gets larger

There is a room that most people never enter. Sometimes it has a university crest above the door. Sometimes it sits behind a corporate badge reader, several laboratories and an annual research budget large enough to sustain a small city. Inside are people allowed to ask difficult questions for a living. Outside are people who must first persuade someone that their questions deserve the expense.

The room has produced extraordinary things. It has also confused the conditions that once made those things possible with a permanent right to decide who may attempt them. Access to expertise became evidence of expertise. The capacity to employ a thousand researchers became a reason to dismiss a company employing twelve. Institutional affiliation became an early answer to a question that should have been settled much later: is the work any good?

That arrangement is becoming harder to defend.

The change is not that everyone suddenly knows everything. It is that a growing share of difficult intellectual work can be attempted without first assembling the institution that used to supply it. A person can ask for a mathematical approach, develop a model, inspect a paper, write an experiment’s analysis code and challenge an explanation in the same afternoon. The quality varies. Some results collapse on inspection. Others survive long enough to become useful. The right to make a serious attempt is spreading faster than the organisations built around rationing it.

My premise is that PhD-level intelligence is becoming a generally available resource. That is a direction of travel and a challenge to institutional design, not a claim that every public model is already a complete researcher. A doctorate contains judgement, persistence, tacit knowledge and responsibility that no benchmark captures in a single number. Yet the phrase names something consequential: work once requiring scarce, expensive specialist attention can increasingly be supplied through a machine, repeatedly, across borders, to people without the expected credentials.

If that continues, the important competitive question changes. An enterprise can no longer answer every challenge by pointing at the brilliance it employs. A university cannot justify every barrier by pointing at the difficulty of its subject. They must show what their organisation adds when a capable outsider can rent substantial parts of the intellectual machinery.

Consider a small laboratory with access to thousands of research agents, and eventually perhaps millions. That is a scenario, not a description of a capability already distributed equally around the world. The agents do not amount to an equal number of independent human scientists. They may share blind spots, repeat one another and overwhelm the people checking their work. Nevertheless, the prospect is enough to expose a weakness in conventional comparisons. Headcount becomes an increasingly poor proxy for the number of credible approaches an organisation can explore.

ASML does not become easy to replicate because an assistant can explain optics. IBM does not lose decades of accumulated work because someone can generate a materials hypothesis. Google combines research, infrastructure and distribution on a scale few organisations can approach. Dyson must put a functioning object into someone’s hands. Each has advantages that survive cheaper reasoning. Each also contains activities whose former exclusivity deserves a new examination. A challenger need not reproduce the whole company to take a valuable piece of its future.

> **Figure 01 · The research atlas (3D)** — Schematic. See the corresponding figure in the web edition.

The scientific examples make this argument harder to dismiss, and easier to exaggerate. OpenAI has announced a proposed resolution of a formulation of the Navier–Stokes problem. Anthropic has reported a formalisation of Fermat’s Last Theorem and progress on a bound concerning zeros of the Riemann zeta function. These are different achievements. One must not be silently substituted for another. A newly proposed proof, a machine-checked version of an existing proof and an improvement within an unsolved conjecture carry different meanings and different burdens of verification.[^S01][^S03][^S04]

In biology, structure prediction, protein design and the interpretation of measurements create another set of possibilities. Here the boundary is physical. An elegant model must face an experiment. A binding result is not a treatment. A treatment that works in one setting is not automatically safe or useful in another. DeepMind’s achievements and the wider field’s progress are remarkable precisely because the distance between a digital proposal and a biological result is real.[^S05][^S07]

An institution seeking comfort can dwell on those limitations indefinitely. A promoter seeking attention can omit them. Neither response is adequate. The limitations tell us where the next sources of power will accumulate. When proposals become cheaper, experiments become more valuable. When plausible explanations multiply, independent checking becomes more valuable. When knowledge travels further, the ability to turn it into a dependable product becomes more valuable. The institution’s task changes before the institution disappears.

That is the challenge I want to put to academia and enterprise. What exactly are you protecting? The conditions for good work, or the position you inherited when fewer people could do it? A laboratory protecting the integrity of its measurements serves a different purpose from a department protecting a familiar hierarchy. A company investing in qualification and service creates a different advantage from a company assuming its engineers’ vocabulary will keep competitors outside. Both may call their advantage expertise. The coming pressure will force the distinction into view.

The pressure will reach industries that rarely appear in glamorous accounts of technological change. Oil and gas has used sophisticated computation for decades; its exposure comes from changes in who can interpret, model and improve operations, not from a sudden first encounter with software. Utilities, manufacturers, drug developers and technical services businesses will confront their own versions. The outcomes can include safer operations, better products, more effective research and lower costs. They can also include concentrated control over the infrastructure on which everyone else depends.

There is no honest timetable for the collapse of an entire industry. There is a serious case for re-examining the timetable of every institution that assumes it has years to decide whether this matters. Scientific announcements do not instantly become commercial displacement. They change expectations, attract capital, alter what smaller teams attempt and expose which parts of an incumbent’s authority can be tested. That sequence is already sufficient to demand action.

The question running through this essay is therefore not whether the old room deserves to be burned down. It is whether its occupants can recognise what happens when the walls cease to determine the reach of the minds outside. The room can become a workshop, a public instrument, a place where more people’s ideas encounter better tests. Or it can spend its remaining authority explaining why the people approaching its door are not qualified to knock.

The future will not be decided by the number of impressive answers a machine can produce. It will be decided by who gets to turn an answer into a credible attempt, who can test that attempt, and who receives the benefit when it works.

## 01 · What becomes available

The phrase “PhD-level intelligence” is powerful because it compresses a social fact into three words. A doctorate signals permission to enter a particular class of conversation. It says that a person has spent years acquiring the background, methods and judgement needed to work near the edge of what is known. When comparable pieces of intellectual performance appear in a widely accessible system, the disruption reaches beyond productivity. It touches the way societies distribute credibility.

The phrase is also dangerous if it becomes a substitute for describing the work. A strong answer to a graduate examination question does not establish the ability to choose a research programme. Writing correct analysis code does not establish that the underlying experiment was well designed. Finding a plausible interpretation in a paper does not establish that the interpretation survives an unfamiliar dataset. The essay’s argument becomes stronger when these distinctions are visible, because institutions cannot escape scrutiny merely by identifying one task machines still perform badly.

A university department does not consist entirely of original discovery. It also reads, teaches, organises evidence, implements methods, reviews calculations and trains people to notice mistakes. An industrial research group spends time documenting designs, translating requirements, comparing alternatives and preparing experiments. Some of this work is difficult; some is repetitive; much combines the two. The economic effect of AI does not require every activity to become autonomous at once. It begins when enough expensive activities become more accessible to change who can assemble a credible programme.

### Capability has an address

We should begin with a distinction that headlines routinely erase. A company’s most advanced research system, the artefact it releases and the public product available to a customer are three different things. A published result can demonstrate that a method is possible while leaving the method inaccessible to nearly everyone. A repository can make a proof inspectable without making the search process that found it affordable. A consumer subscription can offer valuable research assistance without reproducing the capabilities of an internal system operating with extensive tools and compute.

The recent mathematical announcements require this care. OpenAI describes an internal research effort. Anthropic’s accounts of specialised mathematical work describe configurations whose conditions should not be silently transferred to an ordinary public chat session. The existence of a result and the distribution of the capacity to reproduce it are related questions, but one does not answer the other. A reader should be able to inspect both without having to disentangle them from a triumphant sentence.[^S01][^S03][^S04]

This matters to the central premise. Generally available intelligence is not the same as intelligence that has been demonstrated somewhere. The strongest version of the argument asks what must happen between those states. The system must become usable outside the developer’s own research group. Its cost must fall within a practical budget. Its tools must be available. The user must be allowed to access it from the place where they live and work. The outputs must be checkable by someone other than the provider. These are substantive conditions, not footnotes to a story of inevitable democratisation.

> **Figure 02 · Capability is not access (2D)** — Dated comparison. See the corresponding figure in the web edition.

A model may cross the first threshold and fail the others. An excellent assistant can be too expensive for sustained use. A low-priced service can be unavailable in a country where a researcher lives. An apparently open workflow can rely on a database, instrument or licence that restores the old barrier in another form. Even a capable and affordable system may be difficult to use well when a person lacks the background to recognise its mistakes. Access has several layers, and the distance between them is where many future inequalities will live.

This does not diminish the importance of wider availability. It tells us how to measure it. Count completed, checked tasks at an affordable total cost. Ask which users can carry them out without institutional sponsorship. Ask what happens when a subscription ends, a model changes or a laboratory refuses access. A market full of impressive demonstrations can coexist with a narrow distribution of effective research capacity. A serious account of disruption must be able to recognise that outcome rather than defining it away.

### A doctorate is a bundle of activities

Imagine taking an ordinary research week apart. There is the work of understanding what has already been done. There is the work of turning a question into a tractable problem. There is implementation: code, calculations, experimental plans, proofs or designs. There is interpretation of results, including the unpleasant possibility that the original question was poorly chosen. There is communication with colleagues whose expertise differs from one’s own. Finally, there is judgement about what deserves another week, another experiment or another year.

These activities reinforce one another, but they are not interchangeable. A system can be excellent at suggesting code and weak at identifying that the objective function measures the wrong thing. It can propose a surprising molecular structure and fail to appreciate a practical constraint in the laboratory. It can locate an argument’s missing step without knowing whether the larger theorem is worth pursuing. The user’s opportunity lies partly in assigning the right activity to the right tool. The institution’s opportunity lies in rebuilding its workflow around those differences.

> **Figure 03 · What PhD-level work contains (2D)** — Evidence matrix. See the corresponding figure in the web edition.

The phrase PhD-level therefore works best as an invitation to specify a task, a standard and a setting. Which discipline? Which activity? What evidence of reliability? How much assistance? What happens when the inputs are ambiguous or incomplete? These questions do not turn a compelling thesis into a timid one. They prevent the thesis from depending on a claim so broad that any failure can refute it. The competitive pressure arises from particular capabilities becoming widely available, then combining in ways the old organisation was not designed to exploit.

A small company may need excellent computational chemistry for a narrow problem without needing a complete autonomous chemist. A researcher may need someone to explore ten plausible mathematical routes and report why nine failed. A manufacturer may need a reliable way to compare designs before buying another prototype. If those tasks become cheaper and easier to obtain, the user gains something concrete. The industry can change even while the most difficult judgement remains concentrated in experienced people.

There is another distinction inside the word intelligence: performance on a selected attempt versus dependable performance across repeated attempts. A demonstration usually shows what worked. An organisation must budget for what failed, what needed repair and what went undetected. The relevant unit is not the impressive answer. It is the useful result after checking, retries, integration and consequences. That unit may still become dramatically cheaper. We will only know by measuring it properly.

### Distribution changes the kind of competition

When expertise is available mainly through employment, a challenger must hire before it can try. Recruitment requires money, reputation and time. A research proposal must survive several filters before the organisation acquires the ability to explore it. When some expertise can be obtained as a service, the sequence changes. The team can investigate enough to demonstrate promise before it raises the resources needed for a larger commitment. This is one reason the effect may exceed the savings on a particular task.

The change also alters who can combine fields. A capable specialist often knows that another discipline contains something useful but lacks the time or access to investigate it. An assistant can make the first translation, identify concepts and expose where the analogy breaks. The user still needs verification. Yet the cost of entering an unfamiliar literature falls, and with it the cost of discovering that two apparently separate problems share a useful structure. Institutions built around departmental boundaries may find this particularly uncomfortable.

No system eliminates the need to learn. In some cases it increases it. A person who can generate complex work quickly needs enough understanding to supervise a larger field of possibilities. The difference is that learning can become more closely coupled to doing. A question can lead to an explanation, a small implementation, a failed test and a revised understanding in one continuous process. The researcher’s limited attention becomes a resource to protect, rather than a resource consumed by every routine intermediate step.

This is why the worldwide premise matters. The next valuable question need not originate where the largest research budgets are located. It may come from an engineer who understands a local operational problem, a clinician who sees a neglected pattern, a teacher with an unusual mathematical interest or a small business confronting an inefficient process. Their advantage may be proximity to the problem. Wider access to intellectual tools can help them turn that proximity into an investigation someone else can assess.

The argument does not depend on romanticising outsiders. Outsiders can be wrong, careless or overconfident. Established researchers can be generous, inventive and faster to adapt than their challengers. The point is to make the quality of the attempt more decisive than the category of the person making it. A system that lowers the cost of entry also needs stronger ways to reject weak work without restoring affiliation as a convenient substitute for evaluation.

### Availability is a choice as well as a trend

Whether this capacity becomes broadly useful depends on decisions about pricing, access, interoperability, education and physical facilities. A science programme that provides model credits addresses one constraint. It does not automatically provide experimental capacity or long-term independence. An open dataset can lower an entry barrier, but only if its documentation permits meaningful use. A university can welcome external contributors while leaving its practical procedures so opaque that few can participate. The language of openness is easier to supply than the conditions of participation.[^S09]

Enterprises face a parallel choice. They can treat advanced assistants as a cheaper way to preserve the same hierarchy, passing every consequential decision through the same queues. Or they can ask which decisions were centralised because information was expensive and which still require central responsibility. The first approach may save time locally. The second can change the organisation’s speed. Neither follows automatically from purchasing software, and neither can be assessed by counting licences.

The useful test is simple to state and demanding to meet. Can more people attempt valuable work, obtain a fair evaluation and carry a successful result further than before? That test connects capability to access and access to institutional change. It also creates a disciplined way to examine the scientific stories ahead. We are looking for evidence that the frontier of plausible attempts is moving, while keeping the distance between an attempt, a checked result and a benefit in view.

PhD-level intelligence becoming available to anyone is therefore both a technological possibility and an institutional demand. The technology can make more of the work possible. Institutions decide whether the people using it encounter a test, a partner, a customer or another locked door.

## 02 · The equation that changed the conversation

Water is familiar enough to make the mathematical problem sound almost embarrassing. We pour it, pump it, heat it and force it through machinery. Engineers calculate its behaviour well enough to build bridges, turbines and aircraft. Yet a set of equations can be useful in practice while leaving a profound question about its own solutions unsettled. Knowing how to obtain a good approximation in a particular setting is different from proving what every admissible solution can do.

That difference sits at the centre of Navier–Stokes. The equations describe how a fluid’s velocity changes under motion, pressure, viscosity and external force. In the incompressible setting, they also impose a condition that prevents the velocity field from creating or destroying volume locally. The difficult question concerns the relationship between orderly starting conditions and later behaviour. Can smooth data evolve into a breakdown of smoothness? Which assumptions make continued regularity possible, and which leave room for a singularity?

The word singularity invites theatrical pictures: a vortex collapsing, a point flashing, a simulation exploding. Those pictures can be useful for intuition, but they are especially prone to misleading the viewer. A numerical calculation can fail because the algorithm is unstable, because the grid is too coarse or because the quantities exceed the computer’s representation. None of those events, by itself, proves that a solution of the mathematical equations becomes singular. The distinction between a computation failing and a theorem establishing failure is the first thing a responsible visual explanation must preserve.

### Read the claim before the headline

On 8 September 2026, OpenAI announced a proposed solution addressing the breakdown alternatives in the Clay formulation. The accompanying manuscript describes smooth forcing and zero initial velocity, with finite-time growth of the velocity becoming unbounded while kinetic energy remains bounded. It presents results for positive viscosity and claims the relevant whole-space and periodic conclusions. These qualifications define the achievement being claimed; removing them would create a different theorem.[^S19]

The company also released Lean formalisation material. That makes the announcement more substantial than a persuasive essay asserting that an old problem has fallen. It provides an artefact for examination. It does not entitle this article to claim that I have independently verified the entire proof, or that the mathematical community has already completed its scrutiny. The evidence state at the research cut-off is a recent announcement with a manuscript and formal material, not an awarded Millennium Prize.[^S20]

The official problem statement matters because it gives several routes to a resolution. Its existence alternatives are framed for unforced flow; its breakdown alternatives allow appropriately smooth forcing. A reader who assumes every version asks exactly the same question can misunderstand both the result and objections to it. The right response is to put the assumptions beside the conclusion and examine their relationship, not to rely on the familiar name of the problem.[^S22]

There is a broader lesson here for scientific AI. Famous problems arrive in public conversation as names. Research proceeds through statements. Between the name and the statement lie boundary conditions, quantifiers, domains, definitions and exceptions. A system that produces an impressive result within one statement may have done something extraordinary. It has not thereby established every implication that the public associates with the name. Precision is what allows the extraordinary part to survive contact with scrutiny.

### Why bounded energy does not settle everything

For a nonspecialist, the coexistence of bounded energy and unbounded speed can sound contradictory. It is not. An average or an integral can remain controlled while a quantity becomes very large in a progressively smaller region. Consider a simple family of bumps on a line. Make each bump twice as high and sufficiently narrower. The maximum rises, yet the area under a suitable power of the bump need not rise with it. This example is not a fluid solution. It explains why one kind of bound does not automatically supply another.

The corresponding distinction in fluid analysis is between control over the field as a whole and control at its most extreme locations. An energy estimate is powerful information. It may still leave open whether motion concentrates at smaller and smaller scales. A proof must work with the actual equations and their interactions, not merely draw a tall narrow curve. That is where the difficulty resides: showing that the proposed behaviour is compatible with all the constraints simultaneously.

Viscosity tends to smooth variations in velocity. Nonlinear motion can move and intensify structure. Pressure participates in enforcing incompressibility across the field. These processes cannot simply be assigned independent scores and added together. The field changes the conditions of its own evolution. A promising local picture may fail because it creates an unacceptable residual elsewhere, violates the required forcing conditions or cannot be extended across the whole domain. A convincing mechanism is the beginning of the mathematical work, not its completion.

> **Figure 04 · A vortex under scrutiny (3D)** — Mathematical illustration. See the corresponding figure in the web edition.

The figure offers a deliberately finite illustration. It is a way to see rotation and a changing spatial scale, with equations and assumptions exposed in its method note. It does not animate infinity. It does not present the actual proof as a picture. The aim is to help the reader understand why geometry and concentration matter, while leaving the truth of the theorem to the argument and its verification. Visual beauty becomes intellectually useful when it makes the boundary of its own claim visible.

This is also a useful discipline for business leaders. A simulation that looks physically convincing may contain a poor model. A dashboard that moves smoothly may conceal unreliable data. A model’s persuasive presentation can increase confidence faster than it increases evidence. If AI makes it cheap to generate polished analyses, organisations will need a stronger habit of asking what quantity is represented, what assumptions generated it and what independent test would expose an error.

### The research process becomes part of the result

OpenAI’s account describes extensive parallel research and formalisation rather than a single prompt producing a finished proof. The reported scale includes thousands of concurrent agents and a very large volume of generated work. That is evidence about a research process under particular resources. It should not be converted into a claim that an ordinary user can already reproduce the achievement, or that each agent performed like an independent mathematician.[^S01]

The more interesting organisational question is how such a process allocates exploration and rejection. Research contains many routes that do not work. Some fail immediately; others fail after substantial investment. A useful system must identify which partial results can be reused, which ideas duplicate existing attempts and which failures reveal a constraint worth preserving. Generating a larger pile of arguments without organising that feedback would merely move the bottleneck to the people reading them.

Formal tools can change this arrangement because they provide a demanding form of feedback. A candidate step can be rejected without a senior mathematician personally reviewing every line. However, the system still needs an appropriate formal statement, a trustworthy checking process and a way to manage dependencies. The organisation’s skill shifts towards the design of the search and verification environment. It becomes less meaningful to ask whether the system was “just predicting text” than to ask what the complete process could reliably produce.

That last point deserves care. Describing a mechanism does not settle the value of its output. Human researchers also use mechanisms: memory, analogy, calculation, conversation, trial and error. A discovery’s significance depends on what has been established and how it can be checked, not on whether its origin conforms to a preferred image of creativity. At the same time, a useful output does not justify every expansive claim about the system’s understanding. We can recognise an achievement without pretending the philosophical questions have disappeared.

The historical comparison should also avoid an easy injustice. A problem’s long life is not a measure of how long researchers have been doing nothing. Difficult fields accumulate estimates, constructions, counterexamples and methods that narrow the possibilities. A later proof can depend on that work while appearing sudden from outside. The elapsed centuries belong partly to the creation of the language in which a new attempt can succeed. AI’s use of that inheritance is a reason to examine attribution carefully, not a reason to deny the possibility of a new contribution.

### What verification has to accomplish

There are at least three questions a reader should keep separate. Does the formal object establish the formal statement? Does that statement correspond to the mathematical claim people believe is being made? Does the claim satisfy the conditions attached to the famous problem? These questions overlap, but answering one does not make the others disappear. A flawless derivation of a weaker statement would still require a correction to an overbroad headline.

The first question directs attention to the proof system and its trusted foundations. The second directs attention to definitions and translation. The third directs attention to the official formulation and its interpretation. Experts contribute at each stage, and a transparent release lets more of them contribute. An institution that supplies careful independent checking is doing valuable work. An institution that insists its affiliation must precede any consideration of the result is defending a different proposition.

The Clay prize procedure adds another layer. Its rules require qualifying publication, a period of at least two years and general acceptance in the mathematical community before a prize can be considered. That process is not the same thing as the truth of a theorem, but it makes clear why a fresh announcement is not an award. A dated account should preserve the distinction and be revised when the evidence changes.[^S02]

> **Figure 05 · What has actually been checked (2D)** — Evidence timeline. See the corresponding figure in the web edition.

The timeline is intentionally incomplete where the public record is incomplete. It does not fill the gap between release and acceptance with an optimistic arrow labelled inevitable. A reader can see what exists, what has been reported and what has not been established in this article. That is how a strong challenge to institutions protects itself from becoming a promotional document for the companies challenging them.

There is an opportunity for universities here that is easy to overlook. They can organise independent verification as a visible public service. They can provide explanations that identify precisely what a result changes. They can maintain durable records of corrections, competing interpretations and downstream consequences. If AI increases the supply of candidate breakthroughs, the demand for this work increases. The institution’s authority can become more defensible when it rests on the quality of its checking rather than the scarcity of its access.

### A theorem is not a better pump

The temptation to jump from Navier–Stokes to industrial revolution is understandable. Fluid behaviour matters in aircraft, energy systems, chemical processes and countless machines. Yet a foundational theorem does not automatically provide a faster computational method, better measurements or a design that can be manufactured. Its commercial consequences may be indirect, delayed or concentrated in areas that are not obvious at announcement. We should resist promising an operational benefit merely because the equations have industrial applications.

An engineer solving a particular flow problem works with geometry, boundary conditions, approximations, available compute and an acceptable error. A theorem about existence or breakdown addresses a different layer. It may influence understanding without immediately changing that engineer’s tool. Conversely, an improved approximation or optimisation method can create commercial value without resolving the foundational question. Keeping these layers apart allows the article to examine actual routes to disruption rather than relying on a famous equation as a universal symbol.

One plausible route is organisational. If a research process can sustain difficult exploration, retain useful partial results and combine them with formal checking, similar methods may help with other well-specified problems. The relevant transfer is not that every industry contains a Navier–Stokes theorem waiting to fall. It is that parts of research previously limited by the availability of specialist attention may support a different scale of search. The transfer has to be demonstrated in each setting, with the appropriate evaluator.

Another route is educational. A difficult result accompanied by inspectable artefacts can become material that people outside a narrow circle study, question and extend. AI assistance may help them traverse the prerequisites. This will not make advanced analysis effortless. It can make the first serious steps less dependent on admission to a particular institution. The institution then faces a choice between supporting that enlargement and treating it as a threat to the order in which learning is supposed to occur.

A third route is commercial expectation. Investors, founders and research leaders update their beliefs about what smaller teams might attempt. They may allocate resources differently long before a theorem changes a product. Some of those bets will fail. The important point is that competitive pressure can begin with changes in the range of credible attempts, rather than waiting for complete automation of a profession. An incumbent that waits for its entire business to be demonstrably replaceable will have chosen an unnecessarily late moment to respond.

The Navier–Stokes announcement therefore belongs in this essay as a serious, recent claim requiring close reading. Its value to the argument does not come from pretending scrutiny is finished or that fluid engineering has been solved. It comes from the research process it exposes, the artefacts it makes available and the questions it forces institutions to answer about who can now attempt work of this difficulty.

The old defence was that only a small number of places could assemble the necessary intellectual capacity. The emerging defence must be more specific. An institution may possess unusual judgement, physical resources, a valuable research environment or an exceptional ability to verify and explain. Those are substantial advantages. “We are the people who do difficult mathematics” is becoming a less complete answer.

## 03 · The proof becomes a file

Fermat’s Last Theorem is an unusually good test of whether an account of AI progress respects the difference between a celebrated name and a specific achievement. The theorem was not waiting for a language model to supply its first proof. Andrew Wiles’s work, completed with the crucial Taylor–Wiles contribution, was published in 1995. Any account that erases that history in order to make a contemporary announcement sound larger has already weakened its argument.[^S23]

The theorem concerns the impossibility of finding positive whole numbers whose powers satisfy the familiar equation once the exponent exceeds two. Its statement is accessible; its proof draws on deep mathematical machinery. That distance between a simple claim and a complex justification has made it a symbol of mathematical perseverance. In the AI era, it becomes a symbol of something else as well: the possibility of converting a vast structure of reasoning into an artefact that a machine can check.

Anthropic’s September announcement concerns formalisation. The company reports producing a Lean proof over eleven days, building on an extensive human mathematical and open-source foundation. The significance is the translation and completion of a large, interconnected body of argument within a formal system. Calling this the first proof would be false. Calling it mere clerical transcription would miss why the task is difficult and why its automation matters.[^S03]

### Two achievements, two histories

A human mathematical proof is written for readers who bring substantial knowledge to the page. It can omit routine steps, invoke a standard result and rely on shared conventions about notation and definitions. That is a strength of mathematical communication: it makes complicated ideas comprehensible without expanding every consequence to the foundations. It is also one reason translating a proof into a formal system can be much more demanding than retyping its sentences.

A formal development must make the relevant dependencies available in a compatible form. A theorem from one library may use definitions that do not immediately match those in another. An argument described as standard may require considerable construction before the checker can accept it. The formaliser must decide how to represent objects and how to organise reusable results. These choices influence the difficulty of the work and the usefulness of the resulting library.

> **Figure 06 · Fermat, discovery and formalisation (2D)** — Historical timeline. See the corresponding figure in the web edition.

There is no need to diminish one achievement to recognise the other. Discovery expands what is known. Formalisation can strengthen how a result is checked, maintained and reused. Exposition helps people understand why it works. These activities can overlap, and one can stimulate another. A culture that values only the first announcement of a theorem may underinvest in the infrastructure that makes the theorem dependable and usable by a wider community.

That underinvestment has an institutional explanation. Academic rewards often attach more visibly to recognisable intellectual ownership than to maintenance or translation. A long formalisation project may be essential while offering uncertain credit to the people doing it. AI changes the economics of some of that work, but it does not remove the need to decide which infrastructure deserves support. If the newly affordable work becomes abundant, institutions should reconsider the allocation of human effort rather than assuming the old prestige order remains appropriate.

### What the checker checks

Lean’s central contribution is a small kernel that checks proof terms against the system’s rules. Elaborate search and automation can operate around that kernel, but their output must ultimately satisfy it. The separation matters: a creative or unreliable search process can still produce an acceptable proof object. Reliability need not depend on trusting every intermediate suggestion made during the search.[^S24]

The exact statement remains indispensable. A theorem’s name is not its meaning. A file labelled with a famous result could prove a related statement, assume a key conclusion or use definitions that differ from what a reader expects. Proper validation therefore inspects assumptions and checks that the formal statement corresponds to the intended claim. Lean’s own guidance explicitly distinguishes whether a theorem has a valid proof from what the statement means.[^S25]

The released Fermat repository makes this boundary unusually visible. Its final check prints the theorem’s axioms and derives Mathlib’s formulation. The repository describes additional verification through Comparator and an independent kernel implementation. Those are reported checks that readers can inspect; I have not rerun the full resource-intensive build for this essay. The artefact permits stronger scrutiny than a screenshot of a success indicator, while still requiring an honest account of who performed which verification.[^S21][^S26]

> **Figure 07 · Inside a checked proof (2D)** — Dependency diagram. See the corresponding figure in the web edition.

The dependency view is deliberately small. It follows a real portion of the released proof path instead of drawing thousands of arbitrary glowing nodes. Its purpose is to show how a conclusion rests on named intermediate statements and how those statements connect to the final contradiction. The complete development is much larger. A useful visual should make the reader more capable of asking a question about that development, not merely more impressed by its size.

### A new form of intellectual portability

A checked proof can travel differently from a claim supported mainly by reputation. The recipient does not need to know the author personally to inspect the statement and run a suitable checker. This does not eliminate expertise. Understanding the result, selecting useful extensions and interpreting definitions remain demanding. It changes which parts of trust need to be supplied socially and which can be supported by an executable artefact.

That shift is central to the article’s institutional argument. A researcher outside a prestigious department may struggle to obtain attention for a complicated claim. If the work arrives with a well-specified statement and a reproducible checking path, affiliation becomes a less defensible reason to dismiss it. The institution can still reject work that is irrelevant, badly explained or uninteresting. It has less reason to use an author’s address as a substitute for evaluating correctness.

There is a corresponding burden on the outsider. Formal validity does not guarantee significance. One can generate many true statements that nobody needs. One can prove a property of a model whose connection to the real problem is weak. The ability to obtain a checkmark creates a new temptation to mistake measurable completion for intellectual importance. Universities and research communities remain valuable when they help distinguish a technically valid contribution from a consequential one.

For enterprise, the parallel is specification. A company can ask whether a piece of software satisfies a precise property, but first it must decide whether the property captures the behaviour customers require. Formal methods can provide strong guarantees within their scope. They cannot repair a commercial requirement that omits the condition under which a system will be used. The difficult organisational work often lies in translating messy expectations into a statement worth proving.

AI assistance may lower the cost of that translation and of the proof itself. It may help explore edge cases, connect existing lemmas or propose a more tractable specification. The benefit is potentially substantial. It should be assessed through defects prevented, requirements clarified and changes checked, rather than through the number of generated proof lines. More formal text is not automatically a better assurance system, any more than more source code is automatically a better product.

### Credit is part of the infrastructure

The Fermat development rests on a long chain of mathematical contributions and software. The released attribution material acknowledges upstream projects, including the Imperial College London effort led by Kevin Buzzard, other formalisation work and Mathlib. That inheritance is not decorative background. It is part of the explanation for what the contemporary system could do. A narrative in which a company’s model appears alone at the end of history would misunderstand its own evidence.[^S27]

The same point applies to the critique of academia. Universities have built much of the knowledge, trained many of the researchers and supported tools that now make broader participation possible. Their contribution does not grant them a permanent exclusive right to control access. It does mean that a serious argument for change should ask how those public goods will be maintained. A system that consumes a common foundation while starving its upkeep will eventually damage the conditions of its own success.

Attribution also affects incentives. If a later automated project receives all the attention while the people who created the definitions, libraries and underlying ideas become invisible, the market for prestige will send a poor signal. We need ways to credit reusable foundations and careful verification alongside dramatic completion. Institutions can help build those practices. They can also fail to do so, leaving corporate announcements to define the history by default.

The constructive opportunity is a research culture in which more work becomes inspectable, reusable and open to contribution. AI can assist with the labour of formalisation; institutions can supply durable stewardship, education and independent checking. Smaller groups can make contributions whose correctness is less dependent on the standing of their organisation. The resulting competition is demanding because it asks everyone to be more explicit about what they add.

Fermat’s theorem was already proved. The newer story is that a formidable structure of reasoning can be turned into a portable, checkable object through a different organisation of work. That is enough to matter. It changes the economics of verification, the reach of mathematical infrastructure and the terms on which an outsider can ask to be taken seriously. A strong argument needs no historical theft to make that consequence visible.

## 04 · The line through the primes

The prime numbers refuse to settle into a simple repeating pattern. Two, three, five, seven, eleven: the sequence begins innocently and quickly becomes a lesson in the difference between order and regular spacing. Primes are the multiplicative building blocks of whole numbers. Their distribution connects elementary arithmetic to some of the deepest structures in mathematics. The Riemann zeta function is one of the bridges between those worlds.

To draw that bridge, we must first accept that a function can take a complex number as its input. A complex number has a real component and an imaginary component, so its inputs occupy a plane rather than a line. The output is also complex. A complete picture therefore needs more dimensions than an ordinary graph supplies. Any visualisation must choose what to show: magnitude, phase, a slice, contours or some combination. The choice can reveal a structure while hiding another.

The landscape in this chapter plots the magnitude of the zeta function over a finite part of that input plane. Places where the magnitude reaches zero correspond to zeros of the function. The critical line runs through the real coordinate one half. The Riemann hypothesis says that every nontrivial zero lies on that line. The word every is the difficulty. A beautiful picture of several zeros, or a computation of an enormous number of them, cannot by itself settle a statement extending without limit.[^S28]

### What the picture can teach

The first useful lesson is that a zero is a location, not a little object placed on a chart. It is an input at which the function’s output vanishes. When the landscape approaches that location, its height falls. A coarse grid can miss the exact point, leaving a shallow depression where a precise calculation would reach zero. Interpolation makes the surface look continuous, but does not create additional evidence. A rendered valley is a guide to the function, not a certificate of its exact behaviour.

The second lesson is that the axes carry different meanings. Moving horizontally changes the real part of the input. Moving along the other direction changes the imaginary part. Height represents a transformation of the output’s magnitude. It does not represent probability, the density of primes or how much of the conjecture has been solved. Without that explanation, a sophisticated visual can teach a false intuition more effectively than a simple paragraph ever could.

> **Figure 08 · The zeta landscape (3D)** — Computed mathematics. See the corresponding figure in the web edition.

The third lesson concerns scale. We sample a bounded region and state its limits. The pole at one is outside the displayed domain. The height transformation prevents large values from flattening the interesting detail near smaller ones. These are editorial choices about visibility. They are acceptable because the method is disclosed and the reader can see what was computed. They would become misleading if the figure were presented as a faithful picture of the entire mathematical problem.

This is a small example of a larger responsibility. When AI makes computation and visual production easier, it also makes it easier to publish an authoritative-looking object whose relationship to the underlying claim is obscure. A serious scientific article should help readers inspect that relationship. The sophistication lies in exposing what matters, not in making the mechanism disappear behind an impressive surface.

### A stronger result inside an unsolved problem

Anthropic’s August 2026 announcement does not claim a proof of the Riemann hypothesis. It reports an improvement in a lower bound for the proportion of relevant zeros on the critical line, from approximately 41.6 percent to 67.2 percent. The company describes an unreleased research version of Claude, subsequent mathematical assessment and a formalisation. The full conjecture remains beyond the reported result.[^S04]

The distinction between a lower bound and a measured fraction is essential. Suppose a theorem establishes that at least a certain proportion has a property in the relevant limiting sense. It does not establish that the rest lack the property. The actual proportion could be higher. In this case, a stronger guarantee about zeros on the line does not imply that the remaining zeros have been located off it. Nor does it divide the conjecture into a completed portion and a remaining portion like a project-management chart.

The expert note states a constant a little above 67.25 percent, with a limiting error term, for zeros in a growing interval of imaginary height. Its statement also distinguishes the counting conventions in the numerator and denominator. For a general audience, the rounded headline is sufficient to identify the reported advance. For a mathematical reader, the precise statement is available. These two levels of explanation can coexist without pretending the rounded number is the whole theorem.[^S29]

> **Figure 09 · A stronger bound (2D)** — Reported mathematical result. See the corresponding figure in the web edition.

The change from 41.6 to 67.2 is 25.6 percentage points. That arithmetic is straightforward. Its interpretation is not a claim that mathematics has progressed by a corresponding fraction towards the full hypothesis. A stronger bound may open useful questions while leaving the final obstruction untouched. It may rely on a method with a natural ceiling. Scientific progress is not always movement along a single road whose length is known in advance.

This is one reason partial results deserve more respect than public discussion often gives them. A result can be significant without completing the most famous problem nearby. It may connect techniques that were previously studied separately, sharpen a tool or change what later researchers can assume. The value of the work must be judged through those consequences. The prestige of the unsolved problem should attract attention to the details, not replace them.

### The inheritance inside the new step

The accompanying material places the result in a line of human research. Aryan’s work and papers by Baluyot, Goldston, Suriajaya and Turnage-Butterbaugh explore ways of working with zero correlations without assuming the hypothesis. Earlier work by Bombieri supplies another part of the mathematical background. The reported advance combines ingredients from that literature. This is a contribution within an active field, not an arrival in a landscape where everyone else had stopped thinking.[^S30][^S31][^S32]

The paper’s argument can be approached at the level of structure without reproducing its technical proof. It associates zeros with contributions to a mathematical form and uses information about that form to constrain how many zeros must lie on the line. The important methodological point is the move from direct inspection of individual zeros to an inequality controlling a population. The original paper supplies the actual definitions and estimates; the institutional question is how a system found and assembled a productive route through that material.[^S33]

A reader might call this recombination. That label is not, by itself, a criticism. Much valuable mathematics connects existing concepts in a way that yields a previously unavailable conclusion. The relevant questions are whether the conclusion is new, whether the argument is valid and whether the connection is useful. If the work merely restates something already known, novelty claims should be withdrawn. If it establishes something new from existing ingredients, the presence of those ingredients does not make the result disappear.

This is where discussions of AI creativity often become unproductive. One side treats any useful output as proof of unrestricted scientific agency. The other sets a definition of discovery so demanding that most ordinary scientific advances would fail it. We need a vocabulary that distinguishes a new representation, a new theorem, a new experimental result and a useful combination of known methods. Those are different achievements. All can matter to a research institution or an enterprise trying to solve a problem.

The competitive consequence does not wait for agreement on the most exalted definition. A company paying for a difficult analysis cares whether the analysis is correct and useful. A researcher cares whether a lemma advances the argument. A laboratory cares whether a proposed experiment resolves an uncertainty. If AI lowers the cost of producing such contributions, organisations that sold privileged access to the underlying labour face pressure even while philosophers continue debating the nature of invention.

### Failed attempts are part of the economics

The final theorem does not reveal the full cost of finding it. Research explores dead ends. A search process may produce many plausible ideas before identifying one that survives. For a small laboratory, the practical question is whether it can afford that process, including the work required to recognise failure. It is not enough to see that a company with substantial compute eventually obtained a result. The distribution of that capacity depends on the cost of the complete attempt.

There is also a distinction between diversity of output and diversity of thought. Sixty agents can produce sixty differently worded versions of the same mistaken approach. A useful research organisation needs mechanisms that make genuinely different routes available and preserve evidence about why earlier ones failed. Those mechanisms may involve different tools, different decompositions of the problem, independent checks and carefully controlled access to shared conclusions. Merely increasing the number of conversations does not guarantee a broader search.

The human analogy is instructive. A department full of talented people can still share a blind spot because of common training, incentives or assumptions about what counts as a promising question. A machine-based organisation can inherit a comparable pattern at far greater speed. The challenge is not unique to AI, but scale can amplify it. The future small lab will need to be designed for disagreement that produces evidence, rather than for a theatrical chorus of agents assigning one another impressive roles.

In mathematics, a failed route may still yield a useful estimate or a clearer statement of an obstruction. In enterprise research, a failed design may identify a material limit or expose a requirement that was missing. An organisation that records only successes loses part of the value of its research spending. AI’s ability to generate more attempts makes the quality of that record more important. The advantage may belong to the team that learns from failure without forcing every new worker to repeat it.

### Why this result matters outside number theory

There is no responsible direct chain from a stronger zero bound to an immediate claim about the price of oil, the cost of a medicine or the future revenue of a software company. The significance for this essay lies elsewhere. The result suggests a research process capable of navigating a difficult literature, selecting mathematical tools and producing a claim that can be examined through formal and expert methods. Those are activities institutions have traditionally supplied through highly selective human organisations.

If comparable processes become accessible more widely, the effective boundary of the research community can expand. A person with a well-chosen question and modest organisational resources may be able to investigate further before needing a specialist collaborator. A small team may be able to approach a difficult subproblem that previously exceeded its practical reach. These possibilities do not mean that expertise becomes irrelevant. They mean that obtaining some of its productive effects may no longer require reproducing its traditional institutional setting.

The implication for universities is particularly sharp. They should be among the first organisations to make these methods understandable, assessable and teachable. They can train researchers to formulate questions, validate formal statements and interpret computational evidence. If instead they respond mainly by defending the old sequence of credentials and permissions, they risk turning their educational mission into a restriction on the very intellectual participation they claim to support.

For enterprise research, the lesson concerns combinations. A valuable idea may sit between literatures owned by different teams, each too busy to investigate the other’s methods. An assistant that helps traverse those boundaries can create value without discovering a new law of nature. A smaller competitor may be less constrained by internal ownership of disciplines. Its advantage could come from connecting known work to a neglected commercial problem faster than the incumbent can organise a meeting about it.

This is a conditional argument, and it is strongest when the condition remains visible. The capability must work reliably enough, at an affordable total cost, on the relevant problem. The result must survive checking. The user must have access to the tools and evidence required. None of those requirements is trivial. None justifies assuming the organisational arrangements of the last century will remain the natural home of the work.

The line through the zeta landscape therefore carries two stories. One is mathematical: a precise claim about zeros, with a reported stronger bound and an unresolved conjecture. The other is institutional: a changing account of who can attempt difficult intellectual work and how that work can earn credibility. We should admire the first without exaggerating it, and confront the second without waiting for every famous problem to be solved.

## 05 · Biology opens

A protein structure can look like a piece of sculpture. Ribbons fold into loops, sheets pack against helices and coloured atoms gather into a form that appears almost designed for display. The beauty is real. So is the danger of mistaking a finished image for a finished explanation. A structure tells us something about the arrangement of a molecule. It does not, on its own, tell us everything the molecule does in a living system, how it changes over time or whether intervening in its behaviour will help a patient.

The biological story is therefore more demanding than a sequence of triumphant images. It concerns several different kinds of work: predicting structure, proposing new molecules, estimating interactions, interpreting experiments and connecting results to useful interventions. AI can improve one of these activities without solving the others. It can also connect them more effectively, reducing the time researchers spend moving between tools and making sense of their outputs. That second possibility is especially consequential for smaller laboratories.

For years, access to specialist computational work has shaped which biological questions a team could investigate. The constraint was not always the absence of an idea. It could be the difficulty of finding someone who knew the relevant software, could prepare the inputs and had enough time to interpret the result. An assistant that removes part of that obstacle changes the practical reach of a researcher. The institution that previously supplied the obstacle’s solution must then explain what else it contributes.

### From sequence to a testable picture

AlphaFold’s importance begins with structure prediction. Jumper and colleagues’ 2021 work demonstrated a major advance in predicting protein structures, evaluated against experimental references. The result changed the practical starting point for many investigations: a useful structural model could be available before a team obtained its own experimental determination. That is a change in the cost and timing of a particular kind of evidence, not a claim that experimental structural biology became unnecessary.[^S36]

AlphaFold 3 extended the scope towards joint structures involving different kinds of biomolecules, including proteins, nucleic acids, small molecules and modified residues. The published work reports improvements across several interaction-prediction settings. Its benchmarks, sampling choices and confidence measures matter because performance is a property of a defined evaluation. The model’s ability to propose a complex is not equivalent to knowing how that complex behaves in every biological setting.[^S05]

The 2024 Nobel Prize in Chemistry recognised both protein structure prediction and computational protein design, through the work of Demis Hassabis, John Jumper and David Baker. That recognition helps locate the field historically. The achievement is not the first use of computation in biology, nor the first contribution of machine learning to science. It is part of a long relationship among theory, data, algorithms and experiments that reached a new level of practical consequence.[^S06]

> **Figure 10 · A protein, with uncertainty (3D)** — Scientific coordinates. See the corresponding figure in the web edition.

The lysozyme comparison in the figure makes the distinction tangible. We take an experimental structure and an AlphaFold prediction for the matching mature protein sequence, align corresponding backbone atoms and show the resulting geometry. This is a familiar molecule, so the comparison is not a fresh test of generalisation. It is a way to inspect what agreement means. A small average distance after alignment does not imply that every residue is equally reliable or that all biologically relevant states have been represented.[^S39]

Confidence colours deserve particular attention. They are estimates produced by the prediction system. They are not the same as experimental uncertainty, and a high-confidence region does not turn the entire biological interpretation into a certainty. A reader should ask what the confidence measure evaluates, what kinds of error it can reveal and which questions remain outside its scope. The colour scale is useful when it directs attention to uncertainty; it is misleading when it becomes a decorative guarantee.

A predicted structure can nevertheless be extremely valuable. It can help a researcher formulate a more specific hypothesis, identify a region worth investigating or interpret an otherwise puzzling observation. It can make an experiment more informative by narrowing the alternatives. The gain may come from asking a better question, not from eliminating the experiment. In many settings, that is exactly what good intellectual assistance should do.

This changes the value of experimental access. If a team can prepare stronger hypotheses before entering the laboratory, each unit of instrument time can carry more information. Conversely, if proposal generation grows faster than experimental capacity, valuable ideas may accumulate in a queue. The organisation that manages that queue well can become more important. The organisation that merely owns it and makes access difficult can become a more conspicuous target for competition or public reform.

### Designing something that must exist

Protein design reverses part of the direction of the problem. Instead of asking what structure a sequence might form, a researcher wants a molecule with a desired property. Generative systems can propose candidates, and other tools can evaluate aspects of their plausibility. RFdiffusion is a prominent example of the field’s progress, combining a generative approach with experimental work on designed proteins. The experimental component is essential to understanding the claim.[^S07]

The space of possible molecules is too large to explore indiscriminately. A useful design process must choose where to search, which constraints to enforce and which candidates deserve expensive testing. This is an allocation problem as much as a generation problem. The best system is not necessarily the one that produces the largest number of attractive structures. It is the one that directs a limited experimental budget towards candidates from which the team can learn something valuable.

General reasoning models can contribute by coordinating specialist tools. That role is easy to underestimate because it sounds less dramatic than inventing a new molecular method. Yet coordination can be a substantial part of the work: preparing inputs, comparing outputs, diagnosing a failed run, preserving the rationale for a decision and translating between tools with different assumptions. A smaller team may lack enough people to perform all of those functions continuously. Making them more accessible can change the scale of the team’s programme.

The technical difficulty does not vanish into a single interface. If one tool systematically favours a particular kind of candidate, a coordinating model can amplify that bias. If several predictors share training data or assumptions, agreement among them may provide less independent evidence than it appears to. If the scoring function rewards an easily measured proxy, the search may optimise that proxy while neglecting a property that matters later. The coordinator needs a well-designed evaluation environment, not merely a large vocabulary of tools.

> **Figure 11 · The binding interface (3D)** — Scientific coordinates. See the corresponding figure in the web edition.

The binding-interface illustration uses an experimentally determined streptavidin–biotin structure. It is an educational example, not an AI-designed candidate. Showing the ligand and nearby protein atoms helps explain why geometry matters, while the caption refuses to convert geometric proximity into a measurement of affinity. That restraint is part of the visual’s scientific purpose. The figure should make the reader more precise about the claim, rather than encourage a leap from a close fit to a medical benefit.[^S40]

An enterprise whose research advantage consists partly of knowing how to operate a collection of specialist tools should pay attention. That knowledge may remain valuable, but its exclusivity can erode. The firm’s stronger advantages may lie in the quality of its experimental feedback, the relevance of its targets, the breadth of its failure records and its ability to carry a promising result into development. A smaller competitor can become dangerous by improving one of those transitions, even if it never builds the incumbent’s entire organisation.

### Keep the unsuccessful designs in the frame

Anthropic’s reported design campaign produced 354 binders from 1,320 designs, with at least one successful binder for fourteen of fifteen reported targets. The technical report describes different model configurations and experimental arms. One initially selected target was excluded from the reported target set because its experimental data were inconclusive. These distinctions are necessary to interpret the headline: success across targets and success across individual designs are different measures.[^S37]

The campaign also illustrates why a model’s name is an incomplete description of an experiment. The system had access to specialist tools and substantial compute. Different configurations used different schedules and target coverage. External laboratories produced and tested the designs. The result belongs to that complete arrangement. It should not be presented as evidence that any person using an ordinary subscription can obtain the same outcome at the cost of a few prompts.[^S08]

> **Figure 12 · From designs to measured hits (2D)** — Experimental results. See the corresponding figure in the web edition.

Why insist on the denominator? Because a research budget pays for failures as well as successes. If two systems find the same number of useful candidates but one requires far more synthesis and testing, they have different economics. If one succeeds on familiar targets and another succeeds on a difficult new class, their value may differ even when a headline rate is similar. A serious comparison has to preserve the conditions that make the number meaningful.

A failed design is also more than a wasted attempt. It can expose a weakness in a predictor, a constraint in expression or a mismatch between a computational assumption and the experimental setting. The value depends on whether the organisation records enough context to learn from it. A pile of negative outcomes without provenance is difficult to reuse. A structured record of what was attempted, under which conditions and with which measurements can improve the next round of work.

This is an advantage large organisations may already possess in uneven form. They have accumulated experiments, but the data may be scattered across projects, formats and internal boundaries. A smaller laboratory may have less data and a more coherent record. AI increases the importance of that difference because a system can only learn operationally from the information it can access and interpret. The incumbent’s database is not automatically an advantage if nobody can reconstruct what its entries mean.

The constructive response is not to announce that every experiment will be automated. It is to make the research record usable. Preserve failed candidates. Record the measurement process. Distinguish an untested design from a negative result. Keep the version of the model and the conditions under which it selected the candidate. These practices support human science already; AI makes the cost of neglecting them more visible because the next search can move faster than the organisation can recover its own history.

### The quieter disruption in analytical work

The interpretation of instrument output receives less attention than the design of a new molecule. It may prove equally important to everyday research. A laboratory does not only need novel hypotheses. It needs reliable answers about what a sample contains, whether an expected transformation occurred and whether the material is suitable for the next step. Delays in those answers can interrupt the entire sequence of work.

A reported Claude Opus 5 example concerns the processing of raw NMR and LC-MS files. The technical account describes agreement with a contract laboratory’s analysis on the selected samples, including a reported purity estimate of 96.4 percent compared with 96.33 percent. The LC-MS processing time was reported as nineteen minutes. This is a bounded case of analytical assistance, not broad validation across instruments, compounds or laboratory conditions.[^S38]

> **Figure 13 · Reading an instrument (2D)** — Measured example. See the corresponding figure in the web edition.

The commercial implication is worth examining without exaggerating it. If routine analysis becomes easier to obtain, a scientist may spend less time waiting for a specialist to interpret a file. A small laboratory may gain access to work that previously required a dedicated computational colleague. A service provider may be able to deliver results faster or handle more samples. The gain can be meaningful even when no new biological fact is discovered.

The relevant comparison includes the whole workflow. A model’s processing time is not automatically comparable with a laboratory’s elapsed reporting time, which may include queues, other work and review. Nor is a match on one result enough to establish acceptable error across many results. An enterprise should test the assistant on representative cases, including inconvenient files and ambiguous samples, and measure the work required to detect and repair mistakes. The outcome to buy is dependable analytical capacity.

There is a particular opportunity in making the intermediate work inspectable. A report that includes the extracted data, processing steps and reasons for the conclusion can be easier to review than an answer alone. The scientist can challenge a baseline correction, a peak assignment or an interpretation without restarting the entire process. In this setting, transparency is not merely a moral preference. It can reduce the practical cost of supervision and make the tool more useful.

It can also expose a new kind of dependence. If the assistant is the only component capable of reading a proprietary format, the laboratory may exchange dependence on one vendor for dependence on another. A durable improvement would include reusable data access and documented transformations. The question is whether the system leaves the institution more capable after the session ends. A service that repeatedly performs an opaque rescue may be valuable, but it has not necessarily enlarged the user’s independence.

### The patient is further downstream

The distance from a promising molecule to a medicine contains questions that cannot be answered by a structural image. Does the intervention have the intended effect in a relevant biological setting? Can it reach the necessary place? What unwanted effects occur? Can it be manufactured consistently? Does the balance of benefit and harm justify use in people? These are different evidentiary tasks, and success at an earlier one does not predetermine the later answers.

Clinical research is specifically concerned with effects in people. The FDA’s account distinguishes it from preclinical work and describes trials organised around defined questions, populations, measurements and review. A computational design result is therefore several conceptual steps away from a claim of clinical benefit. That separation is not a bureaucratic objection to innovation. It identifies the evidence that must exist before a medical claim becomes trustworthy.[^S34]

There are avoidable delays within this process, and AI may help reduce some of them. Better information management, clearer protocols, improved analysis and faster recognition of an unpromising programme can all matter. But not every period of elapsed time is administrative waste. Some outcomes must be observed. Some rare effects become visible only through larger or longer studies. The task is to distinguish delay that produces knowledge from delay produced by poor organisation.

This distinction creates a fairer challenge to the pharmaceutical industry. It is too easy for a promoter to dismiss every slow step as institutional obstruction. It is equally easy for an incumbent to place all delay under the protection of scientific caution. A serious leader should be able to identify which steps are necessary for evidence, which protect quality and which persist because nobody has redesigned the process. AI gives the organisation a reason to perform that accounting with greater urgency.

Smaller laboratories can benefit without pretending to become complete pharmaceutical companies. They may develop a useful research tool, identify a better experimental route, improve a component of a design process or produce an asset suitable for partnership. Their contribution can be narrow and still valuable. The incumbent may remain the organisation best equipped to undertake expensive development, manufacturing and distribution. The bargaining relationship changes when more credible scientific work can originate outside its walls.

For patients and public funders, the desired outcome is not simply more candidates. It is a better conversion of resources into useful knowledge and, where possible, effective interventions. More candidate generation can even become a burden if the system lacks the capacity to evaluate it. The measure of progress should include negative findings, programmes stopped earlier for good reasons, improved evidence and the accessibility of eventual products. Research abundance has to be connected to a system that can act on what it learns.

### Biology’s institutions have a choice

A university laboratory can become more open to collaborators who arrive with serious computational work. It can provide experimental judgement, facilities and training that make those contributions more valuable. An enterprise can use AI to make experienced researchers more effective while inviting smaller teams to supply ideas or tools. Public organisations can support shared datasets, standards and access to facilities. None of these responses requires surrendering scientific standards. They require making the standards more explicit and the routes to meeting them less dependent on affiliation.

The defensive response is easier to recognise. It treats the ability to run familiar tools as a permanent credential, the possession of data as proof that the data are being used well, and the existence of difficult downstream work as a reason not to reconsider upstream barriers. That response confuses a valid caution with a complete strategy. A laboratory can be right about the limits of prediction and still be wrong about the durability of its own advantage.

Biology offers the clearest version of the article’s argument because it forces intelligence to meet reality. The model can propose; the experiment can refuse. Wider access to the proposing work does not weaken the importance of the refusal. It makes the quality, speed and availability of experimental judgement more consequential. Institutions that understand this can become more useful as intellectual work spreads. Institutions that rely on scarcity alone will find that their most persuasive explanation of the future is also an explanation of why someone else should build it.

## 06 · Discovery beyond the famous problems

The famous problems are useful entrances, but they are poor boundaries for the story. Most economically important research does not end with a theorem whose name appears in a school textbook. It ends with a material that behaves better, an algorithm that wastes less compute, an experiment that distinguishes two explanations or a process that fails less often. If we look only for a machine’s next encounter with a centuries-old conjecture, we may miss the changes that reach ordinary organisations first.

These quieter advances often share a practical feature: someone can specify what a better answer would look like and construct a test that gives useful feedback. The test need not be perfect. It must be informative enough to distinguish progress from fluent improvisation. When candidate generation and evaluation can be connected, research becomes a repeated process of proposing, measuring and selecting. AI can change the scale of that process without making every stage equally fast.

### A catalogue is not a factory

Materials discovery offers a striking example. GNoME’s published work reported a very large expansion in predicted crystal structures and a substantial set predicted to be stable relative to its final reference. The paper also discussed hundreds of structures that had been independently realised experimentally. These are different categories of evidence. Millions of predicted structures do not mean millions of new materials have been manufactured, characterised and turned into useful products.[^S12]

A candidate structure can matter before it becomes a product. It may suggest a region of chemical space worth exploring, expose a pattern or help prioritise experiments. But a material’s practical value depends on more than a predicted stability criterion. It must be possible to make it in an appropriate form, understand its properties under relevant conditions and incorporate it into something useful. The distance between those tasks is where industrial research earns much of its value.

The crystal figure uses an established diamond structure from the Crystallography Open Database. This is not a newly predicted GNoME material. It lets us inspect the basic idea of a repeating unit cell using documented coordinates and symmetry. A repeated pattern of atoms can be described compactly while producing a macroscopic solid. The picture explains one kind of structural information; it does not imply a manufacturing route or a commercial opportunity.[^S43]

> **Figure 14 · A crystal before it is a product (3D)** — Scientific coordinates. See the corresponding figure in the web edition.

The distinction matters because prediction abundance can change the organisation of materials research. A laboratory may have more plausible candidates than it can synthesise. It then needs a principled way to select experiments, including experiments that reveal why predictions fail. The most valuable next sample may be the one that tests an uncertain assumption rather than the one with the highest predicted score. A system designed only to celebrate promising candidates can neglect that learning objective.

For a manufacturer, the relevant question might be narrower still. It may need a coating that survives a particular environment, a component that can be produced with existing equipment or a substitute that reduces dependence on a difficult input. The best answer could be an unglamorous improvement within a known family of materials. AI’s value lies in helping explore the actual design problem, including constraints the public research benchmark may not contain.

A small laboratory can enter through this gap. It does not need to own a complete manufacturing chain to provide a useful screening method, a validated measurement or a specialised material insight. The incumbent can become its customer, partner or competitor. The disruptive unit is often a valuable piece of a process rather than an entire company. We will return to that point because it is the most common error in arguments about whether research giants can be challenged.

### An algorithm can meet its evaluator immediately

Computational problems sometimes offer a faster route from proposal to evidence. A candidate program can be run against a defined test. Its correctness and resource use can be assessed under stated conditions. AlphaEvolve combines generated code with automated evaluation and an evolutionary process that retains promising candidates. The system’s architecture directs attention to the evaluator as much as to the language model.[^S42]

One reported result concerns multiplying two four-by-four matrices with complex entries using forty-eight scalar multiplications, improving on the forty-nine obtained by applying Strassen’s construction in this setting. For this essay, I inspected the released coefficient data and independently checked all 4,096 entries of the associated tensor identity. That verifies the published bilinear construction as represented by those coefficients. It does not establish superior wall-clock speed, numerical conditioning or performance on every hardware platform.[^S44]

> **Figure 15 · The algorithm that earns its keep (2D)** — Verified algebraic identity. See the corresponding figure in the web edition.

The difference of one multiplication can look small until we ask what is being measured. In a mathematical complexity result, reducing the count can establish a new construction. In a practical implementation, additions, memory movement, precision and hardware behaviour also matter. The same algorithm can therefore be mathematically significant without being the best routine for an engineer’s immediate workload. A good comparison keeps the metric attached to the claim.

Google’s account also describes deployment-oriented improvements, including a scheduling heuristic and optimisations within its computing infrastructure. Those cases illustrate another path to value: a modest improvement repeated across a large operating base. The article should treat the reported gains as company claims tied to particular systems, rather than a universal promise about what any agent will achieve. The mechanism is nevertheless clear enough to analyse.[^S10]

If a piece of code runs billions of times, a small improvement can justify substantial research effort. A large company has an advantage because it owns the workload, the measurements and the route to deployment. A smaller team may have an advantage because it can concentrate on one neglected optimisation. Wider access to research assistance can intensify both effects. The technology does not automatically choose between centralisation and competition; the ownership of the evaluator and the deployment path helps determine the outcome.

This suggests a practical question for every technical enterprise: which expensive routines have a trustworthy test but receive little sustained attention? The answer may be hidden in scheduling, configuration, data processing or numerical code. These problems often lack the prestige of a new product, yet they impose recurring costs. A research system that explores them systematically can create value without producing a headline that mentions a major conjecture.

### Rediscovery can be a useful test, if it is called rediscovery

Google’s AI co-scientist work offers a different kind of evidence. In one reported evaluation, researchers gave the system a question whose answer their group had already established experimentally but had not yet made public. The system proposed the relevant explanation independently of that unreleased result. The experiment tested its ability to arrive at a useful hypothesis from the available background. It did not show that the machine had replaced the original experimental programme with a short conversation.[^S13]

That distinction matters more than the dramatic comparison of elapsed times. The human research created the evidence against which the machine’s hypothesis could be recognised as correct. Without that evidence, the output would have remained a candidate explanation. Rediscovery can be a strong evaluation because it gives the test a known answer while limiting direct access to it. It should not be narrated as though the evaluation’s reference work had never been needed.

The same method can be useful inside an enterprise. A company can take a completed investigation, preserve the information available at an earlier stage and test whether an assistant would have proposed an informative next step. It can compare the system’s suggestions with what later proved useful. Such retrospective tests are imperfect: they can leak information or select unusually convenient cases. Designed carefully, they are more informative than asking employees whether a polished answer seems impressive.

A prospective test is stronger for some purposes. The organisation records the question, available information and evaluation criteria before the outcome is known. It then compares what the assistant changes about the real work. Does it reduce wasted experiments? Does it identify a failure earlier? Does it help a less experienced researcher obtain a result that survives review? Those questions connect the promise of research intelligence to the institution’s actual performance.

> **Figure 16 · Where feedback is fast (2D)** — Evidence comparison. See the corresponding figure in the web edition.

The feedback comparison makes no claim that all scientific fields can be placed on a single speed scale. Checking a formal object, running a software test and observing a biological outcome are different activities. Even within one field, a test may be cheap or prohibitively expensive. The useful distinction is whether the evaluator can keep pace with the proposed work, and whether it measures the property the researcher actually needs.

### A correction to my own earlier framing

Readers of my June essay, “Your AI Doesn’t Discover Anything. Here’s the Math That Proves It,” will recognise a tension. That article discussed Wang and Buehler’s framework for distinguishing search within a fixed representation from a verified change in the representation itself. The distinction remains valuable. The headline was too categorical if read as a universal claim that AI systems cannot contribute new scientific knowledge. A formal definition of one kind of discovery does not establish that broader prohibition.[^S41]

The framework asks us to track what the system can represent, which operations are available and how a new commitment is supported by evidence. It is useful precisely because it demands more than an impressive answer. It does not require us to deny the value of a new theorem obtained from existing concepts, an experimentally confirmed design or an algorithm with a checked property. Those achievements may occupy different categories without becoming economically or scientifically trivial.

The distinction I want to preserve is between changing the language of a problem and producing new work within that language. Both can matter. A system that cannot revise its representation may encounter a serious limit on some questions. It may still explore an existing representation more effectively than a particular organisation can. For an enterprise competing on the cost and speed of that exploration, the limit is not a refuge from competition.

There is a temptation in technology criticism to choose a definition that keeps the preferred conclusion safe. If a machine produces a valuable result, redefine the achievement as something short of real intelligence. If it fails, treat the failure as a verdict on every future system. The opposite temptation is to describe every successful result as proof that no meaningful limits remain. Both approaches make it difficult to learn from evidence because they decide the interpretation before the work is examined.

Institutions should resist that habit, including when it appears in their own defence. The relevant question is not whether the new tool deserves a title associated with human prestige. It is whether the tool changes who can perform a consequential task, how much that task costs and how its result can be checked. An institution can lose an exclusive capability before a machine satisfies its preferred philosophical definition of a scientist.

### The evaluator becomes a strategic asset

Across these examples, a pattern emerges. The ability to propose is spreading, while the ability to evaluate remains unevenly distributed. In mathematics, evaluation may rely on formal infrastructure and expert interpretation. In computing, it may rely on a representative workload and reliable measurements. In materials research, it may rely on synthesis and characterisation. In biology, it may require experiments that take time and resources. Each setting creates a different opportunity for entrants and incumbents.

A company that owns a good evaluator can use increasingly capable proposal systems to improve its work. It can also become vulnerable if the evaluator is portable and competitors can obtain it. A public institution can expand participation by making evaluation more accessible, provided it retains the standards that make the result meaningful. A small lab can build its advantage around a neglected test rather than attempting to outspend a large company on general intelligence.

This is a more concrete account of disruption than the claim that machines will become clever and everything will change. It identifies the activities, the evidence and the points at which ownership matters. It also gives organisations something useful to do now: find the important questions that can be evaluated, improve the quality of those evaluations and open a path for more capable attempts to reach them. The next breakthrough may carry a famous name. The next competitive threat may arrive as a better answer to a problem the organisation considered too ordinary to prioritise.

## 07 · The small lab with a million agents

Take a small research company seriously for a moment. It has twelve people, a limited budget and a problem that an established enterprise considers too specialised to prioritise. Its founders understand the customer’s frustration. They do not have a campus, a famous laboratory or a recruitment department capable of assembling every discipline the problem might require. Under the old comparison, those absences settle the question before the investigation begins. The larger organisation has more intelligence on its payroll, so it is presumed to have more capacity to solve the problem.

Now give the small company access to a large number of capable research agents. Some read technical literature. Some build models. Some inspect assumptions. Some write and run bounded tests. Some organise the record of failed approaches. The company still has twelve people. Its capacity to attempt intellectual work no longer scales directly with that number. This is the scenario that should disturb an incumbent whose confidence rests mainly on the size of its research staff.

The million-agent version sharpens the question, but the number should not hypnotise us. A million agents are neither a million PhDs nor a million independent minds. They are a large computational organisation whose output depends on the models, tools, task design, information flows and evaluation process. The useful comparison is not between nominal headcounts. It is between the two organisations’ ability to obtain a valuable result at an acceptable cost and carry it into use.

### Start with a result someone would pay for

Suppose the small company wants to improve a component in an industrial cooling system. This is an illustrative commercial scenario. The team has access to a defined operating envelope and a customer willing to test a qualified design. It needs to understand the literature, model alternatives, evaluate trade-offs and decide which prototypes deserve fabrication. It does not need to discover a new branch of physics. It needs to produce a dependable improvement within the customer’s constraints.

The first advantage of agents is breadth at the early stages. The team can investigate more candidate approaches before committing to one. It can compare modelling assumptions, search for relevant prior art and identify failure modes that a small human team might otherwise overlook. The value depends on the quality of the work, but the organisational change is clear: the team can obtain some of the benefits of a broader research group without permanently employing every specialist.

The second advantage is continuity. An investigation need not stop because the person who understands one tool is occupied elsewhere. An agent can prepare a calculation, document a question or organise the next test while the human team focuses on decisions that need its attention. This does not mean that the system can be left unsupervised indefinitely. It means that scarce human time can be allocated differently, provided the work arrives in a form that is economical to inspect.

The third advantage is the ability to make a better case before raising more resources. The team may be able to show a customer a coherent comparison, a reproducible model and a clear experimental proposal. That can reduce the distance between an idea and a serious commercial conversation. The funding required to begin useful work may fall even when the funding required to finish development remains substantial. Entry becomes possible at a different point in the process.

> **Figure 17 · The small lab, expanded (3D)** — Explicit scenario. See the corresponding figure in the web edition.

The scenario instrument shows why the story cannot stop at more attempts. It assumes each agent produces one attempt in a day, then separates unique candidates from duplicates, checks from unchecked work and accepted candidates from rejected ones. Change the number of agents and the review queue can expand much faster than the useful output. Change the checking capacity and another part of the system becomes decisive. The assumptions are visible because the model is an aid to reasoning, not a forecast.

### The research organisation has to be designed

A badly organised swarm can be worse than a single competent assistant. It can distribute an ambiguous task, collect incompatible outputs and then spend most of its resources reconciling them. It can reward agents for sounding confident, causing uncertainty to disappear from the record. It can allow one mistaken result to become a shared premise that contaminates every downstream attempt. Scale makes these failures more expensive; it does not cure them.

The architecture should follow the task. Independent searches are useful when different routes can be explored without constant communication. Sequential work requires careful hand-offs because a later step depends on the validity of an earlier one. A shared library can prevent duplication, but premature sharing can also make every agent follow the same attractive mistake. A reviewer should have enough context to check the work without being pressured by the prestige of the agent or the apparent consensus around it.

Kim and colleagues’ study of agent systems supports this emphasis on task and organisation. Across its bounded evaluations, collaboration helped in some settings and substantially harmed performance in others. The study does not measure million-agent scientific research. Its relevant lesson is that scaling the number of agents is not a general law of improvement; coordination and task structure interact.[^S18]

> **Figure 18 · More agents, less certainty (2D)** — Study and scenario. See the corresponding figure in the web edition.

One source of failure is correlated error. If agents share a model, training background and prompt, their mistakes may not be independent. Asking several of them the same question can produce a reassuring agreement that contains little new information. A human organisation faces a related problem when all reviewers come from the same intellectual tradition. The remedy is not merely to assign different role names. It is to create different routes to evidence and preserve the possibility of a result being rejected.

A second source is weak decomposition. A difficult question cannot always be divided into a thousand easy ones. The interfaces between subproblems may contain the difficulty. An agent can solve its assigned piece while making assumptions that conflict with the next piece. A coordinator then needs enough understanding to detect the mismatch. If that understanding remains the scarce resource, adding workers can simply create a larger integration burden around the same bottleneck.

A third source is evaluation leakage. If the search can inspect and optimise every detail of the test used to judge it, it may find ways to satisfy the test without improving the intended result. The problem is familiar in software and machine learning, but agentic research makes it more general. A capable system can exploit a poorly chosen objective very effectively. The organisation needs independent checks, representative conditions and a willingness to revise its evaluator when the search reveals its weaknesses.

The small laboratory’s most important employee may therefore be the person who knows what a convincing failure would look like. That person can define a test that stops attractive nonsense from consuming the budget. They can identify which uncertainty matters commercially and which calculation would resolve it. Agents can help with this work, but the organisation cannot outsource responsibility for the definition of success without knowing what it has delegated.

### A million agents have a balance sheet

Computational labour can be elastic without being free. Each attempt consumes model inference, tool execution, storage and sometimes specialist compute. The system also needs coordination, monitoring and recovery from failed runs. Accepted candidates may require human review or physical tests. A meaningful cost comparison includes those items and distinguishes them from the ongoing costs of the institution that hosts the work.

It is misleading to compare a model’s token price with an enterprise’s total R&D budget. The latter may include facilities, equipment, materials, prototypes, field trials and activities that the model does not perform. It is equally misleading to use those physical costs to deny savings in the intellectual work that can change. The appropriate comparison follows the same output through both workflows and accounts for what each organisation actually has to buy or provide.

The cost instrument uses deliberately hypothetical prices. It separates the cost of attempts, initial checking and physical tests. A reader can lower the price of generation and watch the total remain dominated by another stage. That is a useful result. It tells a founder where to invest next and tells an incumbent which asset may become more valuable as intelligence becomes cheaper. The point is not to produce a seductive dollar figure for a million-machine laboratory.

> **Figure 19 · The cost of a checked result (2D)** — Explicit scenario. See the corresponding figure in the web edition.

There is also a difference between average cost and cash required before learning whether the programme works. A small company may face a promising expected return but lack the money to survive the sequence of experiments. Cheaper early investigation can help it make a stronger case for funding. It does not eliminate financing risk. The enterprise may retain an advantage through its ability to fund a long programme and absorb failures that would end a smaller company.

Cloud access can reduce the need to own equipment while creating other dependencies. The team may rent compute in bursts, pay for specialised services and contract experiments. That changes fixed costs into variable costs, which can lower the barrier to entry. It can also expose the team to price changes, access restrictions and service interruptions. The future lab’s economics depend partly on whether it can switch providers and preserve its work when the underlying service changes.

An agent count is particularly unhelpful without a time unit. A million brief tasks, a million continuously active processes and a million specialised workers with large computational budgets are different arrangements. They have different costs and different coordination problems. The article’s scenario treats the count as a way to explore the organisation of work. Any real claim should specify the duration, workload, model configuration and output being counted.

### The scarce asset may be the right question

The large research institution often sees more of the field. The smaller company may see more of a neglected problem. It can be close to a customer whose requirements are too unusual to command attention inside a broad product organisation. It can notice that a supposedly minor inconvenience creates a substantial cost. Wider access to research assistance allows that local knowledge to be combined with a larger body of technical work.

This creates an important asymmetry. A challenger does not need a general superiority in intelligence. It needs a sufficiently good understanding of a valuable question and a way to obtain the work required to answer it. An incumbent may employ better scientists in aggregate and still fail to address the question promptly because its internal incentives point elsewhere. The vulnerability lies in allocation, not necessarily in the quality of the people on either side.

The question must be chosen carefully. A small team that competes on a problem requiring years of exclusive data and expensive qualification may discover that its intellectual reach exceeds its ability to act. Another team may choose an interface where a useful result can be tested and purchased quickly. The difference is strategic. Abundant reasoning expands the menu of possible attempts; it does not tell a company which attempt fits its resources.

This is why the best small laboratory may look less like a miniature university and more like a carefully chosen set of relationships. It has access to model providers, experimental partners, domain advisers and a customer who can evaluate the result. It owns the research record, the problem understanding and a route to commercial value. Its permanent staff may be small because much of the surrounding capacity can be obtained when needed.

The organisation can also remain small for the wrong reasons. If it lacks enough human understanding to supervise the work, it may become dependent on whichever output seems most persuasive. If it outsources every experiment without learning how the measurements are produced, it may fail to recognise a systematic problem. If it treats every accepted computational candidate as a discovery, it may spend its credibility faster than its cash. The strength of a small lab comes from disciplined concentration, not from the absence of people.

### What the incumbent should find uncomfortable

The uncomfortable comparison is not twelve people against ten thousand. It is a narrow, well-organised programme against the fraction of a large institution’s attention that actually reaches the same question. The incumbent’s nominal resources may be enormous. Its usable resources for that problem may be constrained by budgets, internal ownership, incompatible systems or the need to justify work to several layers of management. A challenger can exploit that difference without matching the institution’s total capacity.

An R&D leader should ask how long it takes a credible idea to obtain a fair test. Not a meeting, a presentation or a pilot announcement: a test that can change a decision. If the answer is dominated by internal procedure rather than scientific necessity, the organisation has created an opening. AI-assisted outsiders can make more attempts at that opening while the incumbent is still deciding which department is responsible for it.

The defensive response is to emphasise everything the challenger lacks. Some of those observations will be correct. The challenger may lack production experience, access to field data or the ability to provide service at scale. A useful strategy turns those advantages into better collaboration and faster validation. A weak strategy merely recites them as reasons to ignore the challenger’s progress in the part of the process it can already perform.

The enterprise can deploy the same tools. It can connect them to richer data and stronger experimental resources. It can give small internal teams authority to investigate bounded problems and partner with external groups whose work survives review. In that scenario, broad intelligence strengthens the incumbent’s ability to improve. The threat is not that incumbents are forbidden to adapt. It is that adaptation requires confronting the procedures and incentives that previously made their size feel like sufficient protection.

### The laboratory’s product includes its evidence

For the smaller team, a result becomes more valuable when a customer can inspect how it was obtained. The research record should identify the inputs, methods, failed alternatives, evaluation conditions and remaining uncertainties. This does not require disclosing every commercial secret. It requires enough evidence to support the claim being sold. A customer buying a component or analysis needs to know what has actually been demonstrated.

AI can help prepare that record, but the record must not be a retrospective story assembled to make the final answer look inevitable. Failed attempts and changes of direction should remain visible where they affect interpretation. A later reviewer should be able to distinguish evidence available at the time of a decision from information added after the outcome. The quality of that distinction can determine whether the work is reusable or merely persuasive.

An institution evaluating outside work should ask for the same standard it applies internally. If its own teams rely on undocumented judgement while it demands impossible documentation from newcomers, the standard is serving as a barrier. If it accepts polished external claims without adequate evidence because the provider is fashionable, it has abandoned the standard. The opportunity is to make the quality of the work more decisive on both sides.

A million agents do not eliminate scarcity. They rearrange it. The ability to frame a consequential question, obtain independent evidence, manage a coherent programme and carry a result into the world remains uneven. The small lab becomes powerful when it combines broader intellectual access with a focused command of those remaining tasks. The incumbent remains powerful when it uses its assets to shorten that same path.

That is the future worth taking seriously: many more organisations able to make credible attempts, competing and collaborating through evidence that travels beyond their walls. Some will employ only a handful of people. Some will be the great research enterprises we already know, rebuilt around different assumptions. The institutions most exposed are those that cannot explain why their scale improves the result once scale is no longer the only practical way to obtain the thinking.

## 08 · The researcher outside the gates

The most consequential user of scientific AI may be someone the established research system has never learned to notice. Not an undiscovered genius waiting for a machine to reveal a destined greatness. A capable person with a specific question, useful local knowledge and too little access to the intellectual and physical resources needed to investigate it. There are many reasons such a person remains outside the recognised centres. None establishes that the question is unimportant.

The promise of generally available intelligence is that more of those questions can become serious attempts. The promise is larger than convenience for people already surrounded by experts. It concerns the distribution of research capacity itself. A person should be able to obtain explanations, computational assistance and a path to checking without first acquiring the institutional identity that previously bundled those resources together.

That possibility should not be narrated as an accomplished equality. The world’s access to digital infrastructure remains profoundly uneven. ITU estimated that six billion people used the internet in 2025 while 2.2 billion remained offline. Even among people counted as online, the quality and affordability of access differ. A service delivered through a browser cannot be universally available in practice when the conditions for using the browser are not.[^S45]

### The first gate is material

A research assistant needs more than a momentary connection. Sustained work may require reliable electricity, a suitable device, enough bandwidth to move data and the ability to pay repeatedly. A person can technically access a service and still be unable to afford the experimentation needed to use it well. If every failed attempt has a meaningful cost relative to income, the user’s practical freedom to explore remains constrained.

Payment is another gate. A published price does not guarantee that a person can purchase the service using the financial infrastructure available to them. Currency conversion, billing rules and institutional purchasing procedures can all affect access. The relevant question is therefore not simply whether the provider has a website. It is whether the intended user can sustain a programme of work at a predictable total cost.

Service eligibility also has geography. OpenAI and Anthropic publish country and region lists for their offerings, with product-specific distinctions and additional conditions. The map samples six countries that appear in the inspected API lists. It does not imply that every model, every research configuration or every user in those countries has the same access. Nor does it treat areas not assessed in the figure as having no AI alternatives.[^S47][^S48]

> **Figure 20 · Who can actually use it (2D)** — Dated access map. See the corresponding figure in the web edition.

These distinctions are not reasons to abandon the worldwide premise. They identify the work required to make it real. A public policy that subsidises only a model subscription may leave the user unable to obtain data or run an experiment. A university that offers an online course but restricts all meaningful assessment to enrolled students may widen instruction without widening credibility. Access improves when the path is considered as a whole.

The second gate is language and context. A system may be available in a country while performing unevenly on the languages, documents or technical conventions its users encounter. Translation can help, but it can also erase a distinction that matters in a local measurement or operational record. A broadly deployed research tool needs evaluation against the work people actually do, rather than assuming that success on a dominant-language benchmark represents everyone equally.

The third gate is the right to use relevant evidence. A researcher can have an excellent model and still lack access to the papers, datasets or instrument records needed for the question. Public data can reduce that barrier, but poorly documented files may be difficult to interpret. A dataset without a clear account of how it was produced can create the appearance of openness while leaving the practical advantage with the people who already know its history.

### A programme built around a local problem

Consider a hypothetical engineering group investigating repeated failures in a community water-pumping system. The scenario is illustrative; it is not a reported case or an invented interview. The group knows which failures interrupt service, what spare parts are available and how operators actually use the equipment. It lacks a large research budget. Its immediate goal is to distinguish among several plausible causes and identify a change worth testing.

The first task is to assemble the record. Maintenance notes, operating hours, changes in water conditions and replacement histories may exist in inconsistent forms. An assistant can help organise the material and identify missing fields. The group must still check whether the records correspond to the same equipment and whether dates or units have been confused. The output at this stage is a usable description of what is known, not a diagnosis.

The next task is to compare explanations. An agent can help search relevant literature, explain unfamiliar mechanisms and propose analyses that distinguish them. A local engineer can reject assumptions that do not match the equipment or its operating conditions. The exchange is productive because each side supplies something the other does not automatically possess: broad computational assistance and specific knowledge of the problem’s physical setting.

The group can then develop a bounded test. It might compare operating patterns, inspect a component or arrange a measurement through a nearby facility. The important question is which observation would change the decision. A model that produces a long list of possible causes without helping prioritise evidence has not yet reduced the research burden. A useful assistant helps the team turn uncertainty into an affordable sequence of tests.

> **Figure 21 · A research programme outside the centre (2D)** — Illustrative workflow. See the corresponding figure in the web edition.

If the work produces a credible result, the group needs a route beyond the initial investigation. It may seek a manufacturer’s response, a partner capable of testing a redesign or a public purchaser willing to support a limited trial. The result should include enough evidence for those organisations to assess it. The team’s geographical location should not be treated as an answer to the technical question. Its documentation and the quality of its tests should matter more.

This scenario shows why generally available intelligence can be valuable without replacing every institution. The local group may still need a university laboratory, an experienced specialist or a manufacturer. The difference is the position from which it approaches them. It can arrive with a structured question, a coherent analysis and a proposed test, rather than asking for indefinite expert attention before it can demonstrate why the problem deserves consideration.

The institution’s role can become more focused and more useful. It can provide a difficult measurement, a rigorous review or access to equipment. It can help the group avoid a mistaken conclusion and strengthen the work. The relationship becomes a partnership around a testable question. That is a better use of institutional expertise than requiring every outsider to reproduce the institution’s entire educational pathway before being allowed to contribute.

### Expertise can become more widely practised

There is a risk that the language of universal intelligence encourages people to underestimate learning. A capable assistant can help someone work beyond their current knowledge, but it cannot make the consequences of ignorance disappear. The user needs enough understanding to recognise when a result is uncertain, when a method is inappropriate and when a specialist is required. Wider access should include ways to acquire that understanding through practice.

This creates an opportunity for education to change its sequence. A learner can begin with a meaningful problem, obtain explanations as needed and test a small part of the work. Formal study can then deepen and organise what the learner has encountered. The approach will not suit every topic or every student. It can make advanced work more approachable for people whose circumstances do not permit a conventional full-time route through an institution.

Assessment becomes crucial. If a learner can produce polished work with assistance, the institution needs a better way to determine what the learner understands. That may include oral explanation, examination of choices, reproduction under changed conditions or a demonstration of how errors were detected. The objective is not to preserve an old assignment format by banning the tool. It is to identify the competence the credential is supposed to represent.

The same principle applies outside formal education. A person who presents an AI-assisted analysis should be able to explain its assumptions and respond to a challenge. A customer or collaborator should not need to guess whether the person understands the work. Assistance can broaden participation while leaving responsibility visible. An institution that develops fair ways to assess such work can enlarge its public role rather than surrender it.

### Recognition is a separate gate

Even a correct result can struggle to obtain attention. Reviewers have limited time, and affiliation is an easy sorting device. It provides a rough signal about training, resources and accountability. The problem arises when the signal becomes a substitute for examining an artefact that could be assessed more directly. As the supply of capable outsiders grows, that shortcut becomes both less fair and less informative.

The answer is not to require experts to read every unsolicited document. That would create an impossible burden and reward volume over quality. The answer is to design better entry points: concise statements of the claim, reproducible materials, clear evidence requirements and bounded review processes. A contribution can be filtered rigorously without requiring its author to belong to a preferred institution. The filter should be attached to the work.

AI makes this problem more urgent because it increases the supply of plausible-looking submissions. Some will contain errors that a knowledgeable reader can identify quickly; others will be harder to assess. Institutions need procedures that protect attention while remaining open to genuine contributions. A blanket dismissal of machine-assisted work would throw away useful results. Uncritical acceptance would bury serious research under a larger volume of persuasive noise.

The public record can help. A contributor who preserves code, data, methods and corrections gives reviewers a more economical way to inspect the claim. An institution that records why a submission failed can make the next attempt better. The process should avoid turning every rejection into a permanent judgement on the person. Research advances through revised work, and broader participation requires routes through which correction can lead to renewed consideration.

### Open science needs more than downloadable files

UNESCO’s Recommendation on Open Science treats infrastructure, participation and sustained investment as part of the project, alongside access to knowledge. That wider view is useful here. If advanced intellectual assistance becomes more available, the remaining conditions for effective participation become easier to see. Public investment can support the parts of the research process that a subscription alone will not supply.[^S46]

Shared facilities are one example. A regional laboratory can make specialised measurements available to schools, small firms and independent groups under clear conditions. The facility still needs staff, maintenance and standards. AI may help users prepare better requests and interpret results, reducing some of the burden on those staff. The combination can widen access more effectively than either the physical facility or the digital tool would alone.

Another example is a durable public research record. Data, software and methods should remain accessible when a project ends or a commercial service changes. Maintenance is not a glamorous activity, but it determines whether future users can build on the work. If every new group has to reconstruct the same basic pipeline, the apparent abundance of intelligence is partly consumed by avoidable repetition.

A third example is support for translation between local problems and formal research questions. Communities often know which conditions need improvement without possessing the language used by research institutions. Intermediaries can help formulate the question without taking ownership of it away from the people who understand the setting. AI can assist in that translation. The institution should judge success by whether the resulting work serves the problem, not merely whether it produces a publishable abstraction.

### A wider geography of ambition

The distribution of questions matters because research priorities are shaped by the people and organisations able to pursue them. A problem that is expensive for a small community may be commercially uninteresting to a global company. A local manufacturing constraint may not appear in the benchmarks that guide model development. Wider access to intellectual work can make more of these problems investigable, even if the eventual solution still requires partnership or public support.

This is one of the strongest reasons to resist a future in which every capable researcher must attach themselves to a small number of central platforms. The platform may provide useful tools, but it should not become the only institution able to define worthwhile work. Users need ways to preserve their data, compare systems and carry their methods elsewhere. Otherwise, the old gate has merely moved from a university office to a service account.

The global opportunity also extends to collaboration. A small group can contribute a specialised piece of work to a larger programme without relocating every member to the same city. A researcher can seek criticism from someone with the right expertise rather than someone who happens to occupy the neighbouring office. These possibilities existed before the latest models. More accessible intellectual assistance can lower the preparation and coordination costs that have limited their practical reach.

The outcome will depend partly on whether institutions recognise contributors as partners. A university that invites outside data but reserves all interpretation and credit for itself has opened only one side of the relationship. A company that asks local users to supply valuable knowledge while providing no durable access to the resulting tools may reproduce an old imbalance with new software. Participation should include a fair account of who contributes, who decides and who benefits.

The challenge to established centres is therefore not that geography no longer matters. Physical resources, networks and experience still matter greatly. The challenge is that geography should become a less decisive barrier to the intellectual parts of a serious attempt. When a capable person can obtain help with the analysis, the institution’s refusal to consider the work becomes harder to justify through a simple appeal to where that person was trained or employed.

The promise of PhD-level intelligence for anyone around the world is worth defending precisely because it is not yet fulfilled. It names a larger distribution of possibility. The work ahead is to connect that possibility to devices, knowledge, evidence, facilities, fair assessment and useful outcomes. A society that supplies only the model has supplied a powerful tool. A society that opens the rest of the path has changed who is allowed to become a researcher.

## 09 · The university earns its place again

A university should be among the least surprised institutions in this story. It has helped create the mathematics, computer science, biology and engineering on which the new systems depend. It has trained many of the people building them. It has spent generations explaining that knowledge grows through criticism, evidence and the revision of assumptions. The awkward question is whether it will apply that account of progress to its own organisation.

The university’s strongest defence is its public contribution. Its weakest defence is the claim that difficult knowledge naturally belongs inside its administrative boundaries. Those propositions have often travelled together, making it easy to treat criticism of the boundary as criticism of the contribution. Wider access to intellectual work separates them. We can value scholarship, teaching and laboratories while asking whether every familiar barrier still serves those purposes.

The institution is not a single activity. It teaches, assesses, conducts research, maintains facilities, confers credentials, preserves knowledge and provides a setting in which people can work together over long periods. AI affects each function differently. A serious response begins by separating them. Announcing that the university is either obsolete or immune avoids the harder task of deciding what needs to change.

### Teaching must justify the time it asks for

An explanation is becoming easier to obtain. A student can ask for another example, a different derivation or a patient account of a concept they are embarrassed to admit they do not understand. The quality needs checking, but the supply of explanatory attention is changing. An institution whose teaching model relies heavily on delivering the same information to large groups should examine what the student receives beyond access to that delivery.

The answer can be substantial: a coherent curriculum, informed feedback, disciplined practice, a community of learners and an experienced teacher’s judgement about what matters. None is made irrelevant by a capable assistant. They become more visible as the institution’s real contribution. The lecturer’s authority should rest less on being the only convenient source of an explanation and more on helping students develop the ability to use and challenge explanations well.

MIT OpenCourseWare already demonstrates that sharing advanced educational material need not erase the institution that produced it. Its course materials are freely available without enrolment, while use of the materials does not itself confer an MIT credential. This is an important historical corrective: universities did not wait for generative AI to begin opening knowledge. The new tools can make such material more approachable, while creating a fresh question about how learning is assessed.[^S49]

The difficult issue is the price and structure of the educational experience. If students can obtain competent assistance with routine explanations elsewhere, they will reasonably ask why they must pay for an institution to ration those explanations through an inflexible schedule. The institution should be able to answer in terms of the learning it produces. Tradition is not an adequate substitute for showing that the experience develops capabilities the student would struggle to acquire alone.

A better model can use AI to identify where a learner is confused and reserve human attention for the questions that benefit from it. It can give students more opportunities to practise and receive feedback. It can ask them to compare competing explanations and identify mistakes. The objective is a more demanding education, not a cheaper imitation of a lecture. The institution should measure whether students become better at reasoning, not whether they spend the expected number of hours watching material.

### Assessment must look behind the finished object

When a student can generate a polished essay, program or analysis with assistance, an assignment that judges only the finished object becomes less informative about the student. That is a problem with the assessment, not a complete argument against the tool. The institution has to specify what competence it intends to certify and design a way to observe that competence under contemporary conditions.

For some purposes, unaided work remains useful. A student needs internal knowledge and fluency that cannot be replaced by consulting a system at every step. For other purposes, the ability to use tools responsibly is part of the competence itself. An assessment can examine how the student frames a question, checks an output, responds to changed assumptions and explains a decision. The balance should follow the subject and the learning objective rather than a universal rule about assistance.

Consider a technical project in which the student must defend the choice of model, identify the weakest assumption and reproduce the analysis after a relevant input changes. A generated report alone will not carry the assessment. The student must demonstrate understanding through interaction with the problem. This kind of examination can be demanding for teachers, but AI may also help prepare variants, organise evidence and reduce routine marking. The institution needs to redesign the work on both sides.

The credential then becomes a clearer statement. It can attest that a person has demonstrated specified abilities under defined conditions, including the ability to supervise advanced tools. That is more useful than pretending the labour conditions of the past remain unchanged. Employers should welcome such clarity. A degree that conceals uncertainty about what its holder can do will become less valuable precisely when the market needs better signals.

> **Figure 22 · What the university still does (2D)** — Institutional analysis. See the corresponding figure in the web edition.

### Research time should reach research

The case against unnecessary bureaucracy does not require AI. The UK’s independent review of research bureaucracy, led by Adam Tickell, examined the burden across funders and research organisations and produced recommendations for reducing it. The existence of such a review is evidence that the problem has been recognised within the system. It is not proof that every institution is equally inefficient, or that all administration is waste.[^S50]

AI raises the cost of leaving the problem unresolved. If an initial investigation can be performed more quickly, an approval process designed around a slower pace may become a larger share of the total delay. The institution can find itself spending longer deciding whether a bounded attempt may proceed than the attempt would take. Where that happens, the limiting resource is no longer the availability of intellectual work. It is the organisation’s willingness to let the work reach a test.

Some controls protect important interests: research integrity, appropriate use of funds, the welfare of participants and the reliability of shared facilities. The question is whether the procedure is proportionate to the decision. A small, reversible computational investigation should not necessarily require the same process as a consequential physical intervention. Treating them as equivalent consumes attention that could be devoted to the work that truly needs careful review.

The institution should be able to explain why each gate exists, what evidence it needs and how long it normally takes to act. It should also examine whether repeated requests contain information the organisation already holds. Automation can help reduce duplication, but a machine that fills the same unnecessary form faster leaves the underlying problem intact. The more valuable reform may be to remove the redundant request and make responsibility clearer.

A research leader can begin with a simple measure: elapsed time from a well-formed question to the first informative result. Break that time into intellectual work, necessary review, physical waiting and avoidable administrative delay. The measure will differ across fields, but the exercise exposes where the institution helps and where it obstructs. It is harder to defend a slow system once the delay has a name, an owner and no persuasive connection to quality.

### The doctorate should remain an apprenticeship in judgement

A PhD is not merely a prolonged demonstration that a person can generate difficult text. At its best, it teaches someone to sustain a research question, encounter failure, revise an approach and make a contribution that survives informed criticism. Those are demanding activities. AI can assist with parts of them while also creating ways to avoid learning them. The supervisor’s task becomes more important, not less, when a student can produce plausible work faster than they can fully understand it.

The institution should ask what experiences the student needs in order to develop judgement. Some routine labour may be educational at one stage and unnecessary at another. Writing an analysis from scratch can teach a method; repeating the same mechanical implementation many times may add little. Learning to inspect a generated proof or model can be a serious intellectual exercise if the student must understand why it works and where it fails.

The danger is a doctorate that becomes a management exercise in producing a large volume of machine-assisted output. More papers, longer appendices and more impressive diagrams do not necessarily indicate deeper understanding. The programme needs standards that reward the quality of the question, the validity of the evidence and the student’s command of the work. It should be possible to produce fewer outputs and a stronger researcher.

There is also a labour question. Doctoral researchers often perform substantial work that supports the institution’s research programme. If AI reduces the need for some tasks, the institution should not simply remove the training opportunities through which junior researchers previously acquired expertise. It needs to create new routes into meaningful responsibility. A research system that buys senior-level assistance while neglecting the development of future human judgement may create a delayed shortage of the very expertise needed to supervise it.

### Public goods are a stronger claim than protected access

The shared foundations beneath scientific AI deserve more attention than the image of a model working alone. Mathematical libraries, experimental databases, software, standards and educational material make later work possible. Their maintenance often lacks the visibility of a breakthrough. Yet a system that depends on them cannot remain healthy if the organisations supporting them lose the resources or incentives to continue.

CERN’s Open Data Portal offers an instructive example of public research infrastructure extending beyond a paper. It includes datasets, software, environments and documentation intended to support education and research. The value lies in making a path through the material possible, not merely placing files on a server. AI can help users navigate such resources, but the resources still need expert stewardship and a durable institutional home.[^S51]

> **Figure 23 · The public work beneath the model (2D)** — Attribution network. See the corresponding figure in the web edition.

This is where the university can make a compelling claim on public support. It can preserve knowledge that a commercial provider has little incentive to maintain. It can undertake research whose benefits are diffuse or distant. It can provide independent criticism of systems on which society increasingly depends. It can make facilities and expertise available through arrangements that serve a wider community than its enrolled students or immediate partners.

Those functions require money, and openness does not make them costless. A sustainable model must fund maintenance, documentation and the people who provide assistance. The argument for wider participation should not become an excuse to demand unlimited unpaid service from already stretched researchers. The institution needs to design access in a way that is fair to users and viable for the people doing the work. Clear standards and well-prepared requests can help both sides.

The same principle applies to independent verification. If companies produce more ambitious scientific claims, universities can organise careful assessments and public explanations. That work should receive recognition rather than being treated as secondary to producing another claim. The institution’s authority becomes stronger when people can see the service it performs in distinguishing what has been established from what has merely been announced.

### The challenge should reach governance

Universities cannot adapt solely through individual enthusiasm. A few teachers will redesign courses, a few laboratories will build useful workflows and a few administrators will automate forms. Without changes to incentives and authority, those efforts may remain exceptions. The institution has to decide what it will reward, which procedures it will revise and how it will assess the effects on learning and research.

A governing body should ask for evidence about outcomes rather than a catalogue of AI initiatives. Are students learning more effectively? Are researchers reaching informative results sooner? Can outsiders contribute through a fair process? Are shared resources better maintained? Has the institution reduced unnecessary work, or merely added a new layer of reporting about automation? These questions connect technological adoption to the university’s purpose.

The strongest resistance may arise where a change threatens a familiar source of status. A department can defend a procedure in the language of rigour when the procedure also protects its control over access. A credential can be defended in the language of quality when its meaning has become difficult to specify. These are not accusations against every university or academic. They are organisational possibilities that deserve examination, especially when the institution asks others to trust its capacity for self-criticism.

The answer is not to replace scholarly judgement with a market popularity contest. Some valuable work is difficult to explain, slow to mature or unlikely to attract a large audience. Institutions can protect it precisely because they have responsibilities beyond immediate demand. But protection should be tied to a clear intellectual purpose. The difficulty of a subject should not grant every surrounding procedure immunity from scrutiny.

### A university with a larger public

Imagine a university that treats broader intellectual access as an enlargement of its constituency. Independent learners use its material and can seek meaningful assessment. Small laboratories bring work to shared facilities through transparent procedures. Researchers use agents to reduce routine burdens while preserving the experiences through which judgement develops. Public datasets and proof libraries receive sustained support. Verification and careful explanation count as serious contributions.

Such an institution would still be selective about where it invests scarce resources. It would still reject weak claims and require appropriate preparation. It would still employ people whose expertise is difficult to replace. Its difference would lie in the basis of its authority. It would be valuable because it improves the work that passes through it, not because it can assume that serious work has nowhere else to begin.

That is the challenge I would put to academia. If your purpose is the growth and circulation of knowledge, a world with more capable participants should be an opportunity to fulfil it. If your response is mainly to defend the scarcity through which your authority was once organised, you are revealing a conflict between the mission and the business model. AI does not create that conflict. It makes the conflict harder to conceal.

The university can earn its place again, and in many areas it already does. The task is to make that earned contribution the organising principle of the institution. The coming abundance of intellectual assistance will not make scholarship unnecessary. It will make a poor explanation of scholarship’s institutional form increasingly expensive.

## 10 · The R&D department meets a competitor of another size

An enterprise can be extraordinarily good at research and still misunderstand why its research has been valuable. For years, its strength may have rested on a combination of talented people, accumulated records, expensive equipment, relationships with suppliers and the ability to deliver a working product. Success allowed these contributions to merge into a single reassuring phrase: our unique R&D capability. That phrase becomes dangerous when the price and availability of one contribution change sharply. It conceals the question that a new competitor will ask with much less sentiment: which part do we actually have to reproduce?

The challenger does not need to duplicate the department. It needs to solve a problem that someone will pay to have solved. The problem might occupy a narrow part of the incumbent’s workflow. It might be poorly served because its market was too small to justify an internal team. It might be a persistent inconvenience for customers who have learned to tolerate it. A company with formidable research credentials can leave all three kinds of opening. Its general excellence does not make every local arrangement efficient, and its overall size does not tell us how much effort reaches an individual question.

This is the point at which generally available research intelligence becomes an enterprise problem. The organisation previously compared itself with other organisations capable of assembling a similar department. That comparison selected its competitors before the contest began. If an independent group can rent substantial parts of the intellectual work, assemble public tools and purchase a limited set of experiments, the comparison changes. A procurement manager may meet a supplier that looks implausibly small beside the problem it claims to address. The appropriate response is a demanding test. Dismissing the supplier by counting its employees is an increasingly weak substitute.

### What the department actually contains

Take a research programme apart before declaring it protected or doomed. Someone defines the problem. Someone searches the literature and patents. Someone turns the question into a model, prepares data, writes software, proposes alternatives and decides which alternatives merit experiments. Facilities produce measurements. Engineers interpret failures, revise assumptions and integrate the result with the product. Manufacturing determines whether it can be made repeatedly. A commercial organisation persuades a customer to accept it, supports it and takes responsibility when it fails. The word research often travels casually across this entire sequence.

These activities have different dependencies. A literature search requires access to documents and the ability to distinguish relevant work from plausible distraction. A reliable measurement requires an instrument, a suitable sample, calibration and competent operation. A production change may require agreements across suppliers whose incentives are not aligned. The same model cannot be assumed to remove all three obstacles. Equally, the survival of the third obstacle cannot establish that the first remains a protected institutional skill. The honest unit of analysis is a task connected to a deliverable, with the necessary assets written beside it.

> **Figure 24 · Take the R&D department apart (2D)** — Task and asset map. See the corresponding figure in the web edition.

This decomposition reveals a distinction between owning an asset and needing an asset. A new materials company may need measurements from a particular class of instrument without needing to buy the instrument immediately. A design group may require prototypes without operating its own factory. A software company can need a customer’s historical records without acquiring the customer. Contracting, collaboration and shared facilities are old practices. More capable digital research can increase their importance by enabling a small group to arrive with a much better specified request. The rented experiment becomes more valuable when the customer has already narrowed the uncertainty it is buying the experiment to resolve.

Access is nevertheless a commercial fact, not an assumption to hide in a diagram. Is the facility willing to serve outsiders? What does a usable test cost, including preparation and interpretation? How long is its queue? Can the result be reproduced elsewhere? Who owns the records? If the only realistic provider is controlled by the incumbent, the challenger’s intellectual strength may encounter an effective barrier. That barrier should be named correctly. It is a barrier in facilities, contracts or market structure. Calling it unique intelligence protects the wrong explanation and makes a policy response harder to design.

The incumbent should perform this exercise on itself. For each claimed advantage, identify the output it improves and the evidence that it does so. A proprietary dataset may contain unusual failure histories that materially improve decisions. It may instead contain years of inconsistently labelled records that few employees can retrieve. The first is a valuable resource. The second is an expensive memory problem awaiting repair. The fact that both sit behind a corporate firewall does not make them equally defensible. Secrecy can preserve value; it can also preserve ignorance about the absence of value.

### Where a small competitor enters

Consider an explicitly hypothetical supplier of diagnostic software for industrial equipment. It does not manufacture the equipment. Its proposed product combines an operator’s maintenance records with physical reasoning and a testable explanation of recurring faults. The supplier uses research agents to compare mechanisms, inspect code and prepare candidate models. Its commercial promise is deliberately bounded: improve the ranking of a specified set of maintenance investigations under agreed operating conditions. It must demonstrate that improvement on records it did not use to build the system, then survive a monitored deployment.

If the supplier succeeds, it has not replaced the manufacturer’s research department. It may have displaced a profitable service, reduced demand for a proprietary diagnostic package or become an attractive partner for the manufacturer. Each outcome matters. An incumbent can retain its factories, customers and major inventions while losing control of a valuable decision made around them. The relevant market change may be a redistribution of bargaining power within an industry, rather than the spectacular disappearance of a familiar company. A thesis that notices only extinction will miss much of the disruption it claims to anticipate.

The supplier can also fail for reasons that reveal where the advantage really lies. Its training records may omit rare faults. Its proposed explanations may be indistinguishable from one another given the available sensors. The operator may have no practical way to conduct the recommended inspection. An existing service team may resolve the problem faster because it recognises a pattern that was never recorded. These are meaningful failures. They tell the challenger what it must acquire or learn. They do not establish that a large R&D payroll is the only organisational form capable of eventually addressing the gap.

> **Figure 25 · Win one valuable piece (2D)** — Explicit commercial scenario. See the corresponding figure in the web edition.

A second entry point is a component whose performance can be evaluated independently. A smaller group might offer an improved optimisation routine, a measurement correction or a design tool that fits into an established process. A third is research for questions the incumbent has left below its investment threshold. A fourth is an interface that makes an existing capability usable for customers previously priced out of it. The opportunity is not identical in each case. Component suppliers depend on integration; specialist researchers depend on access to tests; interface businesses depend on continued access to the underlying capability. Their contracts should reflect those dependencies before enthusiasm becomes a valuation.

For buyers, the change could be beneficial even when the new supplier never becomes large. Credible alternatives can make an incumbent explain its prices, unbundle a service or improve a neglected product. A small lab can alter the negotiation without owning the whole market. This is one reason headcount is a poor measure of competitive significance. Another is that a challenger can specialise around a customer’s actual inconvenience while a larger organisation optimises around the economics of its existing portfolio. Intelligence that becomes easier to obtain may reward specificity as much as ambition.

### The internal competitor

There is a second small lab to consider: the team already inside the enterprise. It has access to records, equipment and experienced colleagues, yet lacks the authority to use them together. Its proposals cross functional boundaries. Its budget sits in one department while the benefit would appear in another. A modest experiment requires approval from managers rewarded for protecting current delivery. The organisation describes these obstacles as coordination. Sometimes coordination is genuinely necessary. Sometimes it is the polite name for a system in which nobody gains enough from a new result to bear the inconvenience of permitting it.

Research agents do not automatically repair these incentives. They can make the situation more embarrassing. When a team can prepare a credible experiment in a day but waits through several review cycles for access to an ordinary dataset, the delay becomes visible. When a prototype can be tested against a clear acceptance criterion but receives another request for a strategy presentation, the organisation exposes its own preference. The scarce resource is not always the ability to think. It can be the willingness to let a useful piece of thinking disturb a settled distribution of authority.

This is where the challenge should be sharpest. A company that tells investors it possesses rare intelligence should be able to explain why that intelligence spends so much time obtaining permission to perform reversible tests. A university that celebrates interdisciplinary research should be able to show how a person crosses its administrative boundaries with a good question. Neither institution has to abolish review. It has to connect review with a real risk, an accountable decision and a proportionate cost. Procedure earns its place by improving the work. Repetition alone is not evidence that it does.

An effective internal response gives a team a bounded problem, access to the required inputs and a pre-agreed route to a decision. Before the experiment starts, the organisation names what would count as a useful result, who can authorise the next step and what evidence would stop the work. The team records the cost of data preparation, failed attempts and human checking. It does not receive credit simply for producing a large number of model outputs. This structure creates a fair comparison between the previous process and the proposed one. It also prevents a promising trial from becoming a permanent demonstration with no customer.

### Data becomes an obligation

The enterprise’s records may be its strongest advantage, but their value becomes conditional on use. A model cannot recover a measurement that was never taken. It cannot make two incompatible definitions of failure identical by summarising them eloquently. It can help reconcile records and identify missing information, but the organisation still has to decide what the records mean. This work often lacks the glamour of announcing an advanced research system. It is also where a long-established operator may obtain a durable advantage over a smaller entrant with better access to generic reasoning.

Useful records connect observations with circumstances and outcomes. Which component was installed? Under what conditions did it operate? What changed before the failure? Which intervention was attempted, and what happened afterward? The value lies in those relationships. A collection of documents can contain them without making them easy to recover. An enterprise that systematically preserves the relationships creates a better environment for both people and models. It may also discover that some of its supposedly private expertise was public knowledge wrapped in internal formatting, while its truly unusual knowledge was buried in service notes no research team had examined.

The implication for staff is serious. Experienced engineers should be involved in deciding what is missing and which distinctions matter. Treating them as a temporary source to be extracted and discarded can destroy the very process that keeps the records meaningful. New products, suppliers and operating conditions continually create new exceptions. An organisation needs people capable of noticing when yesterday’s categories no longer describe today’s events. Broadly available analytical ability can make this judgement more valuable. It can also make organisations that suppress inconvenient observations more fragile, because errors can now propagate through more decisions more quickly.

The challenger has a corresponding duty. It should not describe every access restriction as protectionism. Customer records can contain confidential designs, commercially sensitive conditions or information about individuals. A workable proposal specifies the minimum data needed, how the evaluation will be performed and what can leave the customer’s environment. Good constraints can make collaboration possible. A refusal to define those constraints can make legitimate protection indistinguishable from a blanket refusal to compete. The practical aim is a test that reveals whether the proposed improvement works without requiring either party to surrender its entire business.

### Intellectual property after cheaper search

Cheaper research does not settle who owns an invention or who can use it. It changes the volume and distribution of attempts from which ownership disputes and commercial opportunities emerge. A company may encounter more competing designs, more efforts to work around a protected approach and more proposals based on public scientific knowledge. The strategic response cannot be reduced to generating a larger pile of patent applications. A portfolio is valuable through its relation to useful products, enforceable rights and negotiating needs. Producing documents at lower cost does not by itself strengthen any of those relations.

There is also a difference between a secret and a head start. An improvement may remain hard to imitate because its production depends on tacit knowledge, difficult calibration or a supplier network. Another improvement may become easy to reproduce once the principle is public. The latter business needs to recover value through speed, service, integration or continued invention. When research capability spreads, management should revisit which category each advantage occupies. The answer can change without any law changing. More people able to understand and extend an idea can shorten the period during which comprehension itself protects the original developer.

Contracts become especially important for small labs. If a customer pays for a successful experiment, can the lab apply the general method elsewhere? Can it publish a negative result? Does it retain access to enough evidence to defend its claim? Does the customer receive the materials needed to avoid permanent dependence on one supplier? These are business questions to resolve explicitly, with appropriate advice for the relevant jurisdiction. They should not be buried beneath a grand claim that intelligence is now free. The intelligence may be easier to rent while the rights to use its output remain tightly negotiated.

### A better enterprise, not merely a smaller one

The most interesting incumbent response is to turn the company into a better place for research to meet reality. It can expose well-defined problems, make selected test environments available and establish a procurement route for small suppliers. It can pay for independent replication and allow internal teams to challenge existing products. It can distinguish the experiments that need senior review from those that can proceed within an agreed limit. These changes do not require pretending the organisation has no responsibilities. They require designing those responsibilities so they do not protect every inherited inconvenience.

Such an enterprise might become more formidable as research intelligence spreads. Its installed products generate observations. Its facilities test proposals. Its engineers understand integration. Its customers provide a route to useful deployment. If the company combines those strengths with a wider supply of ideas, it can turn outside competition into productive collaboration without assuming that every valuable idea must originate internally. The difficult part is commercial generosity: outsiders need a credible prospect of being paid, retaining appropriate rights and receiving a decision. An invitation to contribute is not an open research system if the organisation captures all the benefit and delays every commitment.

There will be enterprises that make these changes and enterprises that merely buy more software. The difference will not be visible in a count of agent licences. It will appear in the time from a well-specified question to a trustworthy answer, the range of people allowed to propose a test, the quality of failures recorded and the number of useful changes that reach customers. These measures are demanding because they expose the whole process. That is their advantage. A department can make its output look impressive by counting activity. It has a harder time disguising a process that cannot absorb its own results.

The small lab with a million agents is therefore a challenge to enterprise self-knowledge. It asks what the company actually does that is difficult, what it merely used to be able to afford before others could, and what it prevents its own people from doing. The answers will differ by business. Some will confirm substantial advantages. Others will reveal that the institution has mistaken an expensive arrangement for a necessary one. General access to advanced intelligence does not abolish the enterprise. It raises the standard of explanation for why this enterprise, with these boundaries and these permissions, deserves to control this part of the work.

## 11 · ASML, IBM, Google, Dyson

Put four names on the table: ASML, IBM, Google and Dyson. Each can plausibly stand for exceptional technical intelligence. Yet the phrase means something different in each case. ASML connects computation with extraordinary demands on physical manufacture. IBM combines research traditions with methods, systems and enterprise relationships. Google can connect algorithmic research with infrastructure on which a small efficiency improvement has a large operational consequence. Dyson must turn interacting engineering choices into products that people use, maintain and judge. A claim that a small lab can reproduce their intelligence needs to say which of these accomplishments it means.

This comparison is deliberately harder than a contest between a brilliant startup and an inert bureaucracy. These companies are already doing technical work that involves AI, advanced modelling or both. They are not waiting for an essay to inform them that computation matters. Their strengths also do not guarantee that every service, workflow or product remains protected as outside capability improves. The serious question is where a small group could match a meaningful part of the work, and what would still separate that result from a competitive business. Answering it requires more respect for engineering than a slogan about replacement provides.

### ASML: a model meets a wafer

ASML reported €4.7 billion in research and development costs for 2025. Its financial discussion also described using AI in R&D and other operations, and its relationship with Mistral included an approximately 11 percent fully diluted stake following the 2025 investment. These are sufficient reasons to reject the portrait of a company serenely ignoring the change. They do not answer every competitive question. They establish that the starting point is an incumbent actively investing in both physical technology and AI, rather than an institution whose only response is disbelief. [^S15]

The technical relationship is especially instructive. ASML describes computational lithography as modelling the manufacturing process using data from its machines and test wafers. Those models help alter mask patterns to compensate for effects that arise during printing. Its account of metrology describes measurements feeding adjustments to alignment and other aspects of the process. Computation and measurement are therefore intertwined in the public description of the business. An improved model needs contact with the process it claims to improve. A better instrument needs software capable of interpreting and using what it observes. [^S52][^S53]

> **Figure 26 · The machine and its measurements (3D)** — Engineering explanation. See the corresponding figure in the web edition.

Imagine a smaller lab proposing a better way to estimate one class of process error. This is a hypothetical entry point, not a claim about an existing ASML supplier. The lab might use a large number of agents to investigate relevant mathematics, compare candidate algorithms, construct tests and search for failures. It could produce an unusually strong software artefact with a small permanent staff. The next question would be whether the artefact improves decisions on representative measurements under the required operating conditions. An elegant result on convenient public data would be an invitation to investigate, not a finished qualification.

The incumbent’s advantage is partly the ability to ask that next question properly. Which errors matter in production? Which observations are reliable? How does a correction interact with other corrections? Does an apparent improvement disappear when the process changes? These questions arise from a relationship with machines, customers and repeated measurements. A million agents reading the same public literature would not automatically possess those observations. But a smaller lab with a suitable partner might acquire enough evidence to address a bounded problem. The physical barrier is substantial; its size depends on the scope of the claim.

This distinction prevents two symmetrical mistakes. One is to treat a convincing algorithm as a replacement for a lithography business. The other is to treat the difficulty of building the whole business as proof that no software, analytical or service contribution can become contestable. A company can remain indispensable while buying more intellectual work from outside, facing stronger alternatives in adjacent activities or changing how it charges for its own contributions. Competition often reaches the seams before it reaches the centre. The seam is economically interesting precisely because the centre continues to function.

For ASML, a constructive future could involve faster investigation of process anomalies, broader searches across modelling choices and better tools for engineers working with complex evidence. External teams could contribute under carefully defined data and evaluation arrangements. Internal teams could spend less time preparing routine analyses and more time deciding what a surprising measurement implies. Whether any particular proposal delivers those benefits must be established in the relevant workflow. The strategic point is that the company’s physical and informational strengths can complement widely available intelligence. They do not have to remain valuable only by keeping all intellectual work inside.

The public should also resist an apparently flattering but unhelpful description of such a company as pure genius embodied in a machine. The machine embodies accumulated engineering, suppliers, instruments, software, manufacturing practice and customer feedback. Recognising those contributions makes the company’s accomplishment more impressive, not less. It also makes the boundaries of competition intelligible. A smaller lab may match a piece of the reasoning. It still needs a route into the network that allows the reasoning to matter. That route can be blocked, negotiated or deliberately opened. Each choice produces a different industrial future.

### IBM: expertise that can be shared

IBM offers a different test because some of the capability in question has already been released for others to use. In December 2024, IBM Research described a family of open-source models for materials discovery, including models based on molecular strings and graphs, alongside methods for combining their representations. The work used public molecular databases and evaluated prediction tasks. The release illustrates an incumbent making research tools available beyond its own organisation. It does not establish that every user can discover a useful material, manufacture it or earn a return from it. [^S16]

Once a method is available, a small lab can concentrate on a particular material, property or customer requirement. Its distinction may come from a narrow dataset, an informative experiment or a close understanding of an application. In this scenario, the lab’s use of a public model is not evidence that the model developer has become irrelevant. It is evidence that value can be created at several points in the research process. The developer may provide methods and infrastructure. The specialist may provide the question and validation. A manufacturer may provide a route from promising sample to consistent production.

This is an important corrective to the claim that enterprise R&D must always defend itself through exclusivity. An organisation can benefit by making a method widely used, then contributing the integration, experiments or support that make the method useful in demanding settings. That strategy still has to work commercially; openness is not a substitute for an economic model. But it demonstrates that spreading capability and retaining a valuable institutional role are compatible. The choice is not simply between hoarding intelligence and surrendering the organisation’s purpose. The organisation can change which contribution it asks others to pay for.

Now consider the uncomfortable version of the same argument. If a customer can obtain a strong first analysis from public tools, the supplier cannot indefinitely price that analysis as if it required a uniquely assembled team. It must explain what additional work produces the difference between a plausible answer and a dependable result. Perhaps the difference is substantial: integration with existing systems, evidence about rare failures, reproducibility or responsibility for deployment. Those are concrete contributions. If the difference is mostly a familiar brand attached to generic reasoning, the commercial defence becomes harder as customers learn to evaluate alternatives.

This pressure can improve the quality of technical services. A contract can specify the evidence a result must contain, how it will be tested and what happens when the proposed method fails. A small supplier can compete on that evidence. A large supplier can demonstrate the value of its broader experience. The customer benefits when both are forced to identify the uncertainty they remove. The danger is a market in polished reports that are cheaper to produce but no more connected to reality. Better models will not prevent that outcome unless buyers learn to distinguish an informative deliverable from a persuasive one.

For IBM-like research organisations, the institutional opportunity is to connect public methods with durable systems of evaluation. That can include useful datasets, clear benchmarks, experimental partnerships and documentation that lets others understand where a model fails. None of these activities is made unnecessary by a large population of agents. More agents can increase demand for them. A shared method becomes more valuable when its users can compare results and investigate disagreement. The organisation earns a role through stewardship of that process rather than through a claim that no outsider could possibly perform the underlying intellectual work.

There is a broader lesson for corporate research culture. Publishing a useful tool can create capable outsiders. Some will become customers, collaborators or competitors. An institution comfortable only with the first two categories may find openness emotionally more difficult than its strategy documents suggest. Yet the competitor may be the most revealing participant: it identifies an application or commercial arrangement the original organisation did not pursue. The appearance of that competitor is not proof that releasing the method was a mistake. It is evidence that the capacity to use an idea is no longer confined to the institution that helped develop it.

### Google: research with somewhere to land

Google’s AlphaEvolve work illustrates the value of a strong evaluator and a ready place to apply a result. In its 2025 account, Google DeepMind reported improvements in scheduling and computing kernels, alongside mathematical constructions. One scheduling improvement reportedly recovered an average of 0.7 percent of Google’s worldwide compute resources. A 23 percent speedup applied to a particular kernel was reported to reduce the relevant model’s training time by 1 percent. Those are company-reported operational results, and the denominators matter: a faster component is not an equally faster whole system. [^S10]

The rank-48 matrix-multiplication construction discussed earlier belongs to a different category. Its algebraic identity can be checked separately from claims about operational speed. That separation helps explain why the research programme is interesting. Some candidates face a mathematical test. Others face a workload, a hardware configuration and an integration test. The search becomes productive when the evaluator rewards the property that actually matters. A small lab can use the same general principle without having Google’s infrastructure, but it must find a domain in which it can construct a meaningful evaluator and obtain access to a real use.

Google’s scale can turn small local improvements into significant benefits. It also gives researchers unusual opportunities to test ideas against demanding systems. This is a complement to advanced intelligence, not a reason that intelligence must always be confined to a large organisation. A smaller company may have a narrower workload whose requirements are unusually well understood. It may search for an algorithm tailored to that workload, validate it independently and sell the result. The smaller search space can be an advantage if it corresponds to a valuable need that a general platform has little incentive to optimise individually.

The challenge for a new lab is to avoid mistaking a benchmark for a business. A routine can be faster under one input distribution and slower under another. It can reduce arithmetic while increasing memory movement. It can win a test that omits setup costs or error handling. The evaluator must reflect the intended use closely enough that improvement transfers. Agents can help design adversarial tests, but the customer’s system remains the final source of relevant constraints. The lab’s product is therefore partly the improved routine and partly the evidence that the routine is an improvement in the promised setting.

There is an interesting competitive asymmetry here. A platform can distribute a broadly useful improvement very quickly across its own systems. A specialist can focus on an unusual problem with an intensity the platform would not justify. General access to strong models can strengthen both. It can make the platform’s researchers more productive and the specialist’s minimum viable team smaller. The outcome depends on access to customers, interfaces, compute and evaluation, as well as on model capability. A prediction that only the large company wins overlooks specialisation. A prediction that size no longer matters overlooks deployment.

For customers, the constructive future is a market in improvements that can be demonstrated rather than merely asserted. A software component arrives with reproducible tests, clear limits and a method for checking it on the customer’s workload. A research claim includes the unsuccessful alternatives needed to understand how the final result was selected. A supplier can be small without asking the buyer to accept uncertainty on trust. A large supplier can justify its price through broader coverage and reliable integration. This is a higher standard of technical exchange than either brand deference or enthusiasm about autonomous agents alone provides.

Google also complicates any clean division between incumbent and disruptor. In the same account, it can be a frontier research organisation, a provider of tools to others and an operator whose established businesses face new forms of competition. An organisation does not occupy one permanent position in technological change. It can open a field in one place while defending control in another. The analysis should follow specific capabilities and dependencies. Labels such as old enterprise and new lab are useful only until they start concealing the different roles played by the same company.

### Dyson: the product has to survive its owner

Dyson’s public engineering descriptions place motor control, electromagnetic design, thermal and fluid behaviour, structural dynamics, acoustics, testing and manufacturing in the same broad development process. They also describe machine learning and data work spanning laboratory and manufacturing records. These descriptions are company accounts of the work, not independent evidence that a particular product outperforms every alternative. They are enough to show why matching one engineering calculation would be a narrow claim. A consumer product has to reconcile several properties that can pull against one another. [^S17][^S54]

> **Figure 27 · The motor that survives the workshop (3D)** — Engineering model. See the corresponding figure in the web edition.

A generic motor design makes the issue tangible. A proposed change may improve an electromagnetic calculation while making heat harder to remove. A smaller assembly may become more difficult to manufacture consistently. A control strategy may behave well in a clean test and poorly when components vary. The product also has an acoustic character, a weight, a maintenance burden and a price. A useful design search must preserve the relationships among these requirements. A model that optimises the easiest available score can produce a technically impressive answer to a commercially incomplete question.

Now give a small design lab access to advanced research agents. It can investigate more alternatives, automate parts of its modelling, compare materials and prepare better experiments. It can use outside manufacturing and test services where available. It might build a specialised component or a product for an underserved application. The proposition becomes credible when it identifies which properties matter to that application and demonstrates them on physical prototypes. It becomes less credible when it claims that a large number of simulated candidates makes product engineering a solved problem. The number of candidates says little about the realism of their evaluation.

The incumbent’s response need not be to defend every existing development step. Some steps may exist because earlier tools made iteration expensive. More capable assistance may allow engineers to investigate an awkward trade-off earlier, preserve the reasoning behind rejected designs or identify a test that distinguishes competing explanations. A company with strong facilities can benefit from a better supply of candidates. It can also become overwhelmed if the system produces more proposals than its tests can handle. The relevant design question is how to improve selection and learning, not simply how to fill the pipeline faster.

For the challenger, customer understanding may be as important as technical reach. A product that solves a narrower problem well can succeed without winning every specification on a comparison chart. An application may value repairability, quiet operation, a particular form factor or compatibility with an existing process. These are illustrative possibilities, not claims about Dyson’s current product gaps. The wider point is that broadly available research intelligence can enable more groups to explore different definitions of a good product. Competition can expand the range of questions asked, rather than merely accelerate the race along one established metric.

That expansion is a constructive form of disruption. Customers may receive products designed around needs that a large portfolio overlooks. Incumbents may find new component suppliers or identify demand more quickly. Engineers may spend more time evaluating physical behaviour and less time recreating routine analyses. The benefits will depend on whether companies preserve the evidence from failed tests and allow it to change the design. An agent population that repeatedly rediscovers why an old idea failed is expensive activity. A research system that carries the reason forward can make the next attempt materially better.

### Four companies, four different defences

> **Figure 28 · Four research advantages, examined (2D)** — Comparative analysis. See the corresponding figure in the web edition.

The comparison leaves no honest basis for a single numerical score of how exposed these companies are. Their work differs, their assets differ and the relevant challengers would enter at different points. What can be compared is the logic of the defence. Does the company have unusual observations? Can it test a proposal under relevant conditions? Can it integrate the result and deliver it reliably? Can it retain a customer’s trust after a failure? These questions identify contributions that remain meaningful when advanced reasoning becomes easier to obtain.

They also identify claims that deserve less deference. A large research budget is evidence of expenditure. A famous history is evidence of past accomplishment. A department’s size is evidence of how it is staffed. None of those facts alone establishes that an outside team cannot now solve a specified problem more effectively. The company should welcome a well-designed comparison if its advantage is real. It may still choose not to share particular records or facilities, but it should distinguish that commercial choice from a scientific claim about the limits of outsiders. One concerns control of access; the other concerns what someone can do.

The prospect of millions of PhD-level agents makes this distinction unavoidable. At its strongest, the premise suggests that small organisations could command an extraordinary breadth of research attempts. It does not mean they inherit every observation, instrument, manufacturing process and relationship accumulated by the four companies. The interesting future lies between those statements. More groups become capable of making serious proposals. Incumbents face more credible challenges to particular activities. Collaboration becomes possible with partners that previously looked too small. The institutions that prosper will be those that know which contribution they add and are willing to prove it repeatedly.

That is a demanding standard, but it is also an optimistic one. It allows ASML to remain an extraordinary engineering institution while buying or incorporating more outside intelligence. It allows IBM’s shared methods to support independent research businesses. It allows Google’s research to improve large systems while specialists find valuable problems elsewhere. It allows Dyson and smaller design labs to explore different products with better tools. Disruption need not mean that every great company disappears. It can mean that greatness ceases to confer an uncontested right to decide who is qualified to contribute.

## 12 · Five industries after the research bottleneck

The consequences become clearer when we stop asking whether AI will disrupt industry and ask what a customer, engineer or patient could experience differently. A research breakthrough is not yet an industrial future. Between the two sit decisions about what to test, who may test it, who pays and who changes an operating process. Those decisions vary enough that a single story of acceleration becomes misleading. Five industries provide a useful range: oil and gas, pharmaceuticals and biotech, advanced manufacturing and industrial design, energy and utilities, and enterprise software and technical services. They share a new supply of intellectual capability. They do not share a single clock.

The futures below are conditional scenarios grounded in present research and operating practices. They are not claims that particular companies have already made the changes described. In each case, the question is the same: where could a smaller group now make a serious contribution, how could an incumbent use that contribution, what benefit could reach someone outside the research team, and what would still have to be demonstrated? This approach makes room for disruption that helps an industry function better. It also makes it possible to identify the organisations whose resistance protects a genuine requirement and those whose resistance mainly protects an arrangement.

### Oil and gas: opening the interpretation

Oil and gas is an attractive target for the argument because its operations are physically demanding, information is unevenly distributed and mistakes can be costly. It would nevertheless be wrong to portray the sector as untouched by computation. The IEA describes oil and gas companies as early adopters of advanced computing and identifies applications including subsurface analysis, reservoir simulation, maintenance and leak detection. The issue is therefore the next change in who can perform useful analysis and how widely that analysis reaches, not the arrival of the first computer in a previously analogue industry. [^S14]

One opening lies in the interpretation of incomplete evidence. A subsurface model combines observations with assumptions about what exists between them. Two models can explain a set of observations while implying different responses to a proposed intervention. A useful research system should preserve those alternatives, identify which assumptions drive the difference and propose an observation that would distinguish them. Its contribution is not necessarily a single more confident answer. It may be a better account of why the decision remains uncertain and which uncertainty is worth paying to reduce.

> **Figure 29 · Below the oilfield (3D)** — Scientific or educational model. See the corresponding figure in the web edition.

In a plausible future workflow, a specialised lab works with an operator on a bounded maintenance or interpretation problem. The operator supplies an agreed subset of records and access to relevant engineers. Agents help reconcile units, compare mechanisms, inspect model code and prepare alternative explanations. The lab proposes tests with different costs and expected informational value. An operating team chooses a test, obtains the measurement and reviews the result. The new capability enters through a research service whose promise is inspectable. It does not begin with an autonomous system taking control of an installation.

The incumbent benefits if it can investigate more neglected problems, make better use of its own records and direct field work toward the most informative observations. A smaller supplier benefits if it can demonstrate expertise without owning the whole asset base. Workers benefit when the system makes the basis of a recommendation clearer and preserves the evidence from previous interventions. The operator’s customers could benefit through improved reliability or lower avoidable costs. These outcomes are connected, but none follows automatically from a better model. A maintenance recommendation that cannot be acted upon during a realistic service window may have little immediate value.

Methane monitoring provides a concrete example of the distinction between seeing and acting. UNEP’s July 2026 account describes its MARS system combining satellite observations with AI, followed by independent analyst review before notifications are issued. Governments and companies then have to investigate and respond. This is an operating chain in which digital analysis expands what can be noticed while human and institutional actions determine whether an emission is addressed. The AI contribution is real, but it does not turn a detected plume into a repaired piece of equipment by itself. [^S55]

That distinction suggests a better performance measure than the number of alerts generated. Count verified events, false alarms, time to investigation, time to an appropriate intervention and the measurements available afterward. A system that produces many alerts but overwhelms the response team may underperform a more selective one. A system that improves detection and is paired with a clear responsibility to act can produce a useful change. The enterprise’s competence is the whole chain. Its research department cannot claim success solely at the point where the information leaves its screen.

The deeper disruption is that interpretation may become less exclusive. Independent analysts, smaller service companies and public institutions could become better able to challenge an operator’s account of a technical or environmental question. They still need suitable evidence and careful methods. The operator should not be expected to accept every outside claim. It should expect to encounter more outside claims that deserve an answer. The industry’s domain knowledge remains important, but domain knowledge becomes a contribution to a contestable explanation rather than a permanent right to close the discussion.

### Pharmaceuticals and biotech: buying a better next experiment

The biological examples earlier in this essay show why this industry attracts so much attention. Protein prediction and design can expand the set of candidates and hypotheses a team can examine. Analytical assistance can reduce the labour required to interpret some measurements. Yet a candidate’s apparent promise and a useful treatment remain separated by questions about mechanism, exposure, safety, efficacy, manufacture and delivery. Clinical research exists because effects in people cannot simply be read off a molecular representation. The constructive future is a research process that asks better questions and learns more from its experiments, with clinical claims reserved for clinical evidence. [^S34]

A small lab might begin with an assay or a narrowly defined biological problem rather than a claim to replace a pharmaceutical company. Its agents compare literature, inspect public data and propose candidate designs or experimental conditions. The lab purchases or partners for the relevant tests. It preserves the full denominator of attempted designs and records inconclusive outcomes. If a result reproduces, it can become the basis of a collaboration, a tool or a further research programme. The lab’s size no longer tells us much about the sophistication of its initial hypothesis. Its experimental record tells us much more about the hypothesis’s value.

The most productive customer for such a lab might be an incumbent with the resources to pursue the next stages. This is not a concession that disruption has failed. A larger number of credible research suppliers can change which questions receive attention, how early evidence is priced and who has negotiating power. An established company might license a result, fund replication or offer access to a specialised assay. A smaller group might retain a method and serve several customers. The market changes when entry into a meaningful part of the process becomes feasible for organisations that could not previously assemble the required intellectual breadth.

There is also a valuable form of success that produces no promising molecule. An experiment can reveal that an attractive hypothesis was wrong, that an assay was misleading or that a result depends on an unrecognised condition. If the evidence is reliable and available to the next team, that failure can prevent further waste. Cheap candidate generation makes this contribution more important. Without a record of what failed and why, more capable agents can repeatedly propose variants of ideas that already encountered a decisive obstacle. The industry needs a memory of evidence, not merely a growing library of positive stories.

An incumbent could respond by funding experimental capacity and methods that discriminate among candidates earlier. It could structure collaborations so outside teams receive a clear answer and appropriate credit. It could compare research systems on the information obtained per unit of total effort, including failed experiments and human review. These measures are harder to advertise than a large design count. They are closer to the work that determines whether a programme advances. A more impressive search becomes economically useful when it changes the next decision, not when it only makes the list of possibilities longer.

For patients, the eventual benefit could include more attention to questions that were previously neglected and better use of research resources. Affordability, availability and clinical value still depend on decisions beyond discovery. A treatment developed more efficiently can remain expensive or inaccessible. A technically promising programme can be abandoned for commercial reasons. General access to research intelligence therefore creates an opportunity to broaden the research agenda, but it does not decide the agenda’s priorities. Public funders, patient groups, clinicians and companies will still influence which uncertainties receive sustained attention. Their choices become more consequential when more questions are technically approachable.

Progress should be measured at the stage actually reached. A reproducible binding result is evidence about binding. A well-controlled biological experiment is evidence about its measured outcome. A clinical study addresses its specified population and endpoints. Combining these into one generic success rate conceals the very transitions on which medicine depends. The smaller lab can earn a place by being exceptionally clear about those transitions. The incumbent can earn its larger role by carrying the work through them competently. Both claims are stronger than a contest over who can generate the most confident narrative about a molecule.

### Manufacturing and design: more alternatives, fewer repeated mistakes

Advanced manufacturing offers a wide range of bounded problems in which strong research assistance could matter. A team may investigate a recurring defect, compare a set of geometries, improve a process parameter or evaluate a material substitution. The difficulty often lies in interactions among requirements. A change that improves one property can worsen another, and a design that works as a prototype may be difficult to make consistently. The earlier discussion of ASML and Dyson illustrates these relationships at different scales. The general opportunity is to improve the quality and breadth of iteration while preserving contact with physical tests.

A specialised supplier could offer a component or process improvement supported by a transparent comparison. It would state the baseline, the conditions of the test, the properties measured and the limitations. Agents could help prepare the design space, inspect simulations, construct alternative explanations and document results. Physical samples would determine whether the claimed advantage survives manufacture. The customer would then evaluate integration and reliability. A small team might perform the intellectual preparation that previously required a larger department, while buying carefully chosen fabrication and test services. Its business would depend on access to those services at workable cost and speed.

The incumbent’s opportunity is to make its existing experiments more informative. A failed prototype should leave behind more than a status label. Which assumption failed? Was the material different from its specification? Did the test represent intended use? Did the design work physically but fail on cost or assembly? A research system that can retrieve and compare those answers gives the next team a stronger starting point. This is a practical advantage for established manufacturers: they have encountered many failures. It becomes an advantage only if their records preserve enough context to make the failures usable.

The customer’s benefit could be a product that lasts longer, wastes less material, is easier to repair or better fits a specific need. These are not interchangeable achievements. A lower manufacturing cost can coexist with a shorter useful life. A more efficient component can complicate repair. A responsible comparison keeps the customer’s relevant outcomes visible rather than selecting the single metric that makes the design look best. More widely available engineering intelligence should make it possible to explore a broader set of trade-offs, including needs that a mass-market design did not prioritise.

This creates room for smaller companies without requiring them to defeat an incumbent on every dimension. A local supplier may understand a particular installation, climate or maintenance practice unusually well. It can combine that knowledge with advanced modelling and external tests. The large manufacturer may respond with a partnership, a revised product or a more flexible design process. The industry benefits when unusual needs become technically and commercially addressable. The disruptive effect is partly a reduction in the minimum market size needed to justify serious engineering attention, provided the costs of testing and manufacture can also be managed.

The measures should follow the product into use. Track defect rates, rework, material waste, reliability and the cost of maintenance under specified conditions. Keep design-cycle time separate from the time needed to qualify a change. Record whether the improvement persists across production batches and suppliers. A system that creates prototypes quickly but increases downstream failures has moved work around rather than removed it. Conversely, a system that takes slightly longer to compare alternatives but prevents an expensive failure may be highly productive. The relevant unit is the useful, reliable product, not the speed of the most photogenic step.

### Energy and utilities: intelligence inside a physical network

Utilities face a different combination of constraints. Their decisions must reconcile changing demand, variable supply, equipment condition and the limits of a network. Better forecasts and analysis can improve choices, but a forecast is not an extra transformer and an optimisation result is not a new transmission line. The useful opportunity is to make better decisions about the infrastructure that exists, identify where new investment matters most and detect problems early enough to respond. Advanced research assistance can contribute to each activity if its recommendations are tested against the operating conditions that matter.

Dynamic line rating is an instructive existing example of computation connected to physical conditions. The US Department of Energy describes using weather and operational data to update estimates of the capacity of transmission lines. Its account traces research, software, commercial partnerships and utility adoption. This is not proof of an autonomous research agent; it is evidence of the kind of complete route from analysis to operation that new research systems will need. A more capable model would still have to fit into sensing, communication, operating procedures and the physical limits of the equipment. [^S56]

A smaller lab could enter through planning studies, asset diagnostics or methods for identifying bad data. It could compare alternatives on historical records and then operate in a mode that produces recommendations without controlling equipment. Engineers could inspect disagreements with the existing process and decide whether further tests are justified. This staged approach lets a new supplier demonstrate value without asking the utility to accept a broad, undefined change in responsibility. The supplier’s research intelligence helps it prepare a serious proposal. Its evidence and integration determine whether the proposal belongs in an operating system.

The incumbent utility can benefit from a wider set of specialised contributors. Its own staff may have more capacity to investigate anomalous readings, review investment assumptions or examine unusual operating scenarios. A research agent can help organise the analysis, but the utility must decide which scenarios are relevant and what a failure would mean. Rare events deserve particular attention because a system can perform well on ordinary days and still be unsuitable for an important decision. More extensive testing is useful only if it includes the conditions that distinguish a dependable system from an attractive demonstration.

For households and businesses, the potential benefits are concrete: more reliable service, better use of existing assets, lower avoidable losses and more informed investment. The distribution of those benefits depends on the utility’s incentives and the arrangements under which costs are recovered. It should not be assumed that every efficiency gain immediately lowers a bill. Nor should the absence of an immediate price change imply that a reliability improvement has no value. The appropriate assessment follows the stated objective and asks whether the proposed change achieved it without creating an unacceptable cost elsewhere.

This industry also shows why faster intelligence and slower physical change can be complementary. Better analysis can help prioritise a construction programme that still takes time. It can identify a temporary operating improvement while a durable upgrade is prepared. It can expose that a proposed shortcut does not solve the relevant constraint. The value of research is sometimes to prevent an expensive acceleration in the wrong direction. An institution that adopts advanced intelligence well should become more precise about what must happen next, including the steps whose duration cannot be reduced by generating more recommendations.

### Software and technical services: the output is easier to make

Software appears to offer the shortest route from research to use because many artefacts and tests are digital. This makes it an important case, but not a universal proof of effortless productivity. Requirements can be unclear, tests incomplete and existing systems full of assumptions that are poorly documented. A change that compiles may still fail to meet the customer’s need. Technical services face a similar issue: producing a polished analysis is different from resolving the uncertainty for which the customer sought help. As output becomes easier to generate, buyers need better ways to evaluate whether the work is complete.

The evidence already warns against taking subjective speed as a sufficient measure. METR’s early-2025 randomised study of 16 experienced open-source developers and 246 tasks found that access to the studied AI tools increased completion time by 19 percent in that setting. Its February 2026 follow-up reported serious selection and measurement problems and judged the new data an unreliable estimate of the current effect size. The lesson is neither a permanent slowdown nor a proven universal speedup. It is that tools, participants, tasks and measurement design must remain attached to a productivity claim. [^S57][^S58]

The plausible disruptive future is still substantial. A small team can investigate a customer’s specialised workflow, build a prototype, prepare tests and revise it with less routine production labour. Work that was previously too small to interest a large supplier may become commercially viable. A business may obtain a tool adapted to its actual process instead of changing its process to fit a generic package. An independent consultant may bring stronger technical analysis to a narrow problem. The benefit depends on understanding the requirement and maintaining the result after delivery, not simply on how quickly the first version appears.

An incumbent software or services company faces a direct question about pricing. If part of the work becomes much cheaper, what justifies the bill? There may be good answers: dependable integration, maintained systems, difficult customer context or responsibility for a continuing service. Those contributions should be explicit. A pricing model that relies on customers being unable to obtain a competent first attempt may become vulnerable. The supplier can respond by charging for a verified outcome or a useful continuing relationship. It can also try to preserve old prices while lowering its own costs. Competition and customer bargaining power will influence which response survives.

The work of checking becomes more central. A good supplier can show what changed, which tests were performed, which assumptions remain and how the customer can recover if the change fails. A bad supplier can generate more code or more pages than the customer has time to inspect. Broad access to capable agents may increase both kinds of supply. The market’s quality will depend partly on whether buyers reward evidence and maintainability. The smallest team is not automatically the best team, just as the largest team is not automatically the safest choice.

Progress measures should include accepted changes, defects found after delivery, rework, maintenance burden and total cost. Elapsed time and staff effort should be recorded separately, especially when agents run while people do other work. A team may improve the value of its output without reducing every task’s duration. It may also appear faster because it defers necessary work to the customer. Measuring the complete service protects the constructive promise of the technology. It makes it possible to distinguish a genuinely more capable small supplier from an inexpensive source of unfinished work.

### The shared opportunity

> **Figure 30 · Five industrial futures (2D)** — Comparative scenarios. See the corresponding figure in the web edition.

Across all five industries, the small lab’s most credible first move is a bounded contribution with an identifiable buyer or beneficiary. Its intellectual reach can be broad while its commercial promise remains narrow. This is not a failure of ambition. It is how a new organisation earns access to the evidence and relationships required for larger work. The incumbent’s strongest response is to become an effective place to test and use good contributions, wherever they originate. A protected position can make refusal easy. It cannot make refusal a source of improvement.

> **Figure 31 · From an idea to an improvement (2D)** — Capacity scenario. See the corresponding figure in the web edition.

The visual above makes one constraint explicit. Increasing the rate at which candidates arrive does not necessarily increase the rate at which improvements reach users. If testing or deployment is the limiting stage, a larger inflow creates a queue. That queue can contain valuable opportunities, but it also consumes attention and can hide the best proposal among many plausible ones. The organisational response should improve selection, expand the relevant capacity or reduce unnecessary delays. Buying more generation capacity because generation is the easiest stage to buy may leave the final output unchanged.

This is where the future becomes a choice rather than a prediction. Oil and gas companies can use broader intelligence to investigate and respond more effectively. Biotech can learn more from a well-chosen experiment. Manufacturers can carry evidence from failure into better products. Utilities can connect analysis with dependable operation. Software and services firms can make useful work affordable to more customers. Each industry can also choose to protect access, count activity and leave the difficult handoffs untouched. The technology changes what is possible. The institutions determine how much of that possibility reaches the world.

## 13 · Where the clock will not accelerate

There are at least three kinds of delay in research. One is the time required to perform the intellectual work. Another is the time required for the world to produce an observation. The third is the time spent waiting for an organisation to permit, fund or use the work. Advanced AI can affect all three, but through different mechanisms. It can perform part of an analysis, improve the design of an experiment or make an administrative decision easier to assess. It cannot make these mechanisms interchangeable. A company that does not distinguish them will either expect impossible acceleration or defend unnecessary delay as if it were a law of nature.

The distinction matters because the language of caution can be used honestly or opportunistically. A researcher may need to observe a process long enough to answer the question. A manufacturer may need evidence that a part behaves consistently under relevant conditions. A committee may simply have no reason to meet sooner. All three can appear as a long bar on a project schedule. Only one can be removed by changing the committee’s calendar. The institution should be able to explain what is happening during each period and what information would be lost if the period were shortened.

### The unchanged part sets a limit

Consider a deliberately simplified project that takes one hundred units of elapsed time. Half is a serial analytical stage; half is a stage whose duration does not change. If the analysis becomes ten times faster, the project takes fifty-five units, not ten. If the analysis becomes instantaneous, fifty units remain. This is arithmetic, not a forecast for any industry. Its usefulness lies in exposing the missing assumption behind a common claim: a dramatic improvement in one part of a process becomes a dramatic improvement in the whole process only if that part occupies most of the relevant time or changes the other parts as well.

> **Figure 32 · The clocks that remain (2D)** — Elapsed-time scenario. See the corresponding figure in the web edition.

Real projects can be more favourable than this simple example. Better analysis may eliminate an unnecessary experiment, reveal that two stages can proceed together or identify an earlier observation that answers the question. They can also be less favourable. A larger flow of proposals may congest a shared facility. Faster design changes may create more integration work. A team may discover that the apparent analytical stage included unrecorded judgement that must now be performed during review. The right response is to map the actual dependencies and measure the complete process. The simple example is a prompt to investigate them.

A useful schedule therefore distinguishes effort from elapsed time. An analyst may spend four hours preparing an experiment that occupies an instrument for several days. A manager may spend ten minutes approving a request that waited two weeks in a queue. An agent may generate a candidate overnight while its supervisor works on another problem. Adding those durations as if they represented the same resource can produce a false productivity claim. Separate records make it possible to see whether a change saves staff effort, shortens delivery, increases capacity or merely moves work into another person’s queue.

### Observing the thing that matters

Some questions are about what happens over time. A system’s initial performance may not establish its durability. An intervention’s immediate effect may not establish its later consequences. A process that works once may not establish repeatability. A model can help identify early indicators, but the relation between an indicator and the outcome must itself be supported. Otherwise the organisation has shortened the test by changing the question without admitting it. The result can look faster because the difficult part of the claim has quietly disappeared from the evaluation.

This does not mean every conventional test duration is optimal. Historical practice can contain conservative assumptions that better evidence allows us to revise. More informative measurements may reveal a failure earlier. A well-designed accelerated test may help investigate a mechanism under specified assumptions. Research can improve research methods. But the burden is to show that the revised method supports the intended conclusion. A committee should not reject it merely because it differs from tradition. A supplier should not demand acceptance merely because the new method was produced by a more capable model. The test earns its place through the evidence it provides.

The same issue appears in mathematics in a different form. A machine-checkable artefact can make some verification tasks precise and reproducible. It does not eliminate the need to understand whether the formal statement corresponds to the intended mathematical claim. Nor does public release instantly complete independent scrutiny of a large result. The relevant clocks include computational checking, inspection of assumptions and the community’s ability to examine the work. They are not all equally reducible by adding more copies of the system that produced the original artefact. Diversity of scrutiny has a purpose distinct from volume.

### Parallel work needs a reason

Parallelism is powerful when the tasks are sufficiently separable. A group can investigate several candidate explanations, run independent checks or prepare different experiments at once. It is less useful when every task waits for the same missing measurement. A million agents cannot distribute the absence of that measurement into a million fragments and reconstruct it by agreement. They can propose ways to obtain it or reason about decisions under uncertainty. That may be valuable. It is still different from having the observation, and the project schedule should reflect the difference.

Parallel experiments can also compete for resources. They may need the same instrument, specialist or batch of material. They can create a larger volume of results than the team can interpret carefully. If the research system treats the number of simultaneous activities as its main objective, it can reduce the information obtained from each scarce resource. A better objective asks which combination of activities is most likely to resolve the important uncertainty within the available capacity. Sometimes that means running more tests. Sometimes it means waiting for one informative result before committing to a large batch of similar ones.

This is a place where advanced agents could be especially useful without pretending to be omnipotent. They can help compare proposed tests, identify dependencies, expose duplicated work and maintain a record of why a sequence was chosen. They can revisit a plan when new evidence arrives. Their value comes from making the allocation of effort more responsive to information. An organisation that gives them only permission to generate candidates has underused the capability. An organisation that gives them authority to schedule everything without a reliable account of constraints has misunderstood the capability in the opposite direction.

### The delay nobody owns

Many institutional delays occur between formally well-managed stages. The research team finishes its report. Another department must decide whether to test the recommendation. A facility needs a specification in a different format. A procurement process begins only after the technical decision, even though much of its preparation could have happened earlier. Each department can meet its own performance target while the overall project waits. The delay belongs to the relationship between departments, and nobody’s dashboard makes that relationship visible. A more capable analytical system can expose the problem without being authorised to solve it.

The remedy is not a demand that every decision be instantaneous. It is a named owner for each handoff, a clear input, a decision criterion and a reasonable time for response. If the recipient needs more evidence, it should say what evidence and why. If the proposal is outside the organisation’s priorities, it should receive a clear refusal. An indefinite request for another presentation wastes the proposer’s time and conceals the recipient’s decision. Institutions that want more useful research should make decisions legible, including negative ones. A prompt refusal can be more productive than a year of encouraging ambiguity.

There is a political difficulty here. Removing a delay can remove a person’s informal control over access. A process that looks inefficient from the perspective of the project may be useful to someone whose status depends on being consulted. AI does not create this conflict, but it can make the cost harder to hide. When the intellectual preparation becomes faster and cheaper, permission occupies a larger share of the timeline. The institution must then decide whether its procedures exist to improve decisions or to preserve the importance of those who administer them.

### Reversibility changes the appropriate pace

Not every experiment deserves the same approval process. A reversible analysis on a copy of a dataset differs from a change to live equipment. A simulation differs from a clinical intervention. An internal prototype differs from a promise made to a customer. The institution should identify the consequence of an error and match its review to that consequence. Treating all activity as equally consequential can waste scarce scrutiny on trivial work while leaving the genuinely important decision buried in a general queue. It also teaches staff to describe ordinary experiments in language that avoids attention rather than clarifies the risk.

A better process allows low-consequence exploration within explicit boundaries and reserves demanding review for consequential transitions. The boundary itself should be understandable: what data may be used, what systems may be changed, what evidence is required and who is responsible for the next step. This gives researchers room to investigate while making deployment a concrete decision. It can also make an outside collaboration easier, because the small lab knows what it can do without renegotiating the whole relationship. The organisation becomes faster through clearer distinctions, not through a general relaxation of standards.

The benefit is especially important when a model can produce many plausible suggestions. Without a disciplined distinction between exploration and use, teams may either freeze or deploy too casually. Exploration should be cheap enough to reveal mistakes early. Use should require the evidence appropriate to the consequence. A research system that preserves this boundary can be both ambitious and dependable. The ambition lies in how many serious questions it allows people to pursue. The dependability lies in how carefully it decides which answers are ready to matter outside the research environment.

### Speed that leaves a record

An accelerated process should remain reconstructable. A later investigator should be able to identify the question, the data available at the time, the alternatives considered, the tests performed and the reason for the decision. This record is not an administrative ornament. It enables correction when an assumption changes and prevents the same failure from consuming another team’s effort. If acceleration comes at the cost of losing the evidence behind decisions, the organisation may borrow time from its future self. The debt becomes visible when the original team leaves or the system behaves unexpectedly.

The record also makes institutional criticism fairer. It lets an outsider distinguish a necessary observation period from a queue, and a justified rejection from a reflexive refusal. It lets management see which stage is actually limiting useful output. It lets a small lab demonstrate the work behind a result without relying on reputation. These are forms of accountability that can support faster research because they reduce the need for vague reassurance. People can move with confidence when the basis of the decision is available for inspection.

The rapid future therefore does not require pretending that every clock can accelerate equally. It requires refusing to let the slowest unavoidable clock excuse every avoidable delay around it. A long experiment can coexist with rapid preparation, clear decisions and prompt interpretation. A difficult qualification can coexist with open access to preliminary tests. An institution can preserve the time needed for evidence while sharply reducing the time spent protecting its own routines. That is a more exacting demand than simply telling it to move faster, and a more useful one for the people who depend on its work.

## 14 · The market after intelligence gets cheaper

The market does not pay for intelligence in the abstract. It pays for products, services, rights, access and the expectation of future returns. If advanced research becomes cheaper, the consequences travel through those arrangements. A company may lower prices, increase its margin, attempt more projects or spend the saving on the next constraint. A new competitor may enter a market that previously could not support its fixed costs. A customer may obtain a better product at the same price. These are different outcomes, and the distinction matters. A scientific achievement does not contain a complete theory of who will capture its economic value.

The premise of generally available PhD-level intelligence is economically important because it could reduce the cost of making a serious attempt. It could also change the minimum size of the organisation capable of making that attempt. Neither change guarantees a profitable result. But both can alter the set of people who can credibly enter a negotiation. A buyer who previously had two plausible suppliers may encounter a specialised third. An investor who previously required a large permanent team may consider a smaller group with a strong evaluation process and access to experiments. The market begins to change before the incumbent disappears.

### A cheaper input is not an equally cheaper product

Start with a simple accounting example. A product sells for one hundred units and has an allocated cost of eighty. Suppose research represents a quarter of that cost, and a new method halves the research component. The saving is ten units. If the seller passes half of that saving to the customer, the price becomes ninety-five, the cost becomes seventy and the unit margin becomes twenty-five. Research became fifty percent cheaper; the price fell five percent. The example is invented and holds volume and quality fixed. It demonstrates why the share of cost matters before any claim about a final price.

> **Figure 33 · Who receives the savings (2D)** — Explicit economic scenario. See the corresponding figure in the web edition.

Even this illustration is more direct than many real businesses. Research often has a substantial fixed component: the cost is incurred before the company knows how many units it will sell. Dividing that cost across units is useful for accounting, but it does not make research identical to the marginal cost of producing one more unit. A reduction in future research expenditure may affect entry, the range of projects attempted or the return on a new product more than the price of an existing one. The company’s previous research expenditure may already be sunk. Prices do not mechanically rewind to refund it.

This distinction is particularly important when discussing medicines, industrial equipment and software. Their cost structures differ, and the commercial arrangements around them differ again. The essay’s accounting model is not a claim that each industry has the same research share or passes savings through at the same rate. It is a way to expose the assumptions required to make such a claim. Anyone promising an enormous price reduction from cheaper intelligence should identify the relevant cost, its share, the additional costs of using the result and the mechanism that causes the seller to pass the saving onward.

### Entry can matter more than the first price cut

The most consequential effect may be a lower threshold for entry. A specialised service that could not support a large research team may support a small one using rented analytical capability. A product for a narrow application may become worth investigating. A local problem may receive serious attention because the cost of preparing a credible attempt has fallen. These possibilities expand the set of projects that can be tried. They do not require every incumbent to reduce its prices immediately. They can change the alternatives available to customers and the range of needs that suppliers consider commercially addressable.

The arithmetic of a hypothetical small supplier makes the mechanism visible. If a business incurs a fixed development cost and earns a positive contribution on each sale after variable costs, reducing the fixed cost lowers the volume needed to recover that development expenditure, all else equal. But the qualification matters. A lab may replace salaries with substantial compute bills, contract experiments and specialist review. Its initial development cost can fall while its cost per attempt or per customer rises. The business must count those substitutions. Describing agents as an almost free workforce can hide the very expenses that determine whether the smaller organisation is viable.

Access to finance is another constraint. A technically capable group may still need funds before it can demonstrate its result. A buyer may pay only after qualification, while the lab must pay for experiments in advance. A small organisation can therefore be intellectually competitive and financially fragile. Better research tools may improve its proposal without resolving its cash needs. Public grants, customer-funded trials and appropriate commercial partnerships can matter because they bridge a specific gap between a credible attempt and evidence. The policy question is where modest support would allow a useful test that the market otherwise leaves unperformed.

Entry also depends on the buyer’s ability to evaluate a new supplier. If qualification requires an established brand, a long trading history or a scale unrelated to the actual risk, lower research costs may not open the market. Some requirements are justified by the consequences of failure. Others are convenient proxies that exclude a capable small team. A well-designed procurement process asks for evidence appropriate to the promised contribution. It can require insurance, continuity arrangements or independent tests where relevant without assuming that only a large headcount can supply them. The market becomes more contestable when competence has a route to recognition.

### Bargaining over the saving

Once an improvement exists, its economic value is negotiated. The model provider charges for access. The lab charges for its contribution. A facility charges for measurements. The incumbent combines the result with a product and sells it to a customer. If one participant controls a resource that the others cannot readily replace, it may capture a large share of the benefit. The research intelligence can be widely available while a decisive experimental platform, dataset or distribution channel remains concentrated. Lowering one barrier can therefore increase the importance of another.

For a small lab, this makes commercial design part of research strategy. It should understand whether its contribution can be evaluated independently, whether several customers could use it and whether it depends on a single provider’s continued cooperation. A result that is valuable only inside one company’s proprietary environment gives that company a strong negotiating position. A method that can be tested and applied across several settings may give the lab more options. These are not reasons to avoid partnerships. They are reasons to identify the dependency before the lab commits all its resources to a result whose only buyer controls the terms of recognition.

For an incumbent, buying outside intelligence can lower costs and improve products, but it may also make comparison easier. If the same specialist can serve several competitors, an advantage can spread. The incumbent then needs to distinguish itself through integration, service, customer understanding or a continuing ability to improve. This is an ordinary commercial challenge made more frequent by a wider supply of capable research. The company can respond by becoming better at those contributions. It can also try to prevent the supplier from serving anyone else. The choice affects how broadly the original scientific gain reaches the industry.

### Quality, variety and the customer’s time

Price is only one route through which customers benefit. A product can become more reliable. A service can answer a question that was previously too expensive to investigate. A design can be adapted to a neglected need. A research result can make an existing asset more useful. These changes can create value even when the invoice remains similar. The comparison should nevertheless be explicit. A seller should not invoke improved quality vaguely to explain why a cost reduction never appears in either price or a measurable customer outcome.

The customer’s own effort belongs in the calculation. A cheap tool that requires extensive supervision, correction and integration can be expensive to use. A more costly service can be economical if it reliably completes the work and reduces the burden on the buyer. The relevant comparison is total effort and consequence over the period that matters. This is especially important for agent-based services, where a low cost per generated output can coexist with a high cost of checking the output. A market that rewards only the visible generation price may initially encourage suppliers to transfer hidden work to their customers.

There is a constructive response: make the acceptance criterion part of the product. A supplier can define what a completed result includes, how the customer can test it and what support follows. This reduces uncertainty and makes a small organisation easier to buy from. It also creates a stronger comparison with an incumbent whose advantage rests partly on reputation. The customer can ask both suppliers to meet the same requirement. General access to intelligence becomes economically useful when it is accompanied by general access to credible ways of judging the resulting work.

### Employment and the distribution inside the firm

Lower research costs can change employment without eliminating the need for research. Some routine activities may require less staff time. Other activities, including experimental design, verification, integration and customer support, may become more important. A company can use the saving to pursue more projects, improve quality, reduce prices or reduce its workforce. The technology does not choose among these uses. Management, demand, competition and the organisation’s commitments influence the result. Describing every labour effect as either inevitable liberation or inevitable redundancy conceals those decisions.

The transition can be difficult even when total useful output grows. A person whose role centred on a now-cheaper activity may not automatically move into a newly valuable one. Training takes time, and the opportunity to learn depends on how work is organised. If entry-level tasks are removed without creating another route to judgement, an institution can weaken its future supply of experienced researchers. Conversely, agents can help learners attempt more substantial work earlier if supervision and assessment are designed well. The organisation needs a theory of how people develop competence, not just a spreadsheet of hours that might be saved.

The distribution of gains within a firm also matters. A team may produce more while receiving no additional time, pay or autonomy. A company may expect employees to supervise more work without recognising the cognitive burden of checking it. These are choices about work design. A useful measure of productivity should include the quality and sustainability of the process, not simply a temporary increase in output during an intense trial. An institution that uses advanced intelligence to exhaust the people responsible for validating it can damage the very source of reliability on which its commercial promise depends.

### Expectations move before factories do

Markets can react to a change in expected future competition long before an industry’s physical structure changes. A credible demonstration may cause customers to postpone a purchase, suppliers to revise investment or companies to change acquisition priorities. Those reactions can be rational even when widespread deployment remains years or uncertain steps away. They can also overshoot if the demonstration is mistaken for a complete operating system. The analysis must distinguish a change in expectations from a measured change in production, costs or customer outcomes. Both are real events, but they support different claims.

> **Figure 34 · The timing of market change (2D)** — Scenario comparison. See the corresponding figure in the web edition.

There is no defensible universal date on which an entire industry loses its protection. A digital component with a clear evaluator and easy integration can spread quickly. A physical product requiring new facilities and qualification can move differently. An incumbent may adopt the new method before a challenger reaches customers. A challenger may win a narrow segment while the incumbent’s overall business grows. These paths can coexist. The article’s claim is that more forms of competition become technically plausible and that the cost of complacency rises. It is not a timetable for the disappearance of named companies.

This matters for the tone of institutional criticism. A dramatic announcement is not evidence that every manager who asks for a test is obstructing progress. Nor is a slow deployment evidence that the scientific capability is commercially irrelevant. The useful question is what must happen between the result and the claimed economic effect. Name the missing test, investment, contract or adoption decision. Then ask whether the organisation is actively addressing it. This approach is more demanding than a declaration of inevitable disruption because it leaves institutions with specific work to do and gives observers a way to judge their response.

### More attempts can increase total spending

Cheaper attempts may encourage more attempts. A company that can investigate each candidate at lower cost may expand its research programme rather than reduce its budget. New entrants may add demand for compute, experiments and specialist services. A lower price for one activity can therefore coexist with higher total expenditure on the broader process. This is not a paradox. It is a change in quantity as well as cost. The same possibility applies to energy and physical resources: efficiency per task does not establish that total resource use falls when the number and scale of tasks change.

The quality of additional attempts matters. If the new projects address neglected but valuable questions, expanded spending can be productive. If they mostly duplicate one another or chase an evaluator’s weakness, it can be wasteful. Institutions should therefore track what the extra capacity makes possible, not just celebrate that more work is happening. The most useful evidence may be a question that could not previously be investigated, an uncertainty resolved earlier or a failed approach retired with a clear explanation. Those outcomes make the economic argument more concrete than a rising count of tokens, agents or proposals.

The market after cheaper intelligence will consequently be shaped by several simultaneous movements: lower costs for some intellectual tasks, greater demand for complementary assets, more possible entrants and renewed bargaining over who owns the result. Some established services may shrink or disappear. Some companies may become stronger by combining the new capability with assets they already use well. Some small labs may grow; others may supply a narrow contribution indefinitely. The common change is that institutional size and historical prestige become less sufficient as explanations of intellectual exclusivity. Economic value will have to be traced to a contribution that still makes a difference.

## 15 · The new owners of the bottleneck

An institution loses an exclusive advantage. A new provider makes the underlying capability available to millions of people. The story sounds like a transfer of power to those people, and sometimes it is. But another possibility deserves equal attention: the old gate opens onto a road owned by someone else. The researcher can enter, provided the account remains active, the price is affordable, the service is available and the provider permits the intended use. Access has expanded. Control may still be concentrated. An argument about intellectual freedom must examine both changes rather than treating a convenient interface as the end of institutional power.

The distinction begins with the word public. A publicly offered service, a publicly inspectable model, a publicly funded resource and a publicly traded company are different things. A stock-market listing would change the ownership and financing arrangements of a company; it would not by itself make its strongest research system universally available. A model offered through an API can be widely accessible without its weights or training process being open. A research artefact can be published while the configuration that produced it remains internal. The article’s premise concerns the spread of useful capability, and that spread must be assessed directly rather than inferred from any one of these meanings of public.

### The dependency moves

For a small lab, the practical research system may depend on a model provider, a cloud service, public or licensed data, software tools, experimental facilities and a route to customers. Each dependency has a price, a capacity and conditions of use. A change at one point can interrupt the whole process. The lab may be able to replace a model more easily than a specialised instrument, or vice versa. A serious strategy identifies which substitutions are realistic and what evidence would be needed to make them. Independence is a matter of available alternatives, not simply the absence of employees on another organisation’s payroll.

> **Figure 35 · The dependency moves (2D)** — Evidence-backed network. See the corresponding figure in the web edition.

The country lists discussed earlier illustrate a narrow but important boundary. They establish where a provider offers a particular service under stated conditions. They do not establish equal prices relative to income, equal network quality or equal access to an internal research configuration. A person can be formally eligible and practically unable to carry out a large research campaign. Another can afford the model but lack the means to test its proposals. General availability should therefore be measured as a sequence of usable opportunities. The meaningful question is whether a person can obtain a result of sufficient quality, check it and use it for the intended purpose.

This does not make commercial providers villains. Building and operating capable systems requires resources, and a reliable service can dramatically expand what users can do. The concern is structural. If a field comes to depend on a small number of providers whose terms are difficult to negotiate, those providers gain influence over cost, continuity and the direction of use. A society can celebrate the capabilities while still wanting alternatives, portability and independent evaluation. The same standard applied to universities and industrial incumbents should apply to the companies supplying the new intelligence: power should be explained through a contribution and constrained by credible options.

### Compute is a physical resource

The image of millions of agents can obscure the equipment behind them. Computation requires hardware, networking, storage and power, along with the systems that keep the equipment operating. The IEA’s 2025 analysis estimated that data centres consumed about 415 terawatt-hours of electricity in 2024. That figure covers data centres generally, not AI alone, and cannot be converted into an energy cost for an unspecified research agent. It nevertheless establishes the physical context: digital research capacity sits inside an infrastructure with real resource demands and local constraints. [^S61]

A lab that scales its agent population must therefore ask what useful work each additional unit of computation buys. Does it explore a new approach, check an existing result or repeat a correlated attempt? Does the extra computation save a scarce experiment or merely add candidates to a queue? Efficient models and better coordination can improve the answer. They do not make the question disappear. The opportunity to rent a large intellectual workforce should encourage disciplined allocation, especially when the proposed scale is many orders of magnitude larger than the permanent human team.

Resource concentration can also shape which research styles are practical. A group may have access to substantial batch computing but not a low-latency service, or access to an API but not the ability to inspect and modify the model. It may need to keep data in a particular environment. A method that works well under one arrangement may be inconvenient under another. These details influence participation in ways that a headline capability score does not capture. A genuinely broad research ecosystem needs several workable forms of access, because users have different constraints and the scientific questions do not all fit one commercial product.

### Shared resources are an institutional response

There are already examples of public institutions trying to widen access to the surrounding infrastructure. EuroHPC’s Playground access route is designed for eligible European small and medium-sized enterprises and startups, with a short application, technical assessment and available capacity. The US National Science Foundation’s NAIRR provides routes to computing, data, models, software and support for the US research and education community. These programmes have geographic and eligibility boundaries. They are not proof that anyone worldwide can obtain unlimited resources. They illustrate that access can be designed as a public function rather than left entirely to an individual’s ability to purchase a commercial service. [^S59][^S60]

The design of such programmes matters as much as the announcement. An allocation can be generous on paper and difficult to use without technical support. An application can require the institutional credentials that the programme was intended to make less decisive. A resource can be available for a short period that does not match the experiment’s needs. The useful assessment asks who applied, who could complete the process, what work was enabled and why others could not proceed. Measuring nominal capacity alone risks reproducing the same mistake as measuring an enterprise by the size of its R&D budget.

Shared experimental facilities deserve similar attention. If intellectual preparation becomes cheaper while instruments remain inaccessible, public investment in measurement can have a different value than another general-purpose model subscription. A facility can provide calibrated tests, competent operators, documentation and an independent record. It can help a small lab convert a plausible proposal into evidence that another organisation can trust. The bottleneck is specific, so the response should be specific. The aim is not to provide every applicant with every resource. It is to establish fair, understandable routes through which a serious question can receive an appropriate test.

### Openness has several dimensions

Open access to a paper is valuable, but it is only one part of reproducibility. A result may also require data, software, configuration details and documentation of the evaluation. A model’s weights may be available while its training data remain partly unknown. A dataset may be downloadable but poorly described. A proof artefact may be public while rebuilding it requires substantial computing resources. These distinctions do not make openness meaningless. They make it possible to identify what kind of openness has been achieved and what work remains before another person can inspect the claim independently.

The scientific examples in this essay demonstrate why that precision matters. Public coordinates allow a reader to examine a molecular structure. Published coefficients allow an algebraic identity to be checked. A released proof repository allows scrutiny that a company announcement alone cannot support. Each artefact opens a particular part of the process. None automatically answers every question about how the result was produced or how broadly the method performs. An institution that releases useful materials should receive credit for the contribution while still being asked to clarify the remaining limits.

For small labs, preserving their own reproducibility can also reduce dependence. Keep the data transformations, evaluation criteria and results in forms that can be inspected independently of the service that generated the first attempt. Record model versions and configurations where available. Test whether a second method reaches the same substantive conclusion. These practices do not guarantee easy substitution, but they make the cost of substitution visible. A lab that stores its entire scientific memory inside an opaque service may discover that access to its own reasoning has become a commercial dependency.

### The new institution can repeat the old mistake

The frontier AI companies deserve credit for the scientific work discussed here. They also deserve scrutiny when their claims move from a bounded result to a broad statement about research capability. A company can produce an important artefact and still present an incomplete denominator, an unclear comparison or an internal configuration that users cannot obtain. The appropriate response is neither reflexive dismissal nor deference to the brand. It is the same request made of every research institution: specify the claim, provide the relevant evidence and distinguish what has been demonstrated from what remains an expectation.

A smaller lab can repeat institutional arrogance too. It may mistake access to an advanced model for mastery of a field, treat criticism as proof that the establishment is threatened or hide failed attempts behind a successful demonstration. None of those behaviours becomes more scientific because the organisation is young. The challenge to old institutions is compelling only if the new entrants accept a demanding standard themselves. The right to attempt difficult work should be broad. The right to have a claim accepted still depends on evidence.

This is especially important when agent systems generate the appearance of a large intellectual community. Multiple reports can share one underlying mistake. A debate among model instances can be useful without being independent in the way several differently trained researchers or distinct experimental methods might be. The lab should describe the actual sources of variation and the external tests. Otherwise it risks recreating a closed room inside a computer: many voices, one set of assumptions, and an impressive consensus that has not met the world.

### Two possible distributions of power

> **Figure 36 · An open future, a concentrated future (2D)** — Explicit scenarios. See the corresponding figure in the web edition.

In one possible future, capable models are available through several channels, research artefacts are portable, public and commercial facilities offer workable access, and buyers can evaluate new suppliers. Small labs have real alternatives. Established institutions compete through the quality of their questions, tests and implementation. The benefits of scientific progress spread through more products, more useful public knowledge and a wider range of participants. This future is conditional on the surrounding arrangements. The models help make it possible; they do not build those arrangements by themselves.

In another possible future, many people can generate sophisticated proposals, but a small number of organisations control the strongest models, decisive data, experimental capacity and routes to customers. Outsiders can contribute, yet their bargaining position remains weak. The old institutions lose some exclusivity while new providers gain it. There may still be substantial scientific and economic benefits. The claim of intellectual emancipation would nevertheless be incomplete. Participation would have expanded at the stage of proposing work while control remained concentrated at the stages that determine which work becomes credible and useful.

Neither future is a probability estimate, and real systems will contain elements of both. The comparison identifies choices worth making visible. Can a researcher move their work? Can a small supplier obtain an independent test? Can a public facility explain its access decisions? Can a customer compare alternatives without surrendering sensitive information? Can a failed result be preserved and learned from? These questions are practical ways to assess whether the spread of capability is producing a broader distribution of effective agency. They are more informative than a simple count of accounts created or models downloaded.

### The right kind of institutional pressure

The strongest challenge to academia and enterprise is therefore not that they should disappear because an AI company has done something remarkable. It is that they should stop treating their control over access as sufficient proof of their value. The same challenge applies to the AI company if it becomes a new source of dependence. Every institution should be able to show what it contributes, how others can test that contribution and where its power comes from. This standard leaves room for large, expensive and specialised organisations. It removes the presumption that their current boundaries are beyond argument.

Public policy can support that standard through infrastructure, useful access rules, support for independent evaluation and attention to the practical costs of switching providers. Institutional leaders can support it through clear contracts, interoperable records and credible routes for outside contributions. Researchers can support it by making claims inspectable and acknowledging the work beneath their own. These are not glamorous additions to the story of scientific discovery. They are part of what determines whether discovery becomes a broad opportunity or a new concentration of privilege.

The premise that PhD-level intelligence can become generally available should remain ambitious. It should also remain unfinished until access includes the ability to do something consequential with the intelligence. A person needs more than an answer that sounds sophisticated. They need a route to evidence, a way to contest errors and a realistic opportunity to apply a result. The new providers can help build that route. So can the institutions whose authority they challenge. The question is whether each will accept a world in which more people are capable of asking for the route, understanding why it is blocked and proposing a better one.

## 16 · The institution worth rebuilding

The institution worth rebuilding begins with a different answer to a familiar question. When someone asks why the organisation should exist, it cannot stop at its history, its credentials or its control of access. It must identify the work it enables and the evidence that the work matters. A university can educate judgement, maintain shared knowledge and provide a place for difficult questions to meet criticism. An enterprise can connect research with reliable products and real customers. A public laboratory can make expensive measurements available under fair conditions. These are substantial purposes. None requires a presumption that serious intelligence must remain scarce outside the institution.

The practical response should be ambitious enough to change behaviour and narrow enough to evaluate. An organisation does not need to redesign itself in one announcement. It does need to stop using the scale of the transformation as an excuse to postpone every concrete experiment. Choose a real problem. Establish the current process. Give a team the means to attempt an improvement. Decide in advance what evidence would justify continuing, changing course or stopping. Then make the result available to the people who will decide what happens next. The discipline is ordinary. The new capability raises the cost of refusing to practise it.

### Begin with work that somebody needs

The first experiment should have a beneficiary who can say whether the result is useful. That might be an engineer facing recurring failures, a researcher trying to discriminate between hypotheses, a student learning to evaluate an argument or a customer waiting for a reliable software change. The problem should be specific enough that success does not depend on a generous interpretation after the fact. A demonstration chosen only because the model performs well on it teaches little about the institution’s actual work. The purpose is to discover where the new capability changes a decision the organisation genuinely has to make.

The baseline should include hidden effort. Record preparation, access requests, review, failed attempts and the time needed to put the result into use. Separate staff hours from elapsed time. Note what level of quality the existing process achieves and which failures matter. This can reveal that the old process is less well understood than its defenders assume. It can also reveal that the new process saves visible production time while adding substantial checking work. Either finding is useful. The institution should want the comparison to be informative, even when it disappoints the person who sponsored the experiment.

Choose an evaluation that the producing team cannot satisfy merely by changing its language. A proof has a stated claim. A diagnostic tool has a defined set of cases and relevant errors. A design has measurable properties under specified conditions. A service has an agreed deliverable and acceptance process. Some questions will require judgement rather than a single score, but judgement can still be explicit about its reasons. The aim is to prevent the trial from becoming a performance in which every outcome is described as promising. A system that cannot fail its evaluation cannot provide strong evidence of success.

### Give the decision a destination

Before the work begins, identify who can authorise the next step. If the result succeeds, is there capacity for replication, a pilot, procurement or further study? If not, the experiment may still be worthwhile as research, but it should not be advertised as an immediate route to implementation. Many organisations create promising trials without creating a destination for their results. The trial then becomes a recurring presentation rather than a transition. More capable agents can fill that holding area faster. They cannot make it useful without a decision process prepared to receive the work.

The destination should include a stopping rule. A project may fail because the method is inadequate, the data are unsuitable, the benefit is too small or the cost of integration is too high. These are different findings and should be recorded separately. A negative result can justify improving the data, choosing another method or abandoning the commercial proposal. It should not automatically justify a larger version of the same trial. The organisation needs a way to recognise that a particular use of AI did not earn its place without turning that conclusion into a judgement about every possible use.

Successful trials deserve equally careful interpretation. A result on one team, dataset or operating condition may not transfer automatically. Expand along a stated dimension and test the reason for expecting transfer. Perhaps the next team has different records, a less experienced supervisor or a more demanding deployment. The cost of supporting those differences belongs in the assessment. Scaling should mean that useful performance survives a broader set of conditions. It should not mean that the number of accounts or agents has increased while the original success remains the only evidence anyone can cite.

### Rebuild the route into research

For universities, one concrete change is a route through which an outsider can submit a bounded proposal for criticism or access to a test. The route can have eligibility conditions and capacity limits. It should make those conditions understandable and give applicants a clear decision. The institution should evaluate the quality of the question and preparation without assuming that an unfamiliar affiliation disqualifies the work. Advanced intelligence will allow more people to arrive with serious attempts. A university can become the place that helps distinguish those attempts, or the place whose first response is to defend the old entrance requirements.

Teaching should respond to the same change. If an explanation is easy to obtain, the course must help students determine whether it is correct, when it applies and how to use it on a new problem. Students should practise finding errors, comparing approaches and explaining decisions. They should encounter situations in which a polished answer is incomplete. This does not eliminate the need to learn foundational material. It changes how that learning can be demonstrated. A student who cannot recognise a basic mistake is poorly equipped to supervise a powerful research assistant, however fluent the assistant may sound.

Research training should also preserve the experience of uncertainty. A model can produce an answer quickly, but a researcher needs to understand what would change their mind. Students should learn to design a discriminating test, identify an assumption and interpret a failure. They should see that negative results can improve knowledge. They should be credited for making a claim clearer even when the claim becomes less dramatic. These habits are valuable precisely because output is becoming abundant. The university earns its role by developing people who can make that abundance more reliable and more useful.

### Make the enterprise permeable to good work

For enterprise R&D, a concrete change is to publish or circulate well-defined problems with a credible route to evaluation. Some can be offered externally; others can be opened across internal departments. The problem statement should identify the useful outcome, the available evidence and the constraints. It should avoid prescribing the familiar organisational form as part of the solution. A small team should be able to demonstrate a contribution without pretending to possess an entire corporate department. A larger team should be able to show why its additional resources improve the result.

Procurement should support this route with contracts proportionate to the contribution. A preliminary study, a component trial and a continuing operational service require different commitments. Treating each as a fully mature supplier relationship can exclude useful experiments. Treating each as an informal demonstration can leave important responsibilities undefined. The enterprise should make the transition between these stages clear. A small supplier needs to know how a successful test leads to paid work, what evidence the buyer will retain and which rights remain with the supplier. Ambiguity is a poor substitute for partnership.

Inside the company, reward people for using a good outside result as well as producing an internal one. An organisation that publicly celebrates collaboration while privately treating external contribution as a threat will struggle to benefit from wider intelligence. The researcher who identifies a superior method should receive credit for improving the outcome, even if the method makes a previous internal project less central. This is difficult because budgets and reputations are attached to existing work. It is also where the institution demonstrates whether its stated commitment to discovery extends to discoveries that disturb its own arrangements.

### Fund the evidence that the market neglects

Public funders can make a distinctive contribution by supporting the tests, records and shared tools that many projects need but few individual suppliers can profitably maintain. Independent replication, carefully documented negative results and accessible experimental infrastructure can improve an entire field. Their value may not appear in a single product’s revenue. A programme that funds only the most visible new model can leave these complementary needs underserved. The question should be what prevents credible work from becoming reliable knowledge, and whether a shared investment would remove that obstacle for many participants.

Access programmes should be evaluated through the work they enable and the people they reach. Count successful use, not just nominal allocations. Investigate why applicants fail to proceed. Is the problem technical support, preparation, cost, eligibility or the absence of a partner? These findings can guide a more useful intervention than simply expanding the same resource. A researcher with access to computing but no way to obtain a sample has a different problem from one who has a sample but cannot analyse it. General availability becomes practical through attention to those differences.

Funders should also preserve room for questions whose value is not yet easy to express as a near-term commercial outcome. The argument for broader research intelligence is partly that more people can investigate more kinds of questions. If every route to resources demands an immediately profitable application, the new capability may narrow the agenda around the preferences of existing buyers. Public research has a role in maintaining alternatives: basic questions, public-interest measurements and work whose benefits are widely shared. Institutions can become more accountable without requiring every contribution to resemble a startup pitch.

### The small lab’s obligations

A small lab should respond to the opportunity with a narrower promise and a stronger record than its size might lead a buyer to expect. State the problem, the method, the inputs and the evaluation. Preserve the failures needed to understand how the result was selected. Make clear which work was performed by agents, which tools were used and which conclusions received independent checks. This transparency is not an apology for using AI. It is part of making the research reproducible and allowing a customer or collaborator to understand the basis of the claim.

The lab should budget for the whole path to evidence. Compute is one cost. Human review, experimental access, preparation, integration and unsuccessful attempts can be others. A plan that spends everything on generating candidates may leave nothing for the test that gives them meaning. Scale the part of the system that improves useful output. If more agents only increase a queue, improve selection or obtain the missing capacity before expanding again. The ambition should be measured in questions answered and benefits delivered, not in the theatrical size of the digital workforce.

It should also seek criticism early. A domain expert who identifies a fatal assumption can save the lab months of work. A customer who rejects the proposed metric can reveal that the product addresses the wrong problem. A facility operator who questions the protocol can improve the experiment. Treating these responses as establishment hostility would reproduce the worst habits of the institutions the lab hopes to challenge. The new entrant earns credibility by being easier to correct, not by being more certain that its tools have made correction unnecessary.

### A record that makes change visible

> **Figure 37 · What to change next (2D)** — Decision instrument. See the corresponding figure in the web edition.

An organisation can keep a simple record for each serious research attempt: the question, the baseline, the resources used, the evidence obtained, the remaining uncertainty and the next decision. Across attempts, that record reveals patterns. Which questions receive attention? Which teams wait longest for access? Where do promising results stop? Which evaluations produce useful corrections? Which types of failure recur? This is not a request for another administrative reporting burden. It is a proposal to preserve the information already needed to judge whether the research process is working.

The record should be useful to researchers as well as managers. If it exists only to extract performance numbers, people will learn to optimise the numbers. If it helps them recover previous reasoning, find a relevant experiment and understand why a decision was made, it becomes part of the research infrastructure. Agents can assist in maintaining and searching it, but the categories must remain open to revision. A new kind of problem may not fit the old fields. The institution should be able to change its record when the work changes, rather than forcing the work to resemble the form.

This approach also changes how leaders communicate progress. They can describe a result that survived a demanding test, an unnecessary delay removed or an outsider whose proposal became useful evidence. They can explain a failure without pretending it was a success in disguise. They can identify the next constraint and the decision needed to address it. Such accounts may be less spectacular than announcing an enormous agent population. They are more convincing because they connect capability with consequence. Over time, they provide a basis for trusting the institution that does not depend on deference to its name.

The immediate challenge is therefore practical. Open one route that was closed without good reason. Test one claim that has been protected by reputation. Give one promising result a clear destination. Preserve one failure so the next team does not have to discover it again. These actions will not settle the future of academia or enterprise, but they begin to distinguish adaptation from commentary about adaptation. The institutions that make them will learn what the new intelligence can do. The institutions that refuse will continue to debate the future while making it harder for their own people to reach it.

The strongest institutions after this transition may be those that become least possessive about where intelligence originates and most demanding about what it achieves. They can welcome a student, a small lab, an internal team or a group of agents into the same process of question, evidence and use. Their authority will come from making that process work: fair access, exact claims, reliable tests and useful outcomes. That is an institution worth defending. It is also an institution capable of surviving the loss of an exclusive claim to the intelligence in the room.

## Who gets to try

Return to the room at the beginning of this essay. The people inside still have work to do. The equations remain difficult. The experiments still have to be performed. The machines must be built, calibrated and maintained. Patients need evidence, customers need useful products and students need to learn how to judge an answer. Nothing about the spread of advanced intelligence makes those obligations disappear. What changes is the claim that the room’s occupants are the only people capable of making a serious contribution to them.

Outside, the preparation becomes stronger. A person can arrive with an argument that has survived several kinds of criticism, code that can be inspected and a proposed experiment whose purpose is clear. A small lab can approach a company with a narrow result that deserves testing. A student can ask a question that previously required access to several specialists just to formulate. These are possibilities with different degrees of present support, not a claim that every person already enjoys them equally. They describe the direction that gives the premise its force: intellectual reach becomes less tightly coupled to institutional membership.

The scientific results matter because they make that direction harder to dismiss. They also demand exact language. The proposed Navier–Stokes result has a particular formulation and a process of scrutiny. Fermat’s existing proof and its formalisation are different historical achievements. A stronger bound concerning zeta zeros does not settle the Riemann hypothesis. Protein prediction, binding and clinical benefit remain distinct. Preserving those distinctions does not weaken the case for change. It prevents the case from depending on a collection of claims that collapse as soon as a knowledgeable reader examines them.

The institutional challenge survives every one of those distinctions. Even bounded achievements can change the cost of attempting difficult work, the scale of a team needed to attempt it and the confidence with which an outsider can request a test. An enterprise may retain a formidable advantage in manufacturing while losing exclusivity over a valuable analytical task. A university may remain essential to research while losing the ability to treat access to explanations as a sufficient educational offering. The change does not need to be total to be consequential. Partial changes can rearrange an institution’s purpose, income and authority.

The small lab with millions of agents is an especially sharp way to pose the question. It is not a literal equivalence between an agent and an independent scientist. It is a challenge to an old proxy: the belief that the size of the human department reliably measures the breadth of intellectual work an organisation can attempt. If a small group can explore many serious approaches, its limitations must be identified through the quality of those approaches and the evidence they obtain. Counting people will tell us less. Understanding the research system will tell us more.

That shift could help established companies as much as it threatens them. ASML can connect more capable analysis with measurements and manufacture. IBM’s shared methods can support new applications. Google can turn tested improvements into operating gains. Dyson and other engineering companies can explore more possibilities while holding products to demanding physical tests. Smaller competitors can find valuable contributions around and within those systems. The future can contain stronger incumbents, new suppliers, different prices and a wider range of products. It need not resemble a parade of corporate funerals to count as disruption.

It could also help people whose questions have received too little attention. A local technical problem, an unusual research hypothesis or a neglected customer need may become easier to investigate seriously. The gain is not merely that more people receive answers. It is that more people can participate in deciding which questions deserve an attempt. This is the most consequential democratic possibility in generally available advanced intelligence. It remains incomplete wherever access to testing, resources and recognition is reserved for the same narrow group. A wider right to ask must be connected to a workable route to evidence.

> **Figure 38 · Who gets to try (2D)** — Narrative synthesis. See the corresponding figure in the web edition.

The old institutions can help build that route. They have facilities, accumulated records, skilled people and responsibilities that the world still needs. They can make those resources more useful by allowing more good work to reach them. They can create standards that clarify what a claim must show without prescribing who is allowed to make it. They can fund criticism and replication, preserve failures and make decisions legible. These are ways to become more valuable as intelligence spreads. They require confidence in the institution’s contribution rather than anxiety about its exclusivity.

The new institutions face the same test. A model provider should not confuse widespread use with public control. A small lab should not confuse fluent output with discovery. A platform should not treat dependence as proof that its users have no need for alternatives. Every participant benefits from a research culture in which claims can be inspected and corrected. The future becomes less open if the collapse of one gate simply creates a more convenient gate elsewhere. The standard must follow power as power moves.

There will be reasons to proceed carefully, and there will be excuses that borrow the language of care. The difference is whether the delay produces evidence, protects a real responsibility or improves a decision. Institutions should be able to show that difference. Where a long observation is necessary, preserve it. Where a reversible test waits because nobody owns the decision, change the process. Where an outsider lacks a credential but has a serious proposal, examine the proposal. The pace of scientific change makes these distinctions more urgent. It does not make them less important.

My challenge to academia and enterprise is therefore direct: stop asking the world to accept your uniqueness as an inherited fact. Show the question you can formulate, the evidence you can obtain, the judgement you can develop and the useful result you can deliver. Show why your boundary helps the work. If it does, defend it with the evidence. If it does not, open it. The intelligence approaching the door will become increasingly capable of understanding the answer.

The premise is not that every person becomes a complete research institution. It is that access to difficult intellectual work can spread far beyond the institutions that once concentrated it. That possibility deserves more than excitement about impressive machines. It deserves changes in who can try, who can test, who can challenge and who can benefit. The room can remain. Its claim to contain all the minds that matter cannot.

## Sources and reading notes

Research cut-off: 9 September 2026. A source link records provenance, not independent replication. Some sources support figure data or context.

[^S01]: [OpenAI · A solution to the Navier–Stokes problem](https://openai.com/index/navier-stokes-solution/). 8 September 2026. Company announcement; read alongside the paper and formal artefact. Independent acceptance is a separate question.

[^S02]: [Clay Mathematics Institute · Millennium Prize rules](https://www.claymath.org/millennium-problems/rules/). The prize process requires qualifying publication, elapsed time and general acceptance. Announcement does not constitute an award.

[^S03]: [Anthropic · Formalizing Fermat’s Last Theorem](https://www.anthropic.com/research/formalizing-fermats-last-theorem). 4 September 2026. Formalisation of a theorem proved by Wiles in the 1990s, not its first proof.

[^S04]: [Anthropic · Progress on the Riemann zeta function](https://www.anthropic.com/research/riemann-zeta). 10 August 2026; updated 13 August. Reported improvement to a lower bound, not a proof of the Riemann hypothesis.

[^S05]: [Abramson et al. · Accurate structure prediction of biomolecular interactions with AlphaFold 3](https://www.nature.com/articles/s41586-024-07487-w). Nature, 2024. Predictions and benchmark performance are not clinical efficacy.

[^S06]: [Nobel Prize · Chemistry 2024](https://www.nobelprize.org/prizes/chemistry/2024/press-release/). Recognition of computational protein design and protein structure prediction.

[^S07]: [Watson et al. · De novo design of protein structure and function with RFdiffusion](https://www.nature.com/articles/s41586-023-06415-8). Nature, 2023. Generative design with experimental validation under stated conditions.

[^S08]: [Anthropic · Claude accelerates protein design](https://www.anthropic.com/research/Claude-accelerates-protein-design). 18 August 2026. Reported design and wet-lab results; retain denominators and unsuccessful targets.

[^S09]: [Anthropic · Expanding support for scientists](https://www.anthropic.com/news/expanding-support-for-scientists). 27 August 2026. Access programmes do not establish universal availability of internal research systems.

[^S10]: [Google DeepMind · Introducing AlphaEvolve](https://deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/). 14 May 2025. Company-reported algorithmic results and evaluation methods.

[^S11]: [Google DeepMind · AlphaEvolve’s impact](https://deepmind.google/blog/alphaevolve-impact/). 7 May 2026. Deployment claims require attention to baseline and operating context.

[^S12]: [Merchant et al. · Scaling deep learning for materials discovery](https://www.nature.com/articles/s41586-023-06735-9). Nature, 2023. Predicted stability, experimental synthesis and useful products are distinct achievements.

[^S13]: [Google Research · An AI co-scientist](https://research.google/blog/accelerating-scientific-breakthroughs-with-an-ai-co-scientist/). 19 February 2025. Includes hypothesis generation and experimentally assessed examples.

[^S14]: [International Energy Agency · AI for energy optimisation and innovation](https://www.iea.org/reports/energy-and-ai/ai-for-energy-optimisation-and-innovation). 2025. Oil and gas is already using AI; the industry is not an untouched technology enclave.

[^S15]: [ASML · Annual report 2025](https://www.asml.com/en/investors/annual-report/2025/financials). Company financial and operational disclosures. Existing investment and adoption matter to the competitive argument.

[^S16]: [IBM Research · Foundation models for materials](https://research.ibm.com/blog/foundation-models-for-materials). 20 December 2024. An example of incumbent participation in scientific AI.

[^S17]: [Dyson · Research engineering](https://careers.dyson.com/en-gb/what-you-can-do/engineer/research/). Company description of research functions. Not an independent measure of competitive performance.

[^S18]: [Kim et al. · Towards a science of scaling agent systems](https://arxiv.org/abs/2512.08296). Research on bounded agent systems; does not measure a million-agent scientific organisation.

[^S19]: [OpenAI · Navier–Stokes manuscript](https://cdn.openai.com/pdf/32d9f210-8b73-45e0-91bc-82a30aef8a9a/navier-stokes.pdf). 166-page mathematical manuscript. The stated forcing, initial conditions and regularity assumptions define the claim.

[^S20]: [OpenAI · NavierStokesAndEuler](https://github.com/openai/NavierStokesAndEuler). Released formal artefact. This essay does not represent an independent full rebuild of its certificates.

[^S21]: [Anthropic · Fermat formalisation repository](https://github.com/anthropics/fermats-last-theorem). Released proof code; the distinction between theorem statements, dependencies and kernel checking remains essential.

[^S22]: [Charles Fefferman · Existence and smoothness of the Navier–Stokes equation](https://www.claymath.org/wp-content/uploads/2022/06/navierstokes.pdf). Official Clay problem formulation. Its four alternatives have different assumptions about forcing.

[^S23]: [Andrew Wiles · Modular elliptic curves and Fermat’s Last Theorem](https://annals.math.princeton.edu/1995/141-3/p01). Annals of Mathematics, 1995. The original proof predates the AI formalisation by three decades.

[^S24]: [Lean · Language reference](https://lean-lang.org/doc/reference/latest/). The kernel checks proof terms; automation and the trusted core are different components.

[^S25]: [Lean · Validating a proof](https://lean-lang.org/doc/reference/latest/ValidatingProofs/). Explains statement validation, axioms, independent checking and remaining trust assumptions.

[^S26]: [Fermat formalisation · FinalCheck.lean](https://github.com/anthropics/fermats-last-theorem/blob/main/FinalCheck.lean). The final theorem and explicit axiom checks. Read, not independently rebuilt for this essay.

[^S27]: [Fermat formalisation · Attribution](https://github.com/anthropics/fermats-last-theorem/blob/main/ATTRIBUTION.md). Identifies reused material from Imperial College London’s FLT project, flt-regular and Mathlib.

[^S28]: [NIST DLMF · Zeros of the Riemann zeta function](https://dlmf.nist.gov/25.10). Definitions of the critical strip, critical line and hypothesis; finite numerical evidence is distinct from proof.

[^S29]: [Anthropic · Expert note on the zero-proportion result](https://www-cdn.anthropic.com/23455459f8832d06bb175cc0f88d019aed962ef8.pdf). Precise asymptotic statement and counting conventions; the constant is approximately 67.25 percent.

[^S30]: [Aryan · On an extension of the Landau–Gonek formula](https://arxiv.org/abs/1902.05473). Prior mathematical work cited by the 2026 result.

[^S31]: [Baluyot et al. · An unconditional Montgomery theorem](https://arxiv.org/abs/2306.04799). Prior work on pair correlation without assuming the Riemann hypothesis.

[^S32]: [Baluyot et al. · Proportions of simple zeros and critical zeros](https://arxiv.org/abs/2501.14545). 2025 contribution to the line of research used in the later result.

[^S33]: [Claude · Paper on the critical-zero proportion](https://www-cdn.anthropic.com/95c246936988e43127bc6b2ceb7077c1dad2d68e.pdf). Released research manuscript. The essay does not independently certify this proof.

[^S34]: [FDA · Clinical research in drug development](https://www.fda.gov/patients/drug-development-process/step-3-clinical-research). Clinical testing addresses effects in people; computational design is an earlier, distinct step.

[^S35]: [AlphaFold Protein Structure Database](https://alphafold.ebi.ac.uk). Public structure predictions, including the lysozyme example used here. Database access is not a clinical validation.

[^S36]: [Jumper et al. · Highly accurate protein structure prediction with AlphaFold](https://www.nature.com/articles/s41586-021-03819-2). Nature, 2021. Structure prediction evaluated against experimental structures.

[^S37]: [Anthropic · Autonomous de novo protein binder design](https://www-cdn.anthropic.com/30bf50e22a01388bb29bf077ee3f244531594b7a.pdf). Technical report with experimental arms and denominators. Binding does not establish therapeutic benefit.

[^S38]: [Anthropic · Automated processing of NMR and LC-MS data](https://www-cdn.anthropic.com/9f08da5189ac269b3242ca760de9823805c3f5f6.pdf/). A bounded analytical chemistry case. Reported agreement is not broad validation across instruments or compounds.

[^S39]: [RCSB PDB · Hen egg-white lysozyme, 1LYZ](https://www.rcsb.org/structure/1LYZ). Experimental coordinates used in the illustration. The corresponding AlphaFold prediction is P00698, mature residues 19–147.

[^S40]: [RCSB PDB · Streptavidin–biotin, 1STP](https://www.rcsb.org/structure/1STP). Experimental educational example of a binding interface. It is not a product of the AI campaign.

[^S41]: [Wang & Buehler · Self-Revising Discovery Systems for Science](https://arxiv.org/abs/2606.01444). Preprint, revised August 2026. A framework for representational change, not a universal impossibility theorem about AI contributions.

[^S42]: [Novikov et al. · AlphaEvolve: a coding agent for scientific and algorithmic discovery](https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/AlphaEvolve.pdf). 2025 technical paper. Algorithmic identities and operational performance are different claims.

[^S43]: [Crystallography Open Database · Diamond, 9008564](https://www.crystallography.net/cod/9008564.html). Wyckoff, Crystal Structures (1963), via AMCSD. Revision 291735; public-domain coordinates used for the unit-cell illustration.

[^S44]: [Google DeepMind · AlphaEvolve mathematical results](https://github.com/google-deepmind/alphaevolve_results/blob/master/mathematical_results.ipynb). Rank-48 complex matrix-multiplication coefficients inspected and all 4,096 tensor entries independently checked for this essay. No wall-clock benchmark claimed.

[^S45]: [ITU · Facts and Figures 2025](https://www.itu.int/en/mediacentre/Pages/PR-2025-11-17-Facts-and-Figures.aspx). Estimated six billion people online and 2.2 billion offline in 2025; connectivity is not equivalent to effective AI access.

[^S46]: [UNESCO · Recommendation on Open Science](https://www.unesco.org/en/legal-affairs/recommendation-open-science). 2021 framework addressing knowledge, infrastructure, participation and sustained investment.

[^S47]: [OpenAI · API supported countries and territories](https://help.openai.com/en/articles/5347006-openai-api-supported-countries-and-territories). Inspected 9 September 2026. Country listing does not establish access to internal research systems or equal affordability.

[^S48]: [Anthropic · Supported countries and regions](https://www.anthropic.com/supported-countries). Inspected 9 September 2026; API and Claude.ai lists are distinct. Additional eligibility conditions can apply.

[^S49]: [MIT OpenCourseWare · Get started](https://ocw.mit.edu/pages/get-started/). Free course materials without enrolment; using OCW does not confer MIT credit or a certificate.

[^S50]: [UK government · Independent review of research bureaucracy](https://www.gov.uk/government/publications/review-of-research-bureaucracy). Final report July 2022; government response February 2024. Addresses unnecessary burdens across the research system.

[^S51]: [CERN Open Data Portal](https://opendata.cern.ch/). Datasets, software, environments and documentation for education and research; an example of institutional public infrastructure.

[^S52]: [ASML · Computational lithography](https://www.asml.com/en/products/computational-lithography). Models calibrated with machine and wafer data connect computation with manufacturing. Inspected September 2026.

[^S53]: [ASML · Measuring accuracy](https://www.asml.com/en/technology/lithography-principles/measuring-accuracy). Public explanation of metrology, alignment and measurement feedback. The essay’s geometry is educational, not an ASML system model.

[^S54]: [Dyson · Motors and Power Systems](https://careers.dyson.com/en-gb/what-you-can-do/engineer/motors-and-power-systems/). Company description of electromagnetic, control, thermal, mechanical, data and manufacturing work. Not an independent performance evaluation.

[^S55]: [UNEP · AI-supported methane detection and mitigation](https://www.unep.org/news-and-stories/press-release/ai-helping-un-detect-methane-emissions-and-spark-real-reductions). 15 July 2026. MARS combines satellite observations, AI, analyst verification and notifications. Detection and subsequent mitigation are distinct stages.

[^S56]: [US Department of Energy · Smart transmission tools](https://www.energy.gov/cmei/systems/articles/smart-transmission-tools-modernize-americas-power-grid). 13 November 2025. Dynamic line rating connects weather and operational data with transmission limits. It is not itself evidence of an autonomous PhD-level agent.

[^S57]: [METR · Early-2025 AI and experienced developer productivity](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/). Randomised study of 16 developers and 246 tasks: 19% longer task times in this setting. Dated evidence, not an estimate for every tool or present-day workflow.

[^S58]: [METR · Changing the developer productivity experiment](https://metr.org/blog/2026-02-24-uplift-update/). 24 February 2026. Follow-up estimates affected by participant/task selection and time-measurement problems; authors consider the signal unreliable for the current effect size.

[^S59]: [EuroHPC JU · Playground access to AI Factories](https://www.eurohpc-ju.europa.eu/playground-access-ai-factories_en). Access route for eligible European SMEs and startups, with technical assessment and capacity limits. Not universal or unconditional access.

[^S60]: [US NSF · National Artificial Intelligence Research Resource](https://www.nsf.gov/focus-areas/ai/nairr). Shared computing, data, models, software and support for the US research and education community. Inspected September 2026.

[^S61]: [IEA · Energy demand from AI](https://www.iea.org/reports/energy-and-ai/energy-demand-from-ai). 2025 analysis estimates 2024 data-centre electricity at 415 TWh. This covers data centres, not AI alone; future values are scenarios.
