AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: AI’s Hidden Cost: Finding People To Check The Work on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

A report from ThorstenMeyerAI.com argues that AI is making it cheaper to produce mathematical manuscripts, software changes and contract work, while expert review remains slow and limited. The examples point to a growing verification bottleneck, though some figures come from companies that sell code-review tools and should be read with that interest in mind.

A report published this week argues that AI is increasing the supply of work in mathematics, software and contract analysis faster than expert reviewers can check it. Its central example is OpenAI’s publication of 722 mathematical manuscripts, set against the intensive human review of an earlier result from the same programme; the comparison illustrates a possible verification bottleneck, not a measure of review capacity across all fields.

According to the report, OpenAI’s model was given about 4,000 mathematical problems and produced 722 manuscripts across 372 families. Some results were formally checked using Lean, a proof-assistant system. OpenAI cautioned that some results without formal verification could have issues. The report says an earlier counterexample to an Erdős conjecture drew careful scrutiny from five leading mathematicians. These examples show different levels of checking; they do not establish that every manuscript received the same review or that all 722 results are correct.

The report also cites software-industry measurements. Faros AI found teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB’s analysis of 8.1 million pull requests across 4,800 organisations found AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. The source notes that some cited data providers sell code-review tools, a commercial interest readers should weigh when interpreting their findings.

A further example comes from OpenAI’s partnership with contract-software company Ironclad. The source says GPT-6 Astra was evaluated on 11 contracting tasks and met 55% of evaluation criteria on average, an improvement over the previous model. That result is a measure of performance on the stated evaluation, not evidence that the system can independently handle contracts. The remaining criteria illustrate why human review may still be needed before professional work is used.

At a glance
reportWhen: Published this week, according to the s…
The developmentA report has compiled examples from mathematics, software and contract work to argue that AI is increasing the volume of output faster than people can verify it.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Limits AI Output

If AI makes drafting or coding faster but does not make validation equally fast, organisations may gain more output without gaining a matching increase in usable work. Review time becomes a constraint: employees must decide which material to check, how deeply to check it, and who can approve it. That can delay useful work, allow unchecked errors through, or shift attention away from other responsibilities.

The report describes three risks: work may be merged with little or no review, reviewers may give AI-generated work lower priority based on suspicion, or the producer’s own quality filter may effectively replace independent scrutiny. The reported figures do not show that these outcomes are universal. They do raise a practical question for employers: who is accountable for approval when machine-generated output enters a professional workflow?

The argument also has implications for training. Experienced reviewers typically acquire judgement through years of doing the underlying work. If junior staff mainly oversee AI drafts rather than learning to draft, code or prove work themselves, organisations could weaken the future supply of people qualified to review it. The source presents that as a risk, not as an established outcome across industries.

Amazon

code review tools for software development

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evidence Across Three Workflows

The source’s examples cover three different kinds of verification. In mathematics, formal tools such as Lean can check whether a proof follows from its stated assumptions, while human specialists still judge the claim’s framing, importance and meaning. The source uses the phrase “verification abundance, adjudication scarcity” to distinguish automated checking from expert decisions about what a result establishes.

Software offers larger datasets, but the cited measurements have different methods and populations. Faros compares periods of lower and higher AI adoption; LinearB analyses pull requests across thousands of organisations. A peer-reviewed 2026 study cited by the source found 61% of AI-agent pull requests received no human review before being merged or closed. Faros also reported a 31.3% rise in merges with zero review during high-adoption periods. These are separate findings and should not be combined into a single industry-wide rate.

Professional contracting adds a question of responsibility. A system can be evaluated against criteria, but contracts still require people and institutions to decide whether terms suit a particular matter and who is answerable for the final document. Across all three examples, the source’s distinction is between producing an answer and establishing that it is fit for use.

Amazon

AI verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Much Review Is Missing?

The cited sources do not establish a single, comparable measure of the review burden across mathematics, software and contracting. The mathematical example concerns manuscripts from one programme, while the software figures come from different datasets and methods. Some providers cited in the source sell review-related products, and the material does not give enough detail to assess every study’s design or representativeness.

It is also unclear how much of the reported review delay reflects the quality of AI-generated work, reviewer workload, changes in team practices or other factors. The Ironclad evaluation’s task details and the full criteria are not provided here, so the 55% average should not be treated as a general measure of contract accuracy. The source’s concern about weakened training pathways is a forecast; it does not provide longitudinal evidence that the number of qualified reviewers is already falling.

Amazon

mathematical proof assistant software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Track Review and Training Data

The next useful evidence would show whether review time, error rates and human oversight change as AI use spreads, using clearly described samples and consistent definitions. Organisations adopting AI can compare output volume with the time spent reviewing it, while reporting how often work is approved, revised or rejected. Review rates alone are not enough without information about what reviewers checked and what errors were found.

For the examples in this report, the source provides no scheduled follow-up milestone. The open questions are whether formal verification expands beyond selected mathematical results, whether software teams can keep review quality aligned with faster output, and how employers will train junior staff to develop judgement. Until those questions are answered, the report’s central point remains a claim supported by several examples: more generated work can create more demand for human scrutiny.

Amazon

contract review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the report’s main finding?

It argues that AI is making it faster and cheaper to produce work in several fields, while human verification remains limited and can become a bottleneck.

Were all 722 mathematical manuscripts formally checked?

No. The source says some results were checked in Lean and reports OpenAI’s warning that unformalized results could have issues. It does not say that every manuscript received formal verification.

What did the cited software data show?

Faros and LinearB reported higher output or longer review waits in their respective datasets. A separate peer-reviewed 2026 study cited by the source found 61% of AI-agent pull requests received no human review before merging or closing. These figures come from distinct studies and are not one universal rate.

Does the report prove AI-generated work is less reliable?

No single conclusion applies across the examples. LinearB reported lower acceptance for AI-generated changes in its dataset, but that measure alone does not establish why they were accepted less often or the overall reliability of AI work.

Why could AI affect the supply of reviewers?

The report argues that junior workers often learn judgement by doing the work they later review. If AI replaces too much of that practice, employers may need to protect opportunities for hands-on training. The source presents this as a risk, not a confirmed workforce trend.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Gulf: Own the Capital

Thorsten Meyer AI report says Gulf states are using sovereign funds to buy into AI as labor disruption concerns grow.

Who Owns the Robots? The Future of Capital in an Automated Economy

I wonder how the concentration of robot ownership by big corporations will reshape our economy and society in the coming years.

Ubi 2.0: Lessons From Global Cash Pilots

Discover how global cash pilots reveal key lessons for designing effective, scalable UBI programs that overcome political hurdles and meet regional needs.

The license. Why the AI content market pays the brand-name corpus and strands the long tail.

Exploring how licensing practices in AI training data favor major brands, sidelining smaller content sources and impacting the long tail of data providers.