# Research Verification Log

**Course:** Introduction to AI
**Verification pass:** 13 August 2026
**Against:** `docs/production-bible-intro-to-ai.md` (the Claude Chat course design document)

This log records what happened when every load-bearing claim in the production bible was checked
against its source. It exists so that a future maintainer can see which claims were confirmed, which
were corrected, and which were added — without having to redo the work.

Three categories:

- **CONFIRMED** — the bible's claim survived checking unchanged.
- **CORRECTED** — the claim was wrong, imprecise, or missing context that changes its meaning.
- **ADDED** — material not in the bible, introduced because the claim it supports was previously
  asserted without a source, or because a primary source existed behind a secondary one.

---

## 1. Corrections

### 1.1 The Stanford employment figure has a revision history — and the bible quoted only the latest

**Bible:** "Employment of workers aged 22–25 in AI-exposed occupations stands roughly 19% below where
it would be had it tracked their less-exposed peers."

**Verified:** Correct — for the version of *Canaries in the Coal Mine?* revised **12 August 2026**,
using ADP payroll data through June 2026. But the same paper reported approximately **13%** in its
original 2025 working-paper form and approximately **16%** in an intermediate revision.

**Why it matters:** All three numbers circulate as "the Stanford study." A viewer who encounters the
13% figure elsewhere will conclude the course is wrong, and a course that cites 19% without a version
date is making the same error it warns against.

**Action:** Lesson 11 now teaches the revision history as content (§11.2), including a table of the
three figures, and states a citation rule requiring the version date. This was the single most
valuable correction in the pass.

### 1.2 IMF Staff Discussion Note — wrong title

**Bible:** *New Jobs Creation in the AI Age*, SDN/2026/001.

**Verified:** The actual title is ***Bridging Skill Gaps for the Future: New Jobs Creation in the AI
Age***. The bible cited the subtitle alone. Authors: Florence Jaumotte, Jaden Kim, David Koll, Elmer
Li, Longji Li, Giovanni Melina, Alina Song, and Marina Mendes Tavares. Published January 2026.

**Action:** Corrected in Lesson 11 with authors added, and flagged in the endnote as a note that is
"frequently miscited by its subtitle alone."

### 1.3 *Three Lenses on the AI Revolution* — no author in the bible

**Bible:** cited as a bare arXiv number (2510.12859).

**Verified:** Author is **Masoud Makrehchi**. v1 submitted 14 October 2025; v2 13 December 2025.
It is a **preprint** and is not peer reviewed.

**Action:** Author and version dates added in Lesson 12; preprint status flagged in the endnote, as
it is for every preprint in the course.

### 1.4 MCP donation — founding members misattributed

**Bible:** "Anthropic donated MCP to the Agentic AI Foundation under the Linux Foundation, with
OpenAI, Google, Microsoft, AWS, and Block among founding members."

**Verified:** The Agentic AI Foundation was announced **9 December 2025** and **co-founded by
Anthropic, Block, and OpenAI**, with **support from** Google, Microsoft, AWS, Cloudflare, and
Bloomberg. Founding contributions were three projects: Anthropic's MCP, Block's `goose`, and OpenAI's
`AGENTS.md`. Google, Microsoft, and AWS are supporters, not co-founders.

**Action:** Corrected in Lesson 5, with the three donated projects named.

### 1.5 MCP ecosystem counts — a single number where the sources give five

**Bible:** "more than 10,000 MCP servers had been published and SDK downloads passed 97 million."

**Verified:** Both figures appear in the sources, but they are not comparable and neither is a simple
total.

| Measurement | Value | As of | Counting |
| --- | --- | --- | --- |
| Official MCP Registry API | 9,652 latest-server records | 24 May 2026 | Registry entries only |
| PulseMCP directory | 22,311 servers | 16 Jul 2026 | Third-party listing |
| GitHub `mcp-server` topic | 15,926 repositories | 24 May 2026 | Repos, not servers |
| Anthropic | 10,000+ active public servers | 2026 | Active public only |
| SDK downloads | 97M+ **monthly** | Dec 2025 | Python + TypeScript |
| SDK downloads | 150M+ cumulative | 2026 | Reported during 2026 |

None include private enterprise servers.

**Action:** Lesson 5 now presents the divergence itself as the teaching point, with a callout
instructing the reader to state what was counted and when. The 97M figure is labelled **monthly**,
which the bible did not.

### 1.6 Common Sense Media survey — presented one-sidedly

**Bible:** cited 72% ever-used, 52% regular users, and the unacceptable-risk finding.

**Verified:** The headline figures are correct (n = 1,060 U.S. teens aged 13–17, fielded April–May
2025, published 16 July 2025). The regular-use figure is reported by the publisher as "about half" —
the course uses ≈50% rather than 52% to avoid false precision.

**But the bible omitted the countervailing findings from the same survey:** roughly half of teens
distrust AI advice, about 80% say they prioritise real friendships, and trust is markedly *higher*
among younger teens than older ones.

**Why it matters:** Presenting only the alarming half of a survey is the failure mode the course
warns against in Lesson 7 and Lesson 11. It also weakens the argument — the age gradient in trust is
the most actionable finding in the report.

**Action:** Lesson 4 now presents both halves, and the misconception section kills the paired
misconception ("companion products are simply predatory") alongside the intended one.

### 1.7 Dartmouth workshop duration — the sources disagree

**Bible:** "The workshop ran for eight weeks in the summer of 1956."

**Verified:** Accounts differ. It is "usually said to have run for six weeks"; Ray Solomonoff's
contemporaneous notes indicate roughly eight; some institutional summaries say five. Participants
attended intermittently rather than for a fixed term, which is the likeliest source of the
discrepancy.

**Action:** Lesson 1 reports the disagreement rather than picking a number, and uses it as the
lesson's first demonstration of the course's own citation discipline.

### 1.8 Lesson 7 — declined to assert a current model roster

**Bible:** described the 2026 market structure without naming models.

**Verified — and this is an original finding of the pass:** published leaderboards for August 2026
**disagree with one another about which models exist**, listing different version numbers for the
same vendor's flagship within the same month.

**Action:** Lesson 7 asserts only the structural claims (compression at the top, benchmark
saturation, routing as the norm), which are well supported and stable across sources. The
leaderboard disagreement is reported in §7.2 as a methodological lesson, and the self-check turns it
into an exercise. No current model roster is stated, because no reliable one was found.

---

## 2. Confirmed without change

| Claim | Source checked | Result |
| --- | --- | --- |
| Dartmouth founding conjecture, exact wording | Dartmouth College; reference record | Confirmed verbatim |
| Proposal submitted 2 September 1955 by McCarthy, Minsky, Rochester, Shannon | Reference record | Confirmed |
| Gartner: 40% of enterprise apps with task-specific agents by end-2026, from <5% | Gartner press release, 26 Aug 2025 | Confirmed — and it is a **forecast** |
| Character.AI / Google settlement, January 2026 | CNN Business, 7 Jan 2026 | Confirmed; detail added (see §3.9) |
| Hugging Face: 13M users, >2M models, >500K datasets | HF *State of Open Source*, 17 Mar 2026 | Confirmed |
| Hugging Face: top 200 models = 49.6% of downloads | Same | Confirmed |
| Kalai et al. on hallucination incentives | arXiv:2509.04664 | Confirmed, including the "epidemic of penalizing uncertain responses" characterisation |
| MCP released November 2024 as an open standard | Anthropic / MCP blog | Confirmed |
| South Africa draft AI policy, cabinet approval 25 March 2026 | Policy record | Confirmed |
| Three Korean models trending on Hugging Face by February 2026 | HF Spring 2026 report | Confirmed |
| Claude Code: agentic tool; CLAUDE.md read every session; MCP integration | Claude Code documentation | Confirmed verbatim |
| ChatGPT release as index date in labour research | Stanford Digital Economy Lab | Confirmed |

---

## 3. Additions

Material introduced during this pass. Most of it replaces a bible claim that was asserted without a
source, or promotes a secondary citation to the primary work behind it.

### 3.1 METR randomised controlled trial — the largest single addition

Not present in the bible. A randomised controlled trial (arXiv:2507.09089, 10 July 2025) found 16
experienced open-source developers were **19% slower** on 246 real tasks with AI tools available,
while estimating afterwards that they had been **20% faster** — a 39-point perception gap.

This is now load-bearing in **Lesson 10** (§10.4) and **Lesson 14** (§14.1). It is the strongest
available evidence against the course's own subject matter, which is precisely why it belongs.

**Caveats carried with it, all three:** the sample is narrow and close to the worst case for the
tools; the tools studied are those of February–June 2025; and **METR itself labels the result
historical**. Reporting the finding without that last caveat would be as dishonest as suppressing it.

### 3.2 Learning science — the empirical basis for the course's own design

Not present in the bible, which asserted that "reading about AI does not work" without support.

- **Roediger & Karpicke (2006)**, both papers — the testing effect, and the **illusion of
  competence**: most students prefer re-reading to self-testing, a preference that does not track
  effectiveness.
- **Ericsson, Krampe & Tesch-Römer (1993)** — the components of deliberate practice: immediate
  feedback, time for evaluation, repeated performance.

These now underpin **Lesson 14**'s four-step loop and, more consequentially, the **self-check
questions in every lesson SPA** — which are a retrieval-practice instrument rather than decoration.
The parallel between the METR perception gap and the illusion of competence is flagged in Lesson 10's
endnote 4 as this course's observation, not the authors'.

The contested effect size of deliberate practice is disclosed in Lesson 14's bibliography via the
Macnamara & Maitra (2019) replication, rather than hidden.

### 3.3 Primary sources promoted from secondary citations

The bible cited these through consultancy and think-tank summaries. The primary works are now cited
directly, with the secondary sources retained where they add application.

| Claim | Bible cited | Now cited |
| --- | --- | --- |
| Factory re-architecture and the productivity lag | KPMG | **David, *AER* 80(2), 1990** — "The Dynamo and the Computer" |
| Engels' Pause | KPMG | **Allen, *Explorations in Economic History* 46(4), 2009** — with the figures: output per worker +46%, real wages +12%, 1780–1840 |
| Productivity paradox / intangibles | AEI | **Brynjolfsson, Rock & Syverson, *AEJ: Macro* 13(1), 2021** — the productivity J-curve, with TFP 15.9% above official measures by end-2017 |

### 3.4 Technical primary sources for Lesson 3

The bible asserted "a token is roughly ¾ of a word" and described the context window without
citation.

- **OpenAI tokenisation documentation** — ≈4 characters or ≈¾ of a word per token in English.
- **Sennrich, Haddow & Birch (2016)** — subword tokenisation, the ancestor of current schemes.
- **Liu et al., "Lost in the Middle," *TACL* 12 (2024)** — retrieval accuracy is highest at the
  beginning and end of a long context and degrades in the middle. This turns "long chats drift" from
  an assertion into a documented effect and produces a concrete instruction: put what matters at the
  start or the end.
- **Wei et al. (2022)** — chain-of-thought prompting.
- **Bender et al., "Stochastic Parrots," FAccT (2021)** — cited so the "does it understand anything"
  disagreement is *represented* rather than resolved. The course takes no position on the philosophy
  and a firm position on the engineering.

### 3.5 Lewis et al. (2020) — the RAG paper

The bible described the grounding pipeline in Lesson 15 without citing its origin.
**Lewis et al., *NeurIPS* 33 (2020), 9459–9474** is now cited in both Lesson 6 and Lesson 15.

### 3.6 Anthropic, "Building Effective AI Agents"

Not in the bible. Introduces the **workflow vs. agent** distinction and the recommendation to find
the simplest solution possible — "which might mean not building agentic systems at all."

This materially improves Lesson 5, which now closes on a decision rule (§5.4) rather than on the
failure modes. It is the most practically useful addition to Movement II.

### 3.7 Lesson 1 historical apparatus

The bible's Lesson 1 leaned on a single secondary chronology. Added as primary sources: McCulloch &
Pitts (1943), Rosenblatt (1958), Minsky & Papert (1969), Rumelhart, Hinton & Williams (1986),
Newell & Simon (1956), McCarthy on Lisp (1960), Weizenbaum on ELIZA (1966), and the Lighthill report
(1973) with its actual argument — **combinatorial explosion**, and the recommendation against funding
"Category B."

Also added: the Rockefeller Foundation funding detail (**$13,500 requested, $7,500 awarded**), which
usefully punctures the myth of a well-funded moonshot; and the framing of symbolic AI and
connectionism as **two rival programmes** rather than one lineage, which the bible did not make
explicit and without which Lesson 2's milestone table does not cohere.

### 3.8 Exact figures where the bible said "dramatically"

- **AlexNet, ILSVRC-2012:** top-5 error **15.3%** against **26.2%** for second place.
- **Deep Blue, 1997:** won the rematch **3½–2½**, having lost the 1996 match 4–2; ≈200 million
  positions per second.
- **GPT-3:** 175 billion parameters.

### 3.9 Litigation detail

The bible noted the Character.AI/Google settlement. Added from CNN Business (7 January 2026): the
settlement covered the founders as well as the companies, included *Garcia v. Character
Technologies* concerning the death of 14-year-old Sewell Setzer III, plus further cases in New York,
Colorado, and Texas; terms were **confidential with no admission of liability**; and **more suits
have been filed since**. The last two points materially change what the settlement can be said to
establish.

### 3.10 Hugging Face structural findings

The bible used the platform totals and the concentration figure. Added from the same report: China
overtaking the US in downloads with **41%** of the total; Baidu going from zero releases to 100+
during 2025; industry's share of model development falling from ~70% (pre-2022) to **37%** (2025)
while independent developers rose from 17% to **39%**; and robotics datasets growing from 1,145 to
**26,991**.

The producer shift is the most consequential of these for a practitioner — it means provenance and
maintenance guarantees are weakening as production disperses — and it is now in Lesson 8's self-check.

### 3.11 Sovereign AI — institutional corroboration

The bible's Lesson 9 rested largely on trade sources. Added: **Carnegie Endowment, "Early Lessons in
the Pursuit of Sovereign AI" (June 2026)** as institutional corroboration of the layer framing, and
an explicit caution that announced commitments are not spend and that market projections to 2040 are
marketing.

---

## 4. Standing cautions for future maintainers

1. **`Canaries in the Coal Mine?` is under active revision.** Check for a newer version before every
   use. Never cite it without the version date. The figure has moved three times.
2. **Benchmark leaderboards are vendor-reported and mutually inconsistent.** They disagreed about
   which models existed during this pass. Do not treat agreement between two aggregators as
   corroboration — they frequently copy one another.
3. **MCP ecosystem counts differ by an order of magnitude by measurement basis.** Always state what
   was counted and when.
4. **Anthropic's labour research is first-party.** Its methodological caution is a genuine
   contribution *and* it favours a benign reading of a question in which the publisher has a
   commercial interest. Lesson 11 states both; keep it that way.
5. **Five sources in the course are preprints**, not peer-reviewed: arXiv 2509.04664, 2507.09089,
   2507.16078, 2510.12859, 2506.10281. Each is flagged in its endnote. If any is published in a
   peer-reviewed venue, upgrade the citation and the tier badge.
6. **Lesson 15's product specifics are deliberately unwritten.** Everything marked `[PRODUCT: …]`
   must be filled against the current Webspinner build. Nothing in that file describes shipped
   functionality.

---

## 5. Method

Every claim was checked against a live source during the pass. Where a bible claim traced to a
secondary summary, the primary work was located and read in preference. Where sources disagreed, the
disagreement is reported in the lesson rather than resolved silently — Lesson 1 §1.1 (workshop
duration), Lesson 5 §5.2 (ecosystem counts), Lesson 7 §7.2 (leaderboards), and Lesson 11 §11.2
(revision history) are all instances.

Claims that could not be verified were removed rather than softened. No figure appears in the course
that was not traced to a named source on a stated date.
