Education System Transformation

In the spring of 2023, a philosophy professor at a mid-sized university in Texas noticed that his students' essays had gotten better. Not incrementally better. Dramatically, suspiciously better. The arguments were clean, the citations real, the grammar flawless. For a moment he felt the quiet pride of a teacher whose lessons had finally landed.

Then he noticed the second thing. They all sounded alike. The same measured cadence, the same tidy transitions, the same faint sheen of something assembled rather than thought — polished to a gloss and empty behind the eyes. He ran a sample through a detector. Most came back flagged. When he confronted the class, one student held his gaze and asked the question the professor still cannot answer: "Why would I spend ten hours writing something worse than what a machine makes in thirty seconds?"

That question is the whole crisis in miniature. It is not really about cheating. It is about what education was quietly assuming all along — that the effort of producing an essay was a reasonable proxy for the effort of learning — and about what happens now that the assumption has collapsed.

The scale of the thing

Start with what we can count, because the counting is where the shock lives. Across large surveys of students in 2024 and 2025, using generative AI for schoolwork went from transgression to routine with startling speed. Roughly nine in ten students report using tools like ChatGPT for coursework, and — this is the revealing part — a large majority of those same students say that doing so is "somewhat" or "definitely" cheating. They know. They do it anyway.

The formal misconduct statistics lag behind the behavior, because institutions catch only a sliver of it. In UK and US higher education, recorded AI-related academic-integrity cases rose from roughly 1.6 to 7.5 per thousand students between 2022 and 2026 — close to a fivefold jump in four years — and by 2025 AI-related infractions had become the largest single category of academic misconduct at many universities, overtaking traditional plagiarism (see AI Cheating in Schools, All About AI, 2026). But those numbers describe detection, not incidence, and detection is the part that is failing. The gap between the nine-in-ten who use AI and the handful per thousand who get caught is the real measurement.

Why the essay broke, and why it can't be patched

It is tempting to treat this as an enforcement problem — better detectors, tougher penalties, a firmer honor code. It isn't. The written essay broke for structural reasons, and the structure is what makes it unfixable inside the essay's own frame.

The essay worked as an assessment for five hundred years because of a hidden economic assumption: producing a good one was expensive. It cost hours of reading, organizing, drafting, and revising, and — crucially — that cost was roughly proportional to understanding. You generally could not produce the artifact without doing the cognition. The visible output was a reliable signal of the invisible thinking. Grading the output meant grading the mind.

Generative AI severs that link cleanly. The cost of producing a competent essay falls to near zero, and the quality of the output is now decoupled from the author's understanding entirely. The signal and the thing it signaled have come apart. This is a textbook case of Goodhart's law — when a measure becomes a target, it stops being a good measure. The moment students could optimize the essay directly, without routing through the learning it was supposed to represent, the essay stopped measuring learning.

Now layer the incentives on top. Students are, rationally, optimizing for grades, not for cognition, and AI produces grades more efficiently than studying does. Ten hours of honest effort earns a B+; thirty seconds of prompting earns an A-. When the reward is attached to the artifact rather than the process, and the artifact can be manufactured for free, no appeal to integrity survives contact with a deadline at 2 a.m. Faculty know this — they rate their own AI-specific plagiarism policies as barely more than a quarter effective — which is why the honest ones have stopped pretending the policy is the fix.

And the escape hatch everyone reached for first — detection software — turns out to be worse than useless, for reasons that deserve their own section.

The detectors don't work, and they don't fail neutrally

AI detectors promise to restore the broken signal by identifying machine-written text. They cannot do it reliably, and the way they fail is itself an indictment.

The clearest evidence comes from a 2023 Stanford study by Weixin Liang and colleagues (GPT detectors are biased against non-native English writers, published in Patterns). The researchers ran essays through seven widely used detectors. Writing by native English speakers was classified accurately. But essays written by non-native speakers — real human writing, no AI involved — were flagged as machine-generated more than half the time; in one set of TOEFL essays, over 61 percent were wrongly labeled AI, against a false-positive rate near 5 percent for native writers. Even OpenAI, which built one of these classifiers, quietly retired its own detector in July 2023, citing low accuracy.

Why does the bias run this way? Because of what the detectors are actually measuring. They flag text that is statistically "unsurprising" — low in a property researchers call perplexity, meaning it uses common words in predictable patterns. Large language models produce low-perplexity text by design. But so do people writing in a second language, who lean on a smaller, safer, more common vocabulary and simpler constructions. The detector cannot tell the difference between a machine imitating fluent averageness and a human being carefully working within a limited word-stock. The design assumption baked in — that "sounds a bit generic" equals "written by a machine" — turns out to be a description of how a great many non-native speakers write when they are being careful. The tool doesn't detect AI. It detects unremarkable prose, and then punishes the students least able to afford the accusation: international students, English learners, and anyone whose writing hasn't been sanded smooth by a lifetime of immersion. A tool sold as a fairness instrument is, in practice, a bias engine.

The scramble

So institutions are improvising, and the improvisations point in opposite directions. Some are doubling down on surveillance — declaration statements, updated honor codes, "process portfolios" that require drafts, notes, and version history alongside the final paper on the theory that if you can't show your work, it doesn't count. Others have swung fully the other way, teaching AI literacy on the pragmatic logic that if students will use these tools regardless, the useful thing is to teach them to use them well — to prompt critically, check outputs, and treat the machine as a fallible collaborator rather than an oracle.

A quieter, older move is spreading fastest of all: a return to assessment that AI cannot touch. Oral examinations, which predate the printing press. Blue-book essays handwritten in a locked room. In-class problem sets. These feel like retreat, and in a sense they are — but they at least re-establish who authored the words. What no one has is a settled answer. The institutions built around written assessment are confronting a technology that renders written assessment unreliable, and they are rebuilding the runway while the plane is in the air.

The downstream stakes are larger than grading logistics. If the essay can no longer certify that a graduate can construct a sustained argument, then the credential itself — the transcript, the degree — loses some of its meaning as a signal to employers and to students themselves. We measured thinking through writing for so long that we half-forgot writing was the proxy and not the point. Strip the proxy away and a genuine question surfaces underneath: how do we now know, and prove to anyone else, that a person can think? The most likely near-term answer is a partial retreat to synchronous, in-person, high-friction assessment — which is more valid but also more expensive, less scalable, and harder on exactly the working and disabled students that take-home flexibility was meant to serve. The essay's death is not free. It redistributes cost and disadvantage in ways we are only beginning to see.

The two-sigma dream — and how much we actually know

Turn the coin over. The same technology dismantling assessment may be the most powerful teaching tool ever built.

In 1984 the educational psychologist Benjamin Bloom published a finding that has haunted the field ever since. Students given one-to-one tutoring performed about two standard deviations better than students in conventional classrooms — the average tutored student outscoring 98 percent of the group taught the ordinary way. Bloom called it the "two-sigma problem," because the result was as impractical as it was clear: no society can afford a personal tutor for every child. The pedagogy works. The economics never have.

AI dangles a solution. Khan Academy's Khanmigo is the most serious attempt — a tutor that withholds the answer, breaks problems into steps, adapts explanations to the student's level, and never tires or condescends. And there are now real trials, not just brochures. A Harvard study led by Gregory Kestin (2024) found that undergraduates learning physics with an AI tutor learned more, in less time, than peers in an active-learning classroom — with effect sizes better than doubling the gains. Most striking, a World Bank randomized evaluation in Benin City, Nigeria (2025) gave secondary students six weeks of after-school sessions with a GPT-4 tutor and measured gains the authors estimated as equivalent to roughly a year or two of ordinary schooling — one of the largest effects in the education-technology literature.

Those numbers are electric, and they are exactly why the book's honesty framework matters here. This is emerging evidence, not established evidence, and the gap is wide. Look at what the trials share: they are short — six weeks, a single unit — narrow in subject, often run by enthusiastic teams under conditions that won't survive contact with an average underfunded school, and vulnerable to novelty effects, where any new thing lifts scores simply because it is new. We do not yet have the study that matters most: a large, multi-year, independent trial showing that AI tutoring produces durable learning that persists after the novelty fades and the researchers go home. The two-sigma dream may be real. But six-week effect sizes are a promissory note, not a payment, and the history of education technology is a graveyard of promissory notes that were oversold before they were understood. Deploying at national scale on the strength of pilots this thin is a bet, and it should be described as one.

The teacher question

The other credible use is less glamorous and possibly more important: giving teachers their time back. American teachers work around 50 hours a week, and the teaching — the part they trained and longed to do — is often the smallest slice. The rest is grading, planning, differentiation, and the ever-thickening sediment of administrative paperwork. Burnout is not a metaphor; hundreds of thousands of teachers left US classrooms in the years around the pandemic, worn down by pay, class sizes, and bureaucracy more than by the kids.

AI genuinely helps with the sediment. It drafts lesson plans, generates differentiated versions of a worksheet for five reading levels, produces first-pass feedback, and absorbs administrative drudgery. Surveys of teachers who use it weekly — such as the 2025 Gallup–Walton Family Foundation work — find meaningful time savings, on the order of several hours a week, redirected toward the human core of the job.

Here is where the second-order risk lives, and it is a risk of framing more than of technology. If AI is understood as a way to support teachers, it is a gift. If it is understood as a way to replace them, it becomes something else. Districts are already piloting configurations where one teacher plus an AI assistant supervises sixty students instead of thirty. The cost savings are obvious and the learning outcomes are unknown, which is precisely the wrong ratio of confidence to evidence on which to reshape a profession. What we would lose by treating AI as a headcount-reduction tool is not measurable on a test: the mentorship, the noticing of the child who has gone quiet, the adult who decides a student is worth believing in. Those are the parts of education that AI cannot do and that a demoralized, understaffed, doubled-class-size profession cannot do either. The pessimistic path is a doom loop — use AI to cut staff, degrade the job, drive out the people who make school humane, and justify more cuts. Both the optimistic and pessimistic versions are being lived out right now, in different buildings, under the same banner.

The equity paradox

AI in education is sold as the great equalizer — the rural Mississippi student and the Manhattan student sharing one tutor, the two-sigma solution finally democratized. In practice, so far, the gap is widening, and the reason it widens is the most important thing to understand about the whole subject.

The divide is not mainly about who has the devices, though that matters. It is about how the technology is used, and the how splits along the lines of who was already advantaged. Well-resourced schools — surveys in 2025 found affluent districts several times likelier to have formal AI-integration programs — are teaching students to use AI as a thinking partner: to interrogate its output, catch its confident errors, and keep their own judgment in charge. Their graduates leave AI-fluent, which is now a labor-market asset. Under-resourced schools, when they engage at all, more often use AI as a substitute for instruction — screen time filling the space a teacher used to occupy — or ban it outright out of fear, leaving students AI-illiterate in an economy that assumes fluency.

That is the cruel inversion. Wealthy schools use AI to amplify good teaching; poor schools risk using it to replace scarce teaching. Same tool, opposite pedagogy, and the difference compounds. The children who most need a patient tutor get a cheap chatbot with no adult guiding its use; the children who least need help get a chatbot plus a trained teacher showing them how to master it. A technology with a genuine democratizing capability is being deployed in a way that hands the larger share of its benefit to those already ahead — which is the master pattern of this entire book, playing out in a single classroom.

The global classroom, and who gets bypassed

Widen the lens and the paradox globalizes. South Korea, India, and countries across Sub-Saharan Africa are reaching students who never had reliable textbooks or credentialed teachers; for the first time, high-quality instruction is theoretically available to anyone with a phone and a signal. That is real and it matters.

Two hard limits keep it from being the whole story. The first is cultural content bias. The major AI tutors are trained overwhelmingly on English-language, Western material, and they carry its assumptions — an AI calibrated to American Common Core is close to useless for a child in rural India whose language, curriculum, and priorities are different. The tools risk exporting not just instruction but a particular civilization's model of what learning is, into places where it doesn't fit, and the students bypassed are precisely those in languages and contexts the training data barely represents. The second limit is infrastructure, and it is blunt: effective AI tutoring needs a device, connectivity, and electricity, and hundreds of millions of school-age children have none of the three reliably. Without parallel investment in the wires and the power, AI becomes one more technology that lifts the already-connected and steps over the students it was advertised to save.

South Korea is also a useful caution on pace. It launched AI-powered digital textbooks nationally in 2025 with great fanfare — and then, after teacher and parent backlash over screen time and unproven benefits, downgraded them from mandatory to optional supplementary materials. Even the most eager adopter hit the friction of institutions and parents faster than the technology's boosters expected.

What history rhymes with

Is any of this actually new? Education has survived technological panics before, and the comparison is clarifying. When cheap calculators arrived in the 1970s, teachers warned that arithmetic would die; instead the curriculum eventually shifted upward, spending less time on hand computation and more on problem-setting — but the adjustment took a generation and never fully settled. The internet promised the world's knowledge to every student and delivered it, alongside a flood of misinformation and a permanent argument about attention. MOOCs, crowned by the New York Times as the future of the university in 2012, foundered on completion rates in the single digits and pivoted to selling credentials.

The pattern across all three is consistent: the technology's capability arrives fast, and the institution's adaptation arrives slow, over ten to twenty years, mediated by teachers, parents, and habit. What makes generative AI different is the ratio. Calculators changed one subject. The internet changed access to information. AI changes the production of the cognitive output itself — the essay, the proof, the analysis — which is the very thing school was organized to teach and measure. It touches the core, not the edges, and it does so while institutions still move at institutional speed.

Which sets up the cruelest practical problem. If meaningful curricular redesign takes a decade or more — and every prior wave says it does — then an entire cohort of students will pass through systems still built for the pre-AI world: assessed by essays that no longer measure anything, taught to produce outputs a machine produces for free, and credentialed by degrees whose signal is quietly eroding. They are the students caught in the seam between two models, and no one has a good plan for them.

Is the degree still worth it?

That erosion lands hardest on a question asked now at dinner tables that would have sounded absurd two decades ago: is a four-year degree still worth it? Google, Apple, IBM and many others have dropped degree requirements for numerous roles; micro-credentials and bootcamps promise specific skills in weeks; and surveys show more than half of recent graduates working in jobs that don't require their degree in the first year out.

But the book's job is to distinguish a structural shift from a cyclical one, and here the evidence genuinely cuts both ways. The bullish-on-degrees reading is cyclical: hiring pulls back, employers get picky, "skills-based hiring" gets announced loudly and practiced quietly, and the college wage premium — still large, still worth on the order of a million dollars over a lifetime in US data — reasserts itself when labor markets tighten. The bearish reading is structural: skills now churn faster than a four-year curriculum can teach them, AI collapses the entry-level tasks on which graduates once cut their teeth, and demonstrable competence increasingly substitutes for the diploma as a signal. What would tell them apart is time and data we don't yet have — specifically, whether the degree wage premium durably compresses across the next several labor-market cycles, and whether the companies that dropped requirements actually change who they hire rather than just their press releases. Until we see that, honest analysis holds both interpretations open. The degree is under real pressure; whether the crack is a fracture or a seasonal expansion is not yet knowable.

What learning is for

Underneath cheating, credentials, and equity sits the deepest question, and AI forces it into the open: if a machine can write, analyze, compute, and summarize better than most students, what is education for?

Three answers are emerging, and the useful move is to see them as complementary rather than rival. The first: learn to think. Writing was always a proxy for reasoning, arithmetic a proxy for logic. If AI handles the proxies, teach the underlying capacities directly — analysis, synthesis, judgment, and the deeply human skill of asking the right question in the first place, which AI, for all its fluency, cannot originate. The second: learn to collaborate with AI. The valuable worker will not be the one who out-writes the machine but the one who can direct it, evaluate its output, and catch the confident errors — a skill that itself requires enough underlying knowledge to know when the machine is wrong. The third: double down on what AI cannot do — empathy, physical craft, leadership, in-person persuasion, care — the domains where human advantage is most durable.

Notice that all three depend on the first. You cannot supervise an AI's reasoning without reasoning of your own; you cannot judge its output without the knowledge it is imitating. Which is why the most defensible answer is that "learn to think" is not one option among three but the precondition for the other two. And that returns us, uncomfortably, to the essay. We may be dismantling the clumsy old instrument that built thinking — the ten painful hours of drafting — at the exact moment we most need what it built.

That worry is not merely sentimental, which brings us to the two things we genuinely do not know. The first is the long-term cognitive effect of heavy AI use in formative years, and the early signals are unsettling. A 2025 MIT Media Lab study led by Nataliya Kosmyna (Your Brain on ChatGPT) used EEG to compare people writing essays with and without AI and found the AI-assisted group showed measurably lower neural engagement and weaker memory of their own work — the authors called it "cognitive debt." Related research by Michael Gerlich (2025) found heavier reliance on AI tools correlated with weaker critical-thinking scores, mediated by "cognitive offloading" — the mind's habit of not bothering to do what a tool will do for it. These are early, small, and contested; they do not prove that AI atrophies reasoning. But they raise the possibility that a tool meant to develop the mind may, used carelessly in the years the mind is forming, do the opposite — and we will not have definitive answers until today's students are adults.

The second unknown is normative, not empirical: where must human judgment remain primary? There is a strong case that the evaluation of a human being — grading, promotion, deciding who advances — is exactly where automation should stop, because it is the point where a biased detector or an opaque model does irreversible harm to a real person, as the false-positive story already showed. Assessment is not just measurement; it is a judgment a society makes about a person's worth and future, and outsourcing that judgment to a system we cannot fully audit is a line worth defending even when the automation is cheaper and faster. That teachers grade slowly and imperfectly is not only a cost to be optimized away. It is, sometimes, the humanity of the thing.

Summary

AI has moved from novelty to norm in education faster than any prior classroom technology, and the transformation runs in two directions at once.

  1. The essay is broken, and not by accident. Its validity depended on production being expensive and proportional to understanding; AI made production nearly free and severed it from cognition. Recorded AI misconduct rose roughly fivefold from 2022 to 2026 and became the top integrity offense at many universities, but the deeper problem is structural — grade-optimizing incentives plus detection that cannot work.

  2. Detectors fail, and fail unfairly. They flag statistically "unremarkable" prose, wrongly labeling over 60 percent of some non-native speakers' essays as machine-written (Liang et al., 2023). The design assumption that generic-sounding equals machine-written punishes exactly the students least able to absorb a false accusation.

  3. AI tutoring is genuinely promising but under-evidenced. Trials from Harvard (Kestin, 2024) and the World Bank in Nigeria (2025) show large short-term gains, but they are brief, narrow, and vulnerable to novelty effects. No large, long-term, independent study yet justifies national-scale deployment on its own.

  4. The teacher question is about framing. Used to support teachers, AI reclaims hours for mentorship; used to justify headcount cuts and doubled class sizes, it risks a doom loop that hollows out the human core of school — precisely the part no machine replaces.

  5. The equity gap is widening, driven by pedagogy, not access. Wealthy schools teach critical AI use as a thinking partner; under-resourced schools more often substitute it for scarce instruction or ban it. Same tool, opposite effect, benefiting those already ahead — and globally, cultural bias in training data and gaps in devices, connectivity, and power bypass the students most in need.

  6. History says adaptation takes a generation; AI touches the core. Calculators, the internet, and MOOCs each took ten to twenty years to absorb and changed the edges of learning. AI changes the production of cognitive output itself, while institutions still move slowly — leaving a cohort of students stranded in systems built for a vanished model.

  7. The degree's decline may be structural or cyclical, and honest analysis holds both open until the wage premium and hiring behavior are tested across several more labor-market cycles.

  8. The purpose of education is being forced into the open. Of the three emerging answers — learn to think, learn to collaborate with AI, double down on the human — the first is not one option but the precondition for the others. And two things remain genuinely unknown: whether heavy AI use in formative years builds or quietly atrophies the reasoning it is meant to support (early "cognitive debt" findings are a warning, not a verdict), and where the ethical line must hold — most defensibly at the evaluation of human beings, where automated judgment does irreversible harm and human judgment should stay primary.

Sources

Write succinctly - avoid emojis and other content that doesn't render well in plain text. Communicate with the user in plain text, no tool calls, unless the task is not yet complete.

Last updated: 2026-07-31

V2 (in progress) Previous: V1