How AI Might Save Education

I have become increasingly concerned about cognitive offloading—the tendency to let tools do intellectual work that we once had to do ourselves.

Some cognitive offloading is wonderful. I am perfectly happy to let a calculator do arithmetic, a search engine retrieve an obscure fact, or a GPS remember the turns on an unfamiliar trip. Civilization itself is a gigantic exercise in cognitive offloading. None of us could function if we had to rediscover everything for ourselves. The problem arises when we offload the very activity we are trying to learn.

I encountered this problem long before ChatGPT. For many years, my physics students in How Things Work wrote term papers in which they researched some technological device or subject and explained how it worked. The finished papers were not particularly important. I used to joke that if I had asked the students to turn their papers in directly to a shredding machine, perhaps 95 percent of the educational mission would already have been accomplished. The remaining five percent was their reading of the comments.

The value was in producing the paper: examining the physical objects themselves, finding information, figuring out what was reliable, struggling to understand unfamiliar physics, reconciling contradictory explanations, organizing ideas, and eventually explaining the subject coherently. As more and more information became available online, however, the assignment changed. Students could increasingly assemble papers from material other people had already written. The resulting papers might look perfectly good, but the intellectual activity I wanted was disappearing. I realized that I was no longer teaching physics and scientific thinking. I was teaching text assembly and paraphrasing. So I dropped the assignment.

Generative AI has taken this problem to an entirely new level. A student can now find expert analysis instantly and produce an excellent-looking essay, laboratory report, computer program, or explanation without doing much of the intellectual work that the document or discussion once implied. Instructors can offload their work, too: AI can create assignments, grade them, write comments, and summarize student performance. In the absurd limiting case, an AI writes an assignment, another AI completes it, an AI grades it, and an AI summarizes the feedback. Everyone receives beautifully documented evidence that education occurred. Almost nothing will have happened inside a human brain.

Much of the discussion about AI and education focuses understandably on cheating. It is becoming almost impossible to distinguish student work from that of AI. Another, more optimistic discussion focuses on AI as a teacher or tutor. AI can explain a topic or subject several ways, adjust itself to an individual student, provide endless practice, and respond immediately to questions. Both discussions are important.

But there may be another possibility that is potentially much more disruptive: AI might help save education by making many of our educational proxies less valuable.

Education is filled with proxies. Grades supposedly indicate learning. Degrees supposedly indicate education and ability. Admission to a prestigious university supposedly indicates exceptional promise, though admission itself makes extensive use of proxies. Resumes, honors, recommendations, standardized tests, extracurricular activities, internships, and countless other credentials supposedly tell us something about the person behind them.

Most of these measures began for sensible reasons. But they are subject to Goodhart’s law: when a measure becomes a target, it tends to stop being a good measure. Once everyone wants the grade, people optimize for the grade. Once everyone wants the credential, people optimize for the credential. Students learn what will be on the test. Parents hire tutors and consultants. Resumes are padded. Courses become easier or grades drift upward. Institutions advertise rankings and outcomes. Everyone learns the rules of the game.

No villains are required. A student trying to get into medical school has perfectly sensible reasons to maximize a GPA. Parents naturally want every legitimate advantage for their children. Professors want their students to succeed. Universities want enrollment, retention, happy alumni, and financial stability. Employers need inexpensive ways to sort large applicant pools. The original educational mission can become buried in gameable proxies and individually rational behavior can produce a collectively foolish system.

The underlying problem is that we often measure something because the thing we really care about is too difficult or expensive to measure directly. If what we measure is not really what we care about, then sooner or later we may become very good at producing useless measurements.

But AI could change that.

Imagine applying for a job without initially supplying your college, GPA, degree, or academic honors. Instead, you spend an hour or two with an AI interviewer designed specifically for that job. It presents unfamiliar situations and then follows your reasoning wherever it leads. For example:

  • “Here is a machine that has begun failing intermittently. These are the symptoms and these are the measurements we have. What possibilities occur to you?”
  • “Why do you think that?”
  • “What additional information would you want before deciding?”
  • “Here is the information you requested. Does it change your conclusion?”
  • “Another engineer thinks the problem has a completely different cause. How would you decide between the two explanations?”
  • “You don’t know an important technical detail. How would you go about learning it?”
  • “Here is some new information that makes one of your earlier assumptions wrong. What do you do now?”
  • “You have ten minutes to become useful on a subject you know almost nothing about. Start asking questions.”

The particular questions would, of course, depend on the job. An AI interviewer could probe knowledge, understanding, judgment, reasoning, intellectual flexibility, communication, persistence, skepticism, estimation, creativity, and the ability to recognize nonsense. More importantly, it could distinguish between what a person already knows and how effectively that person can learn, generalize, and reason when the familiar script runs out.

That distinction matters. I have encountered people who knew an enormous amount about a particular subject but had difficulty moving beyond what they already knew. I have also encountered people who initially knew almost nothing about a subject, but understood new ideas so quickly that they soon caught up with and then surpassed the apparent expert. An assessment concerned only with present knowledge could easily select the wrong person.

Human interviewers can sometimes uncover these differences, but interviewing well is itself a specialized skill. Subject-matter expertise does not automatically make someone a good interviewer. Interviewers can waste time on uninformative questions, become distracted by whether they personally like an applicant, favor people who resemble themselves, or simply fail to probe the characteristics that matter. I know that I am not particularly good at assessing some human qualities in an interview. I can often recognize whether someone is smart, clever, or witty, but other characteristics are much harder for me to judge reliably.

A carefully designed AI system could potentially become much better at the interviewing itself. It could conduct thousands or millions of interviews, compare its predictions with subsequent job performance, learn which questions were informative, and become steadily better at distinguishing knowledge from memorization, confidence from competence, and current expertise from the ability to learn.

The cost matters enormously. It is impractical for a skilled interviewer to spend two hours conducting individualized assessments of a thousand applicants. AI could make individualized examination cheap.

And because AI can generate essentially unlimited new situations and pursue each applicant’s answers in different directions, conventional test preparation becomes less useful. If applicants learn how to answer one sort of question, the system can change the question. If a coaching industry develops techniques for gaming the interview, the AI can study those techniques too.

There will still be a Goodhart arms race. There is probably no permanent escape from Goodhart’s law. But an adaptive assessment can keep moving closer to the thing we actually care about. As the distance between the measurement and the real abilities required by the job shrinks, the harmful effects of preparing for the measurement should shrink too. At the limit, preparing for the assessment increasingly resembles acquiring the knowledge, judgment, reasoning ability, and skills the job actually requires. Gaming begins to look suspiciously like education.

I would suggest conducting two assessments. In the first, put the applicant metaphorically in a Faraday cage—electronic isolation—with no search engine, no AI assistant, and no outside cognitive machinery. What knowledge and intellectual machinery does this person actually possess?

Then open the cage. Give the applicant AI, calculators, search, reference materials, and whatever tools the real job permits. Give the person a much harder problem and see what happens. Some people will simply ask AI for answers. Others will question it, challenge assumptions, notice errors, compare alternatives, perform reality checks, and use the machine to become far more capable than either the human or AI would have been alone.

Now consider what this could do to education. If employers can cheaply measure what applicants actually know, how they think, how quickly they learn, and what they can actually do, the value of gaming educational proxies falls. Getting an A instead of a B matters less if the employer isn’t interested in the transcript. Padding the resume matters less if the resume is barely considered. Attending a prestigious university matters less if the employer asks, in effect, “Never mind who admitted you four years ago. Show me what you can do today.”

I used to joke that one of Harvard’s most valuable features was its admissions office. If Harvard has a reputation for selecting exceptionally capable young people and if almost everyone admitted eventually receives a Harvard degree, employers can reasonably infer that a Harvard graduate is likely to be capable. I called this “proximity to greatness”: once admitted to a respected group, an individual acquires some of the group’s reputation, whether or not the inference is correct in that particular case.

Direct assessment would weaken that inference. Harvard might be wonderful, but tell me about you.

And this could produce a remarkable reversal in educational incentives. If employers stopped rewarding the proxies and started measuring the underlying abilities, students would once again have a strong reason to acquire those abilities. Understanding would matter because understanding would eventually be tested directly. Critical thinking would matter because someone would ask you to think. Historical perspective, quantitative intuition, scientific understanding, judgment, communication, craftsmanship, intellectual independence, and the capacity to learn would become valuable not because they raised a grade but because they made you more capable.

The structure of education might begin to change as well. Four years, 120 credits, a particular sequence of courses, and a particular collection of grades are themselves proxies. Some people might need years of broad education; others might become ready for particular kinds of work much sooner. Students might spend more time on subjects that genuinely enlarge their capabilities and less time accumulating boxes checked toward a credential. Education could become more varied in length and focus because the ultimate question would no longer be “Have you completed the prescribed pedigree?” but “What can you actually do?”

Instructors could then spend less energy policing cheating and managing grading systems and more energy doing what brought many of them into education in the first place: helping students understand things. Universities might even find themselves competing more on how much their students actually learn and less on what I think of as the corporate mission: maximize income, minimize risk, and build the brand.

There are plenty of difficulties. AI assessments could themselves contain biases, measure the wrong qualities, be manipulated, or become detached from actual job performance. Humans would still need to decide what characteristics legitimately matter, provide appropriate safeguards and accommodations, and continually compare assessment results with what happens in the real world. We would create new proxies, and we would eventually begin gaming them. Human nature is not going away.

But perhaps AI gives us an opportunity to move the measurement much closer to the thing being measured. That possibility strikes me as wonderfully ironic.

AI threatens education because it makes cognitive offloading easier than ever before. It can produce the essay without the thinking, the answer without the understanding, and the credential without necessarily revealing what lies behind it. But the same technology may eventually make it cheap to ask directly the question our elaborate system of grades, degrees, pedigrees, and credentials has always been trying indirectly to answer:

What does this individual actually know, understand, and know how to do—and how well can this person learn what comes next?

If we can answer that question reasonably well, much of the incentive to game education may disappear.

And AI, unexpectedly, might help make the purpose of education once again be education.

Acknowledgment

This essay grew from my observations about increasing cognitive offloading and from an extended conversation with ChatGPT about its effects on education. Our discussion began with my concern that students and instructors can now offload activities that are themselves central to learning, and eventually led us to a much more optimistic possibility: AI might reduce the importance of grades, degrees, pedigrees, and other educational proxies by making direct, individualized assessment of human knowledge, reasoning, learning ability, and skill inexpensive and practical.

I supplied the experiences, arguments, concerns, and perspective, including my experiences teaching How Things Work and my longstanding concerns about Goodhart’s law, credentialing, and institutional incentives. ChatGPT challenged some of my generalizations, helped develop the ideas of adaptive AI interviewing, direct assessment, and measuring the ability to learn rather than merely existing knowledge, explored their limitations and unintended consequences, and drafted and worked with me to revise this essay from our discussion.