TL;DR: Here’s the TL;DR version, if you are impatient, busy, or simply aren’t sure you want to spend that much time on this article.
Exams look precise, standardized, and objective, but they are often much noisier measures of learning than we admit. They test only a small, instructor-selected subset of what was taught, and they frequently measure things that may have little to do with the actual competence we care about: fast recall, comfort with formal testing, written expression, performance under artificial time pressure, and the luck of getting questions that line up with what a student knows best. That can disadvantage students who are highly competent in practical, diagnostic, oral, spatial, or real-world problem-solving settings, as well as students who simply perform badly in formal exam conditions.
The more important question is not whether exams are easy to standardize or hard to cheat on, but whether they resemble the way people will actually have to use their knowledge. If someone’s future work requires diagnosing faults, consulting manuals, researching evidence, collaborating, revising, or solving problems with tools and references available, then an isolated, closed-book, timed written exam may be measuring the wrong thing. AI has made this problem more visible because institutions are responding by bringing back handwritten exams and blue books. But before we make old assessments harder to cheat on, we should ask whether those assessments were good measures of competence in the first place.
The standard does not need to be lowered. In fact, the argument is for better evidence: assess the thing you actually care about, under conditions that make sense for that skill, and use enough different evidence over time that a student’s grade is not determined by a tiny sample of course content or by what happened during three hours on one particular day.
It is September.
Students are heading back to school, instructors are getting their courses underway, and once again there is a lot of discussion about exams. This time, of course, AI is part of the reason. Apparently, “blue books” are making a comeback in some universities — those little paper booklets students use to write exams by hand. The logic is pretty obvious: if the student is sitting in a room with a pencil and a paper booklet, they can’t ask ChatGPT to write the answer for them. Fair enough. That may indeed solve one particular problem.
The thing is though, I think we are skipping over a much more fundamental question:
Why are we giving the exam in the first place?
I have been asking questions like this for a VERY long time. They did not always make me popular. Back in 2017, I gave a talk called Grades and the Random Factor: How Randomness Affects Assessment. The premise was actually pretty simple.
We have long been confident that our “comprehensive” final exams provide a pedagogically sound assessment of what students have learned throughout a course, but is that really true? Suppose I “cover” a 400-page textbook, spend 30 or 40 hours lecturing on the material, add tutorials, discussions, exercises and whatever else I do during the semester, and then give my students a 100-question multiple choice final exam. What exactly have I measured? I certainly haven’t measured their mastery of everything we did in the course. I have measured their performance on the 100 questions I happened to ask. Those are NOT the same thing.
There is an old story about someone looking for their lost keys under a streetlight. A passerby asks where they lost them. “Over there,” they say, pointing somewhere else. “Then why are you looking here?” “Because the light is better here.” We do this all the time, and not just in education. We measure the things that are relatively easy to measure, and then, over time, we begin to forget that ease of measurement is one of the reasons we chose them. Exams are standardized. They produce nice, clean numbers. Multiple-choice exams are cheap to administer to large groups, and can be graded automatically. All of those things may be administratively useful, but NONE of them tells us whether the exam is actually measuring what we claim it is measuring.
Even supposedly “randomized” exams don’t escape this problem. Random questions are only random within the universe we have constructed for them. Someone chose what went into the question bank. Someone decided which topics mattered enough to include. Someone decided how much detail each question would address. Someone wrote the questions, selected the distractors, and decided what counted as the correct answer. If the learning-management system randomly chooses 50 questions from a bank of 500, or even 5,000, the selection may be random, but the dataset certainly isn’t. Worse, when I have examined textbook-supplied question banks closely, I have often found multiple questions that are essentially variations on the same thing. If those are overrepresented in the bank, then they are also more likely to be overrepresented on the exam. Are the questions valid? Are they reliable? Has anyone checked? Or are we simply assuming that because the system can generate a nice-looking test, it must therefore be a good one?
There is another problem with the claim that a grade represents mastery. Imagine six students who each know about 75% of the course material. We would probably call them all “B students.” The problem is that they may not know the SAME 75%. One student may know the first three quarters of the course extremely well and know almost nothing about the last quarter. Another may know roughly three quarters of every topic. Another may have gaps scattered throughout. Another may be exceptionally strong in the material I happen to think is most important and weak in things I barely care about. If I give all six students the same exam, their grades will depend quite heavily on which subset of the course I chose to put on that exam. In my 2017 talk I illustrated this visually: six students could all reasonably be described as knowing 75% of the material, but once we change the distribution of questions, their grades can change dramatically. The number we assign at the end looks precise. The measurement really isn’t.
These images compare what 75% “right” might look like on an exam where content is drawn equally from all teaching units. vs an exam where content has a stronger emphasis on later material from the course.
Each colour represents a main topic and each square represents a correct answer to an exam question.
Then there is the issue of standardization. We often talk about standardized testing as though giving everyone the same test under the same conditions automatically makes the assessment fair. It doesn’t. It makes the CONDITIONS the same. That is not the same thing. A traditional timed, written exam privileges a very particular collection of abilities: fast recall, working memory, written expression, comfort with formal testing, the ability to function under artificial time pressure, and often the ability to figure out what the instructor thinks is important. Those may all be useful abilities in some situations, but are they actually part of the thing we are supposed to be assessing?
Think about plumbers, electricians, aircraft maintenance technicians, nurses, programmers, mechanics, designers — the list could go on for pages. Someone may be exceptionally good at walking into a real situation, spotting what is wrong, diagnosing the problem, finding the information they need, choosing an appropriate response, and carrying it out safely and effectively. They may even work extremely well under genuine pressure. They may ALSO be terrible at writing about it in a formal exam. Why should their ability to write about solving a problem be allowed to stand in for their ability to actually SOLVE the problem? Conversely, someone may write perfectly well, think quickly, and cope quite capably with real-life pressure, but fall apart in the artificial environment of a formal high-stakes exam. Again, what are we actually measuring?
This does not mean that time pressure is never appropriate. If I am training someone who genuinely needs to make good decisions quickly under pressure, then by all means, TEST THAT. If an aircraft technician, emergency-room nurse, or power-system operator needs to recognize a dangerous situation rapidly and respond correctly, then their ability to do that is part of the competence we care about. The key is that the pressure itself is part of the learning outcome. If, on the other hand, what I want to know is whether someone can analyze a complex problem, find relevant information, evaluate evidence, arrive at a defensible conclusion, and communicate that conclusion clearly, then why would I take away their references, isolate them from everyone else, give them an arbitrary two- or three-hour limit, and insist that they produce the answer from memory? Is that how they will be expected to solve problems after they graduate? If not, why are we testing them that way?
One of the strangest traditions in formal education is the idea that looking things up somehow contaminates evidence of competence. In most professions, NOT looking something up when you are unsure can be irresponsible. Engineers consult standards. Physicians consult references. Programmers look up documentation. Aircraft maintenance personnel use manuals and checklists. Academics look things up CONSTANTLY. We consult papers, books, notes, colleagues, databases, and now AI. We revise things. We ask questions. We get feedback. We try again. Yet somehow we have decided that the best way to discover whether a student is competent is often to remove most of those tools and see what they can produce in one sitting.
I eventually allowed students in my classes to resubmit almost anything. If the point of an assignment was to demonstrate mastery of some concept or skill, and the first attempt showed that they had not quite mastered it yet, what exactly was wrong with letting them learn from the feedback, fix what they missed, and demonstrate that mastery later? If they have no opportunity to correct their mistakes, then what we have created is effectively just another test. In most real-life situations, people submit drafts, proposed solutions, prototypes, plans, designs, code, reports, and all kinds of other work, receive feedback, and then improve it. Why should education be LESS forgiving of iteration than the world we claim to be preparing students for?
High-stakes exams add yet another layer of randomness. We spend an entire semester gathering evidence of what a student can do, and then make a substantial part of their final grade depend on how they perform during a few hours on ONE particular day. Are they sick? Did their child keep them awake all night? Did they just get bad news from home? Do they have three exams in two days? Did they happen to study the parts of the course that I chose to put on THIS exam? None of those things necessarily tells us much about whether they understand chemistry, databases, accounting, Shakespeare, or thermodynamics, but all of them can have a substantial effect on the grade.
I experienced this myself. In my first year of university I took an upgrading chemistry course. I had straight A’s throughout the semester and on the midterm. Since it was an upgrading course, I reasoned that I clearly understood the material and only really needed to pass the final, so I spent my limited study time on other courses. I ended up with a C. I only realized afterward that the course still counted toward my GPA. Did that C suddenly reveal my TRUE level of chemistry competence? Of course not. My work all semester had already demonstrated that I knew the material. The C reflected how I performed on one particular exam, on one particular day, after making one particular decision about how to allocate my study time.
Whenever I raise these kinds of questions, someone will eventually suggest that changing the way we assess students means lowering standards. It does NOT. I am not arguing that we should make things easier simply because students would prefer that. I am arguing that we should make our assessments more ACCURATE. If a student is supposed to be able to repair an electrical system, make them repair one. If they need to diagnose faults, give them faults to diagnose. If they need to write, assess their writing. If they need to recall critical information instantly, assess that. If they need to work collaboratively, assess their ability to collaborate. If they need to research, LET THEM RESEARCH. Giving students more than one way to demonstrate competence does not lower the standard; it broadens the body of evidence we are willing to consider. It remains our responsibility to decide whether that evidence meets the criteria.
And now, of course, we have generative AI. Suddenly there is enormous pressure to make our old assessments “AI-proof.” Hence the return of handwritten exams and blue books. Perhaps handwritten exams really ARE the right tool for some learning outcomes. I have no objection to them merely because they are old technology. Sometimes an old tool is still the right tool. What bothers me is the assumption that because AI has made an old assessment easier to cheat on, the obvious solution is to find a more secure way to administer the SAME assessment. Maybe the problem is AI. Maybe the problem is the assignment. Those possibilities are not mutually exclusive.
Cheating matters. Authorship matters. If we are going to award someone a credential, we need good evidence that the person receiving it can actually do the things that credential claims they can do. I agree completely. In fact, that is precisely WHY I think we should be asking harder questions about assessment. If ChatGPT can produce a competent answer to the assignment, perhaps we should ask what evidence that assignment was really giving us in the first place. If the only way we can establish competence is by removing all tools, all collaboration, all references, all opportunities for revision, and putting the student under surveillance for three hours, perhaps we need to ask whether we are assessing the competence we think we are.
Most teachers genuinely want their students to learn. I have believed that throughout my teaching career. I also know that many instructors are overworked, underpaid, teaching classes that are too large, and required to work within institutional rules they did not design. Sometimes a multiple-choice exam is used because the instructor genuinely has no realistic alternative. I understand that. What I object to is pretending that administrative practicality somehow transforms the assessment into a pedagogically ideal measurement instrument.
Formal education is full of practices that have become nearly invisible through familiarity. We give midterms. We give finals. We impose time limits. We prohibit resources. We select a tiny subset of what we taught. We turn the result into a percentage. Then we behave as though that percentage tells us something much more precise than it actually does. As another school year begins, perhaps this is a good time for every instructor to look at each assessment in their course and ask a few uncomfortable questions: Why am I doing this? What am I actually trying to find out? Does this assessment give me good evidence of that? Are the restrictions part of the skill I am assessing, or are they simply traditions attached to the assessment? What else is this task measuring that I DIDN’T intend to measure?
And perhaps the most useful question of all is this one:
If exams did not already exist, would I invent THIS exam as the best possible way to find out what my students have learned?
If the answer is no, perhaps the blue book isn’t the problem.
Perhaps the exam is.
If you’re interested in more on this, feel free to look at the slides I prepared for a talk some years ago.
