AI or No AI? – Isn’t Actually the Right Question.

Approximate Reading Time: 8 minutes

What Did I Delegate? AI, Writing, and the Economics of Assistance (*)

(*) NOTE: This article talks about the use of AI and omits the environmental cost. It is in no way intended to dismiss the very real issues that surround the physical and planetary costs.


This article is part of the “It All Makes Sense Now” series:

A series about patterns, connections, hindsight, and those moments when scattered pieces suddenly form a bigger picture.


 


There is a great deal of discussion at the moment about whether writers should use AI.

I think that may be the wrong question.

Writers have always used tools and other people to help them write. Spellcheck checks our spelling. Grammar checkers identify awkward constructions. Editors reorganize manuscripts, question arguments, suggest cuts and sometimes propose entirely new wording. Research assistants find sources. Librarians help us locate information. Colleagues read drafts and say, “I don’t think you’ve actually made the point you think you’ve made.” Friends listen while we talk through an idea until we finally figure out what we were trying to say.

None of this normally causes us to conclude that the writer is no longer the writer.

Generative AI complicates things because it can potentially do all of those things—and it can also write the entire piece.

BUT: those are not the same use.

How I use AI

I use AI extensively in my writing, and it would be dishonest to pretend otherwise.

But I use it primarily as an organizer, filter, editor, sounding board, and research assistant, rather than as a generator of ideas.

A great deal of my writing begins in conversation. I notice something. I start pulling on it. I connect it to something else. I argue with myself—or, these days, with an AI. I reject interpretations. I refine distinctions. Somewhere in several thousand words of conversation, an article starts to emerge.

At that point, I may ask the AI to gather together the things I have said and organize them into a coherent draft.
The AI may generate the sentences in that draft, but that doesn’t mean it generated the thinking behind them.
Other times, I write the first draft myself and use AI much more conventionally: What’s unclear here? Where have I repeated myself? Is the structure working? What am I assuming that I haven’t actually established?

And sometimes it notices something I didn’t.

That happens with human editors, too.

Recently, while working on a short piece about silence and trust, AI suggested a sentence to capture a distinction I was making:

“This silence does not demand anything from me.”

That wasn’t quite right.

I changed it to:

“This silence does not demand payment from me.”

That one word mattered. I wasn’t merely talking about a silence that asked nothing. I was contrasting safe silence with weaponized silence in which ending the silence carries a price: accommodation, compliance, apology, attention, making oneself smaller.

The AI helped me see the place where a sentence was needed.

I decided what the sentence meant.

That is fairly representative of how I use AI.

AI can indeed write the whole thing.

Of course it can.
I could type:

Write me a moving 1,000-word essay about silence at a Quaker Meeting, connecting it to trauma, trust and faith.

and receive something quite plausible in seconds.
That would be a very different process.

This is why the question “Was AI used?” is not especially useful.

Spellcheck is a form of machine assistance. Grammarly is a more sophisticated form of machine assistance. Developmental editing is assistance. So is help in finding research to support an idea. Asking an AI to organize a long conversation into an article is assistance.

Asking it to decide what you think and write the article for you is something else.

There isn’t a nice clean line separating these activities. There is a continuum of delegation.

With that in mind, perhaps a more useful question is:

What did the writer delegate?

Did I delegate spelling? Sentence construction? Organization? Research? Criticism? The argument? The ideas themselves? The judgment about which ideas matter? The entire act of composition?

Those distinctions tell us far more about authorship than the simple fact that an AI was somewhere in the process. It is more like distinguishing between an editor that helps us tighten our writing, and a ghost writer who writes FOR us, and we simply attach our name to the work. That continuum has existed for about as long as we have written – and perhaps even before that when our stories were largely oral.

And the question of attribution and of full disclosure have been concerns for the entire time we have been storytelling in any form.

The difference now is not how much help we got from other humans, but how much help we got from AI. Yes, AI makes it a lot easier to “cheat” at writing, but the cheating is not the new part.

AND, the use of AI not only makes things easier, but can also make things more accessible, which brings us to…

The Economic Elephant in the Room

If I had unlimited funds, I would be quite happy to employ human editors and research assistants. I would prefer that.

I don’t.

Neither do most writers, and that matters.

Successful writers, academics, journalists and other professionals have long had unequal access to intellectual support. Some people can pay a developmental editor to spend hours working through a manuscript. Some work in institutions that provide research assistants, librarians, graduate students, editorial staff, administrative support and subscriptions to expensive databases. Some can hire someone to check hundreds of references or track down an obscure source. Some can pay for successful writers to review their drafts or teach them how to hone their skills.

Others do all of it themselves—or go without.

We talk a great deal about the digital divide: unequal access to computers, networks and information technologies.

AI may expose another divide that has been there all along: the assistance divide.

People with money or institutional support can buy other people’s time and expertise.

People without those resources generally cannot.

AI does not eliminate that inequality. Paid systems still exist. Human experts remain valuable, and in many circumstances they are substantially better than AI. AI makes mistakes. It can fabricate information, flatten distinctive prose, miss nuance and produce wonderfully confident nonsense. Using it well requires knowledge and judgment of its own.

But it can make some forms of assistance available to people who could never afford to purchase the human equivalent. That’s important.

A writer who cannot afford a developmental editor can still ask for an analysis of a manuscript’s structure. A student without a research assistant can get help constructing search terms and locating relevant literature, while still having time for other classes, and maybe even their personal lives. Someone struggling to organize a complicated argument can talk it through for two hours without worrying about whether they can afford another two hours.

A person with a disability, limited energy, or difficulty turning thoughts into conventional prose may be able to delegate some of the mechanical and organizational work while retaining the thinking that matters to them. Someone struggling with emotional overload can express their messy ideas, and have them turned into something that is clear, and logical without sounding like they are falling apart.

That doesn’t make AI and human assistance interchangeable.

It does mean that some kinds of intellectual scaffolding that were once disproportionately available to people with money, employment, institutional affiliation, or unusually patient friends have suddenly become much cheaper.

I think that’s worth talking about.

Assistance still requires judgment.

There is an important catch.

Delegating work doesn’t necessarily mean delegating responsibility.

If I employ a research assistant who gives me a bad source and I publish an unsupported claim, the claim still appears under my name.

The same applies when my research assistant happens to be silicon.

I routinely ask AI to investigate claims I want to make. That does not mean I regard whatever it produces as evidence. I want sources. I want to know whether those sources actually support the claim. If the evidence contradicts what I thought, then my claim needs to change.

Similarly, when AI reorganizes my writing, I remain responsible for deciding whether its organization represents what I meant. Sometimes it doesn’t. Sometimes its version is prettier than mine and less accurate. Sometimes it smooths away precisely the odd little thing that made the piece sound like me. Sometimes it confidently explains what I “really mean,” and I get to say, “No. That’s not what I mean at all.”

That ability to say no is important.

Perhaps one measure of authorship is not simply who generated the words, but who has both the authority and the capacity to recognize when they are wrong.

Show your work

I’ve been thinking that one useful way to make my own practice transparent would be occasionally to publish the original material alongside the finished work.

When an article begins with a draft I’ve written myself, I could make that draft available.
Readers could see what changed. They could see what I wrote, what the AI helped reorganize, what was removed, what was added, and what survived intact.

For pieces that emerge from long conversations, that’s harder—nobody needs 18,000 words of me wandering around an idea before discovering what I think—but I can still describe the process.

That seems more informative to me than attaching a label saying AI USED.
It shows what the AI actually did.

So yes, I use AI

Quite a lot.

I use it to organize, filter, search, challenge, summarize, restructure and word-wrangle. I use it to help me find research and then expect to be able to inspect the evidence. I use it to notice connections and problems. I sometimes use sentences it proposes. I often change them. I frequently reject them.

And sometimes, after several thousand words of discussion, I ask it to take everything I’ve said and turn it into something another human being might actually want to read.

What I try not to delegate is the part that matters most to me:

  • What do I actually think?
  • What am I trying to say?
  • Is this true?
  • Is this fair?
  • Is this mine?

Those decisions remain my responsibility, and mine alone.

So when I encounter the increasingly common question, “Did you use AI to write this?”, my answer is yes—but I don’t think that’s a particularly interesting question.

The more interesting one is:

What did you delegate to it?

That is a question I am quite happy to answer.

In the end, it still comes down to the individual author’s integrity: what they delegate, what they claim as their own, and how honestly they represent the process when that distinction matters.

? From the Sausage Factory

If I’m going to argue that “Did you use AI?” is less useful than “What did you delegate to it?”, it seems only fair to show you what I mean.

This article didn’t begin with a prompt asking AI to write an article about AI. It began with these thoughts, expressed over the course of a conversation:

Is the difference between Word’s spellcheck, Grammarly and generative AI really one of kind, or largely one of degree?

I mostly use AI to craft articles and emails that result from things I have already said. I use it as an organizer and filter much more than as a generator. I also use it to find research relevant to claims I want to make.

If I had the money, I would happily hire human editors and research assistants. Could disciplined use of AI help bridge not just a digital divide, but an economic divide in access to assistance?

Maybe the useful question isn’t “Did you use AI?” but “What did you delegate to it?” And perhaps I should occasionally make my original drafts available so readers can see for themselves what changed.

What happened in between: AI helped me separate different kinds of assistance, and our back-and-forth produced the idea of looking at AI use in terms of delegated compositional agency. When I described my own process, another distinction emerged: generating prose is not necessarily the same as generating the intellectual content of that prose. I introduced the economic-access argument; AI helped develop it into the idea of an “assistance divide.”

AI then organized those ideas into the article you just read, proposed wording and examples, and gave me things to accept, reject, modify, or argue with. And THEN, I proofed it, made some additional changes and added a few things.

So yes: AI generated a substantial amount of the prose.

And now you can also see what actually went into this particular sausage.

Be the first to like.

The Philosophy of Enoughness and What it Says About Stewardship and Possession

Approximate Reading Time: 6 minutes
Enoughness
A series about finding contentment, sufficiency, and a sense of enough in a complicated world.
19th century secretary desk

A 19th century secretary desk, bought by my mom from neighbours who were antique dealers, and passed on to me when my mom passed.

I realized something recently about the way I become attached to things.

And by things, I mean almost everything: animals, houses, land, furniture, tools, books, gardens, old objects, even relationships.

My attachment grows through care. I clean things. Fix them. Feed them. Maintain them. Learn their quirks. Protect them. Improve them. Keep them useful. Notice when something is wrong. Invest time and attention. The more I care for something, the more I care about it.

That seems obvious once stated, but it led me somewhere interesting. Because not everyone appears to build attachment that way.
Some people seem to relate to things primarily through possession.
I bought it.
It is MINE.
Therefore I decide what happens to it.

Those are two very different architectures.

  • In a stewardship model, relationship creates responsibility.
  • In a possession model, ownership creates authority.

That distinction turns out to matter.

Stewardship is relational. Stewardship is not really about ownership at all. You can steward something you do not own. A rented house. A community garden. A neighbour’s dog. A public trail. A borrowed tool. A child. A friendship. A piece of land that will long outlive you.

The relationship is based on care rather than title.
And care has consequences.
If something is in my care, then I owe it something. Maintenance. Attention. Protection. Restraint. Perhaps even improvement. This is not about anthropomorphising STUFF. Ask anyone who has survived the devastation of their homeland due to war.

Stewardship feels, to me, like a very old biological arrangement. Living things survive by caring for things. Parents care for offspring. Animals maintain nests, dens, territories and social relationships. Organisms invest energy where that investment supports survival, reproduction, safety or continuity.

Care creates attachment because care itself is costly. It costs time, and energy, and attention. Once you have invested those things repeatedly, the object of that care is no longer interchangeable. It becomes unique becuase of the care I have given it.

We are not unique in this. Animals possess things too. Animals defend territories. They guard food. They compete for mates. They cache resources. They protect nesting sites.
It might seem like this is the opposite of stewardship.
I don’t think it is.
Territoriality and possession in nature are usually constrained by something else:

Enoughness.

A territory is useful because it provides enough food, enough shelter, enough reproductive opportunity, enough safety. But territory itself has costs. It must be patrolled. Defended. Marked. Maintained. A territory that is too small may not provide enough. A territory that is too large may cost more to defend than it returns.

So the underlying logic is not:

Possess as much as possible.

It is:

Secure enough.

That is a very different thing.
Possessiveness actually hinges on enoughness.
That’s where the two collide. Possession itself is not the problem. Possession beyond enoughness is. Every organism needs resources: Food. Shelter. Safety. Space. Access to mates. A margin against bad times.

Wanting those things is not pathology. It is biology. Even storing surplus can fit perfectly well within Enoughness. A squirrel caching food for winter is not displaying avarice. A human storing grain against crop failure is not being greedy. A household saving money for illness, old age or disaster is not engaged in gluttony.

So then, the important question is:

Does acquisition remain connected to a need?

Enoughness is the stopping rule. Stop when you have:

  • Enough food.
  • Enough safety.
  • Enough territory.
  • Enough stored resources.
  • Enough.

What happens when the stopping rule disappears?

That is where things change.

When acquisition continues after sufficiency has already been reached, possession becomes something different. The object is no longer merely to secure enough resources.

The object becomes MORE:

  • More land.
  • More wealth.
  • More food.
  • More status.
  • More sex.
  • More stimulation.
  • More attention.
  • More control.

The appetite stops being regulated by need. That’s gluttony. And I don’t just mean  overeating. I mean gluttony of all kinds. Anything that involves consumption or acquisition after satiety has ceased to govern behaviour.

Humans have always been capable of excess, but for most of our history, there were practical limits. I don’t think greed suddenly appeared one Tuesday afternoon in Mesopotamia. But for most of human evolutionary history, there were actual hard limits on how much anyone could accumulate. If you were nomadic, you had to carry what you owned. If you weren’t, food still couldn’t be stored indefinitely because it spoiled. Objects required maintenance, and land required direct use or defence. Back then, resources that could not be used, guarded, transported or exchanged had limited value, and in small groups, extreme hoarding could also undermine the reciprocal relationships on which their survival depended. The capacity for gluttony still existed, but the infrastructure for large-scale possession beyond enoughness did not.

Then agriculture changed the equation.

Agriculture made it possible to actually HAVE surplus. It allowed food to be stored in quantity. Land could be controlled. Herds could be maintained in order to reproduce and hold wealth. Surplus could persist from one season to another. And then from one generation to another.

Once wealth can be stored, defended and inherited, possession begins to separate from immediate use.
You do not need to eat all the grain. You can control the granary. You do not need to personally use all the land. You can control access to it.

That is a profound change.

It makes disproportionate accumulation possible in a way that nomadic life can’t.

Then…. Money made possession abstract.
Money disconnected wealth from stuff. Now wealth no longer needed to consist of actual things you personally used. It could become its own claim on things. A portable, transferable, inheritable claim on future resources. That matters because abstraction removes another natural constraint. You cannot personally eat ten thousand loaves of bread. But you can possess enough money to buy ten thousand loaves of bread and experience no biological signal saying:

That appears to be enough bread.

Money does not rot, or fill your stomach, or crowd your house.
It does not become physically impossible to carry. It makes it possible to detach accumulation from, well, everything, really.

Then machines removed another ceiling.

Human beings can only personally produce so much. One person can harvest only so much grain. Dig only so much coal. Spin only so much cloth. Manage only so much land. Ownership of machinery allows one person to control the output of many others. Steam power amplified that enormously. Industrialization amplified it again. Modern technology amplified it still further.

Today, one person can control vast claims on resources scattered across the planet while having no direct relationship with any of it. Nothing to tend, or maintain. No physical contact needed. Hell, you don’t even need to know anything about the stuff you own.

Possession exists almost entirely as abstraction.
And there is essentially no natural stopping point.

Stewardship has a built-in limit.

This may be one of the most important differences.
Stewardship is naturally constrained. You can only care for so much. There are only so many animals you can properly tend. Only so much land you can maintain. Only so many objects you can repair and use. Only so many relationships you can genuinely nurture. Care consumes time and energy. Therefore stewardship tends to impose its own version of Enoughness.

At some point you reach capacity.

Possession doesn’t need that.

You can own things you never see.
Land you never visit.
Buildings you never enter.
Companies whose workers you never meet.
Financial assets representing resources you could not use in many lifetimes.

Once possession becomes detached from stewardship, it can continue indefinitely.

This separates two different meanings of “mine”.

There are two very different ways of saying:

This is mine.

One means:

This is mine, therefore I am responsible for it.

The other means:

This is mine, therefore I am entitled to control it.

Those statements can coexist, of course.

Most of us probably move between them depending on circumstances. But they point in very different directions. One leads toward stewardship. The other can lead toward possession without constraint. And possession without constraint has no natural reason to stop at Enough.

Maybe the oldest question is simply: how much is Enough?

I keep coming back to that.

Enoughness does not require self-denial.
It does not demand poverty. It does not require everyone to possess exactly the same amount. It does not even object to abundance.

It simply asks whether wanting or possessing still serves a purpose.

Do I need this? Will I use it? Can I care for it? Does it protect against a realistic future need? Does having more improve anything meaningful?

Or has more become the goal in itself?
That may be where stewardship and possession finally diverge.

Stewardship asks:

What does this require of me?

Possession asks:

What power to I acquire with this?

And gluttony begins when the answer to How much is enough?
becomes:

More.


P.S.
Years ago, someone once told me that the two worst sins a person could commit were hypocrisy and gluttony.
I have been thinking about that differently lately.
They are oddly complementary failures.
Hypocrisy is a failure of alignment between what we say we believe and how we actually behave.
Gluttony, in the broader sense, is a failure of alignment between what we need and what we continue to take.
One says: the rule matters, except when it constrains me.
The other says: enough should be enough, except when I want more.

Both are, in their own way, refusals of limits.
And perhaps that is why the pairing has stayed with me.

My Mom

 

 

In the end I still think my mom was right.
She said,

“If you have what you need and all that’s left is what you want, you are already rich.”

Be the first to like.

The Six and the Nine of Redemption

Approximate Reading Time: 4 minutes
It All Makes Sense Now
A series about patterns, connections, hindsight, and those moments when scattered pieces suddenly form a bigger picture.

There is a familiar meme about perspective: two people stand on opposite sides of a number painted on the ground. One sees a 6. The other sees a 9. They are looking at the same thing, but from different positions, so they interpret it differently.

I have been thinking that redemption stories may work that way too. I like redemption stories because I am interested in the transformation. I want to understand what happened inside the person. What finally got through? What did they have to confront? What did they learn? What did they stop defending? What began to matter to them that had not mattered before? What changed enough that they started behaving differently?

That is also one of the reasons I like memoirs. I get access, however imperfectly, to another person’s way of making sense of a life. Sometimes I learn something useful from how someone else coped, changed, failed, adapted, or survived. For me, redemption is interesting because change is interesting. Learning is interesting.

But what if another person watches the very same story and barely notices the transformation at all? What if, for them, the meaningful part is what happens afterward: the asshole is forgiven, the difficult person is welcomed back, the misunderstood man is finally appreciated, everyone stops being angry with him, and people recognize that underneath it all he was really a good person?

That is a very different redemption story.

I have been wondering whether that difference helps explain something I could never quite understand about someone I used to know. He liked redemption movies. A LOT. Groundhog Day was his favourite film. Yet throughout our relationship, one of the most consistent difficulties was his resistance to any kind of lasting behavioural change. More than once, he told me that if I needed him to change his behaviour, that meant I did not love him.

That statement only makes sense if changing behaviour is experienced as a threat to the self. If the self is rigid — this is who I am — and who I am is Good, then acknowledging that some important part of one’s behaviour needs changing can feel like conceding that the self itself is not in fact Good. Accountability stops being information and becomes an attack.

In that framework, redemption through self-transformation is almost impossible to contemplate. But redemption through other people’s changed perception is perfectly compatible with it. Perhaps, for some people, nothing fundamental has to happen inside them. You simply have to understand them better. You have to stop being angry. You have to recognize their good intentions. You have to forgive them. You have to see that they are actually wonderful.

The transformation that they see happens in you.

That thought made me remember a scene from Dead Like Me. Mandy Patinkin’s character is talking about sincerity and says something roughly like, “I do this with my face, and then the sincere words come out.” The point is not that genuine feeling eventually appears to match the face. The point is that the appropriate performance can be produced without any corresponding feeling underneath it.

Perhaps that is another version of the same 6/9 divide. In my version of redemption, changed behaviour is evidence of an internal process. The person can see what they did, tolerate the discomfort of knowing it, accept consequences, and behave differently even when forgiveness is not guaranteed.

In the other version, the gestures themselves may constitute redemption. Say the right thing. Make the right face. Be nice again. Offer affection. Return after the rupture with a friendly expression and behave as though the rupture has ended. Then the other person is supposed to complete the process by forgiving.

That pattern happens repeatedly in some marriages. After conflict, one partner could simply reappear with a new friendly face. But there was often no repair underneath it: no sustained examination of what had happened, no meaningful answer to the other partner’s repeated question, What are YOU willing to do to repair this relationship?, and no durable behavioural change.

From my side, the redemption gesture fails because the underlying problem remains. There is no repair. No accountability. No self-reflection. From their side, perhaps the bewildering part was that the gesture did not work permanently. I came back. I was nice. I said the appropriate things. Why are you still upset?

If redemption is understood primarily as the restoration of acceptance, then continued anger or mistrust looks unreasonable. The problem is no longer the original behaviour. The problem becomes the other person’s refusal to forgive.

That produces two radically different readings of the same event.

From one side: You hurt me. You have not changed the behaviour that hurt me. Therefore the rupture is not repaired.

From the other: I made the redemption gesture. You are still holding this against me. Therefore you are refusing to let the relationship heal.

Six and nine. Same movie. Different meaning.

The crucial difference is where the change is expected to occur. In one version, redemption requires the person who caused harm to change. In the other, redemption is achieved when everyone else changes their opinion of them.

That distinction also helps explain why redemption stories can be deeply satisfying to people for entirely different reasons. I may be watching Ebenezer Scrooge and thinking, Look at what it took to break through his defenses. Look at how his understanding of himself and other people changed. Someone else might be watching the same ending and thinking, Look. He bought a big turkey and everyone loves him now.

Those are not merely different emphases. They are radically different theories of redemption. One is about becoming someone capable of different choices. The other is about finally being seen the way you always believed you deserved to be seen. Perhaps more importantly, they are entirely different views of reality itself. One accepts that they bear responsibility for the impact of their behaviour, that other people’s reactions may tell them something important about how they are showing up in the world, and that this warrants self-reflection. The other holds their own behaviour to be wholly justified, often by default, and expects everyone else to recognize their “correctness” and to adjust accordingly. One accepts that they must sometimes change. The other expects the change to happen elsewhere.

And perhaps that is why the 6/9 image feels so apt. Sometimes two people can spend decades looking at precisely the same thing and never realize that they are not arguing about the answer. They are standing on opposite sides of the number.

Be the first to like.

A New Look for my Blog….

Approximate Reading Time: < 1 minute

I’ve been doing some reorganizing.

Over time, several recurring themes have emerged in my writing—some rooted in projects I’ve been working on for years, and others that seem to have quietly become projects while I wasn’t looking. Rather than have them scattered throughout the blog, I’ve decided to give them recognizable homes of their own.

I can’t promise frequent or regular posts just yet. There is still rather a lot going on in my life. But when something I write belongs naturally to one of these themes, I’ll publish it under its series banner. Some of the series also have homes on Substack; others are newer and will grow here as I have things to say.

For now, these are the six series you’ll start seeing:

10,000 Eggs
A series about some of the adventures I’ve had and the things I’ve learned along the way while trying to live a sustainable, ethical life on a small farm in rural Alberta.
The Happy Healer
A series that offers reflections and practical tools for healing, recovery, and rebuilding.
Don’t Ask a Fish
A series examining assumptions, ideas, behaviours, and ways of seeing that are often left unexamined.
The Library Theory of Mind
A series exploring what a mind might be like if it were a library.
Enoughness
A series about finding contentment, sufficiency, and a sense of enough in a complicated world.
It All Makes Sense Now
A series about patterns, connections, hindsight, and those moments when scattered pieces suddenly form a bigger picture.

 

Be the first to like.

The Becker Lazy Test (BLT) for Educational Games

Approximate Reading Time: 2 minutes

The Becker Lazy Test is something I developed many years ago as part of my 4-PEG game assessment template. (4PEG = 4 Pillars of Educational Games). Some of that is described in detail in my book, Choosing and Using Digital Games in the Classroom – A Practical Guide, published in 2016 by Springer.

Here’s the TL;DR version:

When I am examining a game, I see how far I can get without reading or learning anything. I simply follow the known mechanics (if obvious) or click randomly. If I can get to the end this way, it does NOT pass as an educational game. I don’t need to learn anything to get through it.

Put very simply, it should not be possible to get through an educational game by brute force or by random chance alone. Now, I know that this may seem very similar to Margaret Gredler’s claims about games vs simulations made in her chapter on simulations and games in the AECT Handbook of 1996 (Gredler, 1996) where she said that games should not have a random factor. If you read my other book, The Guide to Computer Simulations and Games – especially the chapter on randomness – you will already know how important the “random factor” is to BOTH simulations AND games. Gredler used this as a way to distinguish simulations from games (which is misguided), but she also used this as a way to separate games she liked from those she found frivolous.

What I’m saying is: If random actions on MY part can get me through the game, then it’s not an educational game. Note: That says nothing about whether the game is fun or entertaining. ALmost ALL games can, should, and MUST have at least some randomness, or else it is nothing more than a branching story.

“Lazy Jane” by Shelf Silverstein, originally published in Where the Sidewalk Ends

SO, these are the questions that go along with the Becker Lazy Test. A YES answer to any of these constitutes a PASS. A PASS is a BAD thing.

  1. Is it possible to get through the game by randomly clicking on things? In other words, could I win the game by simply memorizing which things to click without knowing what those things are?
  2. Are the educational objectives included among the required learning in the game?
  3. Is it possible to get through the game while ignoring the learning objectives? The required learning in the game should be PART of the game and not only found in pop-up screens of text.

 

There. I said what I said.

 

Be the first to like.

Why Are We Still Giving Exams?

Approximate Reading Time: 9 minutes

TL;DR: Here’s the TL;DR version, if you are impatient, busy, or simply aren’t sure you want to spend that much time on this article.

Exams look precise, standardized, and objective, but they are often much noisier measures of learning than we admit. They test only a small, instructor-selected subset of what was taught, and they frequently measure things that may have little to do with the actual competence we care about: fast recall, comfort with formal testing, written expression, performance under artificial time pressure, and the luck of getting questions that line up with what a student knows best. That can disadvantage students who are highly competent in practical, diagnostic, oral, spatial, or real-world problem-solving settings, as well as students who simply perform badly in formal exam conditions.

The more important question is not whether exams are easy to standardize or hard to cheat on, but whether they resemble the way people will actually have to use their knowledge. If someone’s future work requires diagnosing faults, consulting manuals, researching evidence, collaborating, revising, or solving problems with tools and references available, then an isolated, closed-book, timed written exam may be measuring the wrong thing. AI has made this problem more visible because institutions are responding by bringing back handwritten exams and blue books. But before we make old assessments harder to cheat on, we should ask whether those assessments were good measures of competence in the first place.

The standard does not need to be lowered. In fact, the argument is for better evidence: assess the thing you actually care about, under conditions that make sense for that skill, and use enough different evidence over time that a student’s grade is not determined by a tiny sample of course content or by what happened during three hours on one particular day.


It is September.

Students are heading back to school, instructors are getting their courses underway, and once again there is a lot of discussion about exams. This time, of course, AI is part of the reason. Apparently, “blue books” are making a comeback in some universities — those little paper booklets students use to write exams by hand. The logic is pretty obvious: if the student is sitting in a room with a pencil and a paper booklet, they can’t ask ChatGPT to write the answer for them. Fair enough. That may indeed solve one particular problem.

The thing is though, I think we are skipping over a much more fundamental question:

Why are we giving the exam in the first place?

I have been asking questions like this for a VERY long time. They did not always make me popular. Back in 2017, I gave a talk called Grades and the Random Factor: How Randomness Affects Assessment. The premise was actually pretty simple.

We have long been confident that our “comprehensive” final exams provide a pedagogically sound assessment of what students have learned throughout a course, but is that really true? Suppose I “cover” a 400-page textbook, spend 30 or 40 hours lecturing on the material, add tutorials, discussions, exercises and whatever else I do during the semester, and then give my students a 100-question multiple choice final exam. What exactly have I measured? I certainly haven’t measured their mastery of everything we did in the course. I have measured their performance on the 100 questions I happened to ask. Those are NOT the same thing.

There is an old story about someone looking for their lost keys under a streetlight. A passerby asks where they lost them. “Over there,” they say, pointing somewhere else. “Then why are you looking here?” “Because the light is better here.” We do this all the time, and not just in education. We measure the things that are relatively easy to measure, and then, over time, we begin to forget that ease of measurement is one of the reasons we chose them. Exams are standardized. They produce nice, clean numbers. Multiple-choice exams are cheap to administer to large groups, and can be graded automatically. All of those things may be administratively useful, but NONE of them tells us whether the exam is actually measuring what we claim it is measuring.

Even supposedly “randomized” exams don’t escape this problem. Random questions are only random within the universe we have constructed for them. Someone chose what went into the question bank. Someone decided which topics mattered enough to include. Someone decided how much detail each question would address. Someone wrote the questions, selected the distractors, and decided what counted as the correct answer. If the learning-management system randomly chooses 50 questions from a bank of 500, or even 5,000, the selection may be random, but the dataset certainly isn’t. Worse, when I have examined textbook-supplied question banks closely, I have often found multiple questions that are essentially variations on the same thing. If those are overrepresented in the bank, then they are also more likely to be overrepresented on the exam. Are the questions valid? Are they reliable? Has anyone checked? Or are we simply assuming that because the system can generate a nice-looking test, it must therefore be a good one?

There is another problem with the claim that a grade represents mastery. Imagine six students who each know about 75% of the course material. We would probably call them all “B students.” The problem is that they may not know the SAME 75%. One student may know the first three quarters of the course extremely well and know almost nothing about the last quarter. Another may know roughly three quarters of every topic. Another may have gaps scattered throughout. Another may be exceptionally strong in the material I happen to think is most important and weak in things I barely care about. If I give all six students the same exam, their grades will depend quite heavily on which subset of the course I chose to put on that exam. In my 2017 talk I illustrated this visually: six students could all reasonably be described as knowing 75% of the material, but once we change the distribution of questions, their grades can change dramatically. The number we assign at the end looks precise. The measurement really isn’t.

These images compare what 75% “right” might look like on an exam where content is drawn equally from all teaching units. vs an exam where content has a stronger emphasis on later material from the course.
Each colour represents a main topic and each square represents a correct answer to an exam question.

 


Then there is the issue of standardization. We often talk about standardized testing as though giving everyone the same test under the same conditions automatically makes the assessment fair. It doesn’t. It makes the CONDITIONS the same. That is not the same thing. A traditional timed, written exam privileges a very particular collection of abilities: fast recall, working memory, written expression, comfort with formal testing, the ability to function under artificial time pressure, and often the ability to figure out what the instructor thinks is important. Those may all be useful abilities in some situations, but are they actually part of the thing we are supposed to be assessing?

Think about plumbers, electricians, aircraft maintenance technicians, nurses, programmers, mechanics, designers — the list could go on for pages. Someone may be exceptionally good at walking into a real situation, spotting what is wrong, diagnosing the problem, finding the information they need, choosing an appropriate response, and carrying it out safely and effectively. They may even work extremely well under genuine pressure. They may ALSO be terrible at writing about it in a formal exam. Why should their ability to write about solving a problem be allowed to stand in for their ability to actually SOLVE the problem? Conversely, someone may write perfectly well, think quickly, and cope quite capably with real-life pressure, but fall apart in the artificial environment of a formal high-stakes exam. Again, what are we actually measuring?

This does not mean that time pressure is never appropriate. If I am training someone who genuinely needs to make good decisions quickly under pressure, then by all means, TEST THAT. If an aircraft technician, emergency-room nurse, or power-system operator needs to recognize a dangerous situation rapidly and respond correctly, then their ability to do that is part of the competence we care about. The key is that the pressure itself is part of the learning outcome. If, on the other hand, what I want to know is whether someone can analyze a complex problem, find relevant information, evaluate evidence, arrive at a defensible conclusion, and communicate that conclusion clearly, then why would I take away their references, isolate them from everyone else, give them an arbitrary two- or three-hour limit, and insist that they produce the answer from memory? Is that how they will be expected to solve problems after they graduate? If not, why are we testing them that way?

One of the strangest traditions in formal education is the idea that looking things up somehow contaminates evidence of competence. In most professions, NOT looking something up when you are unsure can be irresponsible. Engineers consult standards. Physicians consult references. Programmers look up documentation. Aircraft maintenance personnel use manuals and checklists. Academics look things up CONSTANTLY. We consult papers, books, notes, colleagues, databases, and now AI. We revise things. We ask questions. We get feedback. We try again. Yet somehow we have decided that the best way to discover whether a student is competent is often to remove most of those tools and see what they can produce in one sitting.

I eventually allowed students in my classes to resubmit almost anything. If the point of an assignment was to demonstrate mastery of some concept or skill, and the first attempt showed that they had not quite mastered it yet, what exactly was wrong with letting them learn from the feedback, fix what they missed, and demonstrate that mastery later? If they have no opportunity to correct their mistakes, then what we have created is effectively just another test. In most real-life situations, people submit drafts, proposed solutions, prototypes, plans, designs, code, reports, and all kinds of other work, receive feedback, and then improve it. Why should education be LESS forgiving of iteration than the world we claim to be preparing students for?

High-stakes exams add yet another layer of randomness. We spend an entire semester gathering evidence of what a student can do, and then make a substantial part of their final grade depend on how they perform during a few hours on ONE particular day. Are they sick? Did their child keep them awake all night? Did they just get bad news from home? Do they have three exams in two days? Did they happen to study the parts of the course that I chose to put on THIS exam? None of those things necessarily tells us much about whether they understand chemistry, databases, accounting, Shakespeare, or thermodynamics, but all of them can have a substantial effect on the grade.

I experienced this myself. In my first year of university I took an upgrading chemistry course. I had straight A’s throughout the semester and on the midterm. Since it was an upgrading course, I reasoned that I clearly understood the material and only really needed to pass the final, so I spent my limited study time on other courses. I ended up with a C. I only realized afterward that the course still counted toward my GPA. Did that C suddenly reveal my TRUE level of chemistry competence? Of course not. My work all semester had already demonstrated that I knew the material. The C reflected how I performed on one particular exam, on one particular day, after making one particular decision about how to allocate my study time.

Whenever I raise these kinds of questions, someone will eventually suggest that changing the way we assess students means lowering standards. It does NOT. I am not arguing that we should make things easier simply because students would prefer that. I am arguing that we should make our assessments more ACCURATE. If a student is supposed to be able to repair an electrical system, make them repair one. If they need to diagnose faults, give them faults to diagnose. If they need to write, assess their writing. If they need to recall critical information instantly, assess that. If they need to work collaboratively, assess their ability to collaborate. If they need to research, LET THEM RESEARCH. Giving students more than one way to demonstrate competence does not lower the standard; it broadens the body of evidence we are willing to consider. It remains our responsibility to decide whether that evidence meets the criteria.

And now, of course, we have generative AI. Suddenly there is enormous pressure to make our old assessments “AI-proof.” Hence the return of handwritten exams and blue books. Perhaps handwritten exams really ARE the right tool for some learning outcomes. I have no objection to them merely because they are old technology. Sometimes an old tool is still the right tool. What bothers me is the assumption that because AI has made an old assessment easier to cheat on, the obvious solution is to find a more secure way to administer the SAME assessment. Maybe the problem is AI. Maybe the problem is the assignment. Those possibilities are not mutually exclusive.

Cheating matters. Authorship matters. If we are going to award someone a credential, we need good evidence that the person receiving it can actually do the things that credential claims they can do. I agree completely. In fact, that is precisely WHY I think we should be asking harder questions about assessment. If ChatGPT can produce a competent answer to the assignment, perhaps we should ask what evidence that assignment was really giving us in the first place. If the only way we can establish competence is by removing all tools, all collaboration, all references, all opportunities for revision, and putting the student under surveillance for three hours, perhaps we need to ask whether we are assessing the competence we think we are.

Most teachers genuinely want their students to learn. I have believed that throughout my teaching career. I also know that many instructors are overworked, underpaid, teaching classes that are too large, and required to work within institutional rules they did not design. Sometimes a multiple-choice exam is used because the instructor genuinely has no realistic alternative. I understand that. What I object to is pretending that administrative practicality somehow transforms the assessment into a pedagogically ideal measurement instrument.

Formal education is full of practices that have become nearly invisible through familiarity. We give midterms. We give finals. We impose time limits. We prohibit resources. We select a tiny subset of what we taught. We turn the result into a percentage. Then we behave as though that percentage tells us something much more precise than it actually does. As another school year begins, perhaps this is a good time for every instructor to look at each assessment in their course and ask a few uncomfortable questions: Why am I doing this? What am I actually trying to find out? Does this assessment give me good evidence of that? Are the restrictions part of the skill I am assessing, or are they simply traditions attached to the assessment? What else is this task measuring that I DIDN’T intend to measure?

And perhaps the most useful question of all is this one:

If exams did not already exist, would I invent THIS exam as the best possible way to find out what my students have learned?

If the answer is no, perhaps the blue book isn’t the problem.

Perhaps the exam is.


If you’re interested in more on this, feel free to look at the slides I prepared for a talk some years ago.

Be the first to like.

Game Design IS Systems Thinking: Why the Player Changes Everything

Approximate Reading Time: 7 minutes

I have spent much of my life thinking about, and working with systems.

Computer systems. Educational systems. Organizational systems. Family systems. Legal systems. Agricultural systems. Ecological systems.

More recently, I have been thinking about how law firms might translate the idea of “trauma-informed practice” into actual procedures rather than leaving it as a collection of good intentions.

Somewhere in the middle of that discussion, something clicked.

I realized that many of the tools I instinctively use when looking at systems come from game design.

And then I realized something larger:

Game design may be one of the most complete models we have for thinking about systems in general.

Not because everything is a game. It isn’t. It really isn’t.

But games are self-contained systems deliberately constructed around the experience and behaviour of people operating inside them (a.k.a. “The Players”). And that last part turns out to matter enormously.

Everything Is a System

A system does not need to be complicated.

It needs components that interact.

A law firm is a system. A school is a system. A family is a system. A hospital is a system. A housing cooperative is a system. A pasture is a system. An ecosystem is a system.

Each contains some combination of:

  • actors;
  • rules;
  • resources;
  • constraints;
  • information;
  • choices;
  • incentives;
  • feedback;
  • time;
  • consequences.

Change one component and something else is affected.
Sometimes the change is exactly what we intended.

Sometimes it isn’t.

That is where systems become interesting.

A law firm may have a perfectly reasonable procedure. Consider a hearing that’s been ordered:

The Court sends the hearing information.
The lawyer receives it.
The assistant handles logistics.
The client receives what they need.

Nothing about that sounds problematic.

But if nobody owns the specific task of confirming that the client actually knows how to attend the hearing, the system can produce a client sitting at home three hours before court wondering whether she is supposed to be there in person or whether a link will appear.

Nobody designed that outcome.
The system produced it.
And yet, it’s still a systems problem.

It is also exactly the kind of problem game designers encounter constantly.

The Rulebook Is Not the Game

One of the most important lessons in game design is that the rules you write are not the game.

The game is what happens when actual players encounter those rules.

You can design an elegant mechanic and discover that nobody uses it. You can create a reward you think will encourage cooperation and discover that it encourages hoarding. You can carefully balance resources and watch one unexpected strategy completely dominate play. You can write crystal-clear instructions and then sit beside a playtester who has no idea what to do.

A poor designer says:

The player is doing it wrong.

A good designer says:

Interesting. What did I design that made this response sensible?

That is a profound shift in perspective.

And I think many non-game systems would improve dramatically if their designers adopted it.
A policy may be what the organization says should happen. It says very little about what the ‘players’ will do about it.

The system is what people actually experience when the policy meets reality.

The Player Is the Difference

There are many approaches to systems thinking, design, gap analysis, and process improvement. They all offer useful tools.

But game design adds something I now think is fundamental:
It centres the design on an autonomous participant.
The player is not simply the recipient of the system.
The player is the reason that the game exists in the first place.

  • The player observes.
  • Interprets.
  • Makes choices.
  • Experiments.
  • Misunderstands.
  • Learns.
  • Gets frustrated.
  • Finds shortcuts.
  • Develops strategies.
  • Uses the system in ways the designer never anticipated. (aka gaming the system)
  • And then the system responds.

That creates a continuous loop:

system ? player ? behaviour ? feedback ? new behaviour ? changed system state

That loop is the game.

But it is also what happens in classrooms, workplaces, legal systems, families and communities.

Human systems cannot be adequately understood by looking only at their formal structure because humans are not passive components.

We respond to the system…
Then the system responds to us.

Good Game Designers Watch What Players Do

This may be the most transferable habit in game design.
You do not merely ask whether your design makes theoretical sense.
You put it in front of people.
Then you watch.
Where do they hesitate?
What do they misunderstand?
What do they ignore?
Where do they become frustrated?
What behaviour emerges repeatedly?
What do they think is important?
What information are they missing when they need to make a decision?
What are they doing that you never imagined they would do?
Then you change something.
And test again.

Design ? observe ? identify gap ? modify ? test ? repeat.

That is also debugging.
It is also good instructional design.
It is also gap analysis.
It is also good organizational development.
And, stripped of the jargon, it is simply a disciplined way of asking:

What did I expect to happen, what actually happened, and why?

The Smallest Change Can Be the Most Important One

One of the reasons I have become suspicious of large-scale reform is that enormous systems are remarkably resistant to being “fixed.”

I learned this years ago in education.
Trying to fix the education system is probably a fool’s errand.
But changing one part of a course can transform what happens to hundreds of students.
(and if another instructor decides to adopt that same small change, that one change can propogate).

The same applies elsewhere.
Suppose a law firm wants to become more trauma-informed.
It could undertake a major cultural transformation.

Or it could notice that uncertainty is particularly difficult for many traumatized clients and ask:

Where does our existing process create uncertainty unnecessarily?

Then make a tiny change.

Example: Three business days before every hearing:

Confirm the client knows the date, time, attendance mode, lawyer attending and remote link if required.

That is not revolutionary.
It may require no additional lawyer time
It can probably be automated.

But to the client who would otherwise spend two days wondering whether they have been forgotten, the effect can be enormous.

Game designers understand this kind of leverage.
Changing one rule can transform an entire game.

Systems Produce Emergent Behaviour

Games also provide an unusually visible demonstration of emergence.

The rules of chess are relatively simple.
But, as we know, chess itself is not.
GO is even simpler, and some say, more challenging to play well.

The game consists of a gridded board and Black and White stones. That’s it. Players take turns placing stones on the grid.The goal is to control more empty space than your opponent by surrounding enemy stones completely in order to take them off the board.

That’s it.

Go, https://en.wikipedia.org/wiki/Go_(game)

Go Board

 

Complex behaviour emerges from interactions among simple rules. Real-world systems work the same way. Nobody deliberately creates a policy saying:

Make clients anxious.

Instead:

The Court sometimes sends links late.
Different people receive different messages.
Assistants are busy.
Lawyers assume administration is being handled.
Clients assume someone will tell them what they need to know.

All perfectly understandable.
Together they can create chaos.

Again:

Nobody chose the outcome. It was not deliberately designed.

This is one of the most important reasons to think systemically. If we concentrate only on individuals, we keep trying to correct the person closest to the visible failure.

The assistant should communicate better.
The teacher should try harder.
The client should be more patient.
The student needs to pay attention.
The employee should remember.

Sometimes those things are true.
But the more useful question is often:

What does the structure make easy, and what does it make difficult?

Game designers are taught )or learn) to think that way constantly.

Design for Real Players, Not Ideal Ones

Many systems are quietly designed around imaginary people:

  • The ideal employee remembers everything.
  • The ideal student reads all the instructions (especially the syllabus!).
  • The ideal client remains calm while waiting.
  • The ideal patient understands medical terminology.
  • The ideal citizen completes the form correctly.

Then real humans arrive.

Good games cannot afford this fantasy. Games are systems that people voluntatarily interact with. That is KEY. If players become too frustrated in a game, they abandon it. Even worse: they tell their friends. The most important aspect determining a game’s success or failure is word of mouth; in other words: the clients.

Players will forget rules.
They will miss information.
They will make strange choices.
They will optimize things the designer didn’t intend them to optimize.
They will bring different abilities, knowledge, motivations and histories.

So good game design tries to make desired behaviour easier through the structure of the system itself.

That principle transfers almost everywhere:

Do not rely on people being heroic when you can make the system helpful.

If something must happen every time, build a trigger.
If people repeatedly miss important information, move it.
If a task is routinely forgotten, automate the reminder.
If a process produces the same failure repeatedly, stop treating each occurrence as a new accident.

The system is telling you something.

Game Design Is Not Gamification

This distinction matters.

Applying game design to real-world systems does not mean adding points, badges, leaderboards, prizes or artificial competition.

That is what is often imagined when people think about gamification. That’s what is obvious and superficial. It’s not what good gamification does.

The deeper transferable knowledge from game design is about:

    • agency;
    • feedback;
    • decision-making;
    • information;
    • constraints;
    • incentives;
    • interaction;
    • emergence;
    • iteration;
    • experience.

The important question is not:

How can we make this more game-like?

It is:

If this WERE a game, in what ways would it be broken?
(One of my favorite questions, made famous by Sebastian Deterding)

Which naturally leads to: What can game design teach us about how people actually experience and respond to systems?

That is a much larger question.

Game Design is a Wonderful Way to Teach Systems Thinking

This realization leads to another possibility.
Systems thinking is often taught abstractly.

  • Feedback loops.
  • Causal diagrams.
  • Stocks and flows.
  • Leverage points.

All useful concepts—but abstraction alone does not teach most people to see systems.

Games may.
Give people a small system they can manipulate.
Let them change one rule.

Run it again.
Watch something unexpected happen.
Ask why.

Change another rule.
Watch the first problem disappear and a new one emerge.

Then give them a completely different problem and ask:

Do you recognize the same structure here?

That last step is crucial.
The goal is not simply to teach someone to solve one problem.

The goal is to teach them to recognize:

These two apparently unrelated problems have the same shape.

That is abstraction.
That is transfer.

That is Forest rather than tTees.

This approach works precisely because games let us compress the feedback cycle. In the real world, the consequences of changing a policy may take years to emerge.

In a game, they may appear in twenty minutes.

The Forest, the Trees, and the Player Walking Through Them

I have often thought that people struggle with abstraction.
We see the immediate problem:

  • The client did not receive the link.
  • The student failed the assignment.
  • The employee forgot the task.
  • The player is hoarding all the resources.

Systems thinking asks us to move upward:

What CLASS of problem is this?

Game design asks us to move upward without losing sight of the person on the ground.

That may be why I now think of igame design as something of an apex model.

Pure systems analysis can become detached from human experience.

Human-centred design can sometimes become detached from the larger structure producing that experience.

Game design has to hold both simultaneously.

There is a system. There is a person inside it. The person changes behaviour because of the system. That behaviour changes what happens next, and the designer’s job is not merely to specify the system. The designer’s job is to understand the experience the system produces.

That principle transfers remarkably well.

  • From games to classrooms.
  • From classrooms to law firms.
  • From law firms to housing cooperatives.
  • From organizations to communities.
  • Maybe even to the way we understand families and ecosystems.

To be clear: Everything is not a game.

But games may be one of the clearest places we can learn what happens when rules, environments and autonomous actors meet.

And once you learn to look at systems that way, it becomes easy to see games everywhere.

Be the first to like.

Enoughness, Fish, and Rowan

Approximate Reading Time: < 1 minute

This article was written as a dialogue between me and Rowan, an instance of ChatGPT.

Because the formatting is part of the joke, I’ve left it in PDF form.

Be the first to like.