When generating organized prose, synthesizing sources, and structuring arguments shift from high-friction, laborious tasks to instant outputs, the traditional grading model shatters. Evaluating students on their ability to execute the manual labor of thinking when a machine can do the heavy lifting is the equivalent of grading accountants on their manual ledger-addition speed after the invention of Excel.
I use ChatGPT all the time during a lot of my creative process. To me, at this point, it’s like having an Ironman suit for a lot of things related to research and writing. It can be like having a lightspeed librarian on tap at all times. We’ve let the cat out of the bag. The cat can run a hell of a lot faster than we can, and the cat took the bag with it. Not so ironically, most LLM users are no more skilled than Tony Stark was when he first strapped into his earliest suit. He was clumsy and dangerous. The suit multiplies Tony Stark; it does not supply his curiosity, engineering sense, values, goals, or capacity to recognize when the suit is malfunctioning. In unskilled hands, an LLM can produce fluent nonsense at extraordinary speed. The danger is not that people will use it. The danger is that they will mistake generated language for understanding. And I get that. But what about those students who are as skilled with LLMs as Tony Stark was with the suit by the last Avengers: Endgame?
It has demonstrated to me that the old education and grading model we’ve been using over the last century or so depended greatly on the absence of a tool like this. And yet, here we are. Standing at a crossroads. Why would we not want our students using these things?
Some might blather on about LLMs keeping us from doing all of the heavy lifting of learning the process or understanding how whatever it is that is being taught functions in the real world. But maybe we have come to that point. Maybe we no longer need to understand how things worked/functioned in a world absent of LLMs simply because we still can.
I do recognize that at this time, pretty much all of our teaching faculty is teaching from a pre-LLM model of education. And that is a problem in this new age that we have now entered. It is akin to someone still wanting to keep running the Pony Express after the wires crossed the continental U.S. While riders already out on the trail took a couple of weeks to finish delivering the mail that was already in transit, the operational life of the Pony Express ceased the very moment the wires crossed the continent. Our teachers are, functionally, those riders who were sent out the day before the final wires were connected, and today is the day after. We no longer make every accountant calculate columns by hand before we turn over the keys to Excel, or make every navigator use celestial navigation before allowing them to use GPS.
And to me, this is where education feels stuck. Many instructors are not unintelligent or malicious; they were trained to teach, evaluate, and identify mastery in a world where producing organized prose, summarizing sources, and answering conventional questions required substantial individual labor. LLMs have made those outputs cheap. The old signals of effort and competence no longer mean what they meant.
Clearly, we are at a critical dividing line. The student who treats the LLM as an oracle to do their thinking for them will indeed stunt their own cognitive growth. But the students who treat it as an exosuit—amplifying their own intent, iterating wildly, and applying rigorous human judgment to the output—are operating on an entirely different plane.
If education insists on clinging to the Pony Express model of evaluating academic work by hand rather than evaluating how well a student can pilot the technology toward an authentic, creative, and critical destination, it risks leaving students entirely unprepared for the world they are actually graduating into.