AI influences us all today, both directly through the use of generative AI and indirectly through algorithms that act as recommendation systems in our lives.
Yesterday, I could have asked ChatGPT a question, read a Google summary influenced by AI, accepted a Grammarly suggestion, watched something recommended by an algorithm, and then sat down and written something completely in my own words.
Now hand that writing to an AI detector.
What exactly are we asking it to detect?
That is what I mean when I say AI detectors stopped working yesterday. Not that every detector suddenly became technically useless, but that the distinction they are trying to make is becoming increasingly blurry.
AI usage reminds me of a time in middle school. We were learning how to create bibliographies, this was before AI, and I asked my teacher how to cite my parents as a source. Of course, my teacher explained that I wouldn't cite them the same way I would cite an academic source.
But that didn't mean they hadn't influenced how I thought.
We have always absorbed ideas from teachers, parents, friends, books, and conversations. AI is quickly becoming another layer of that environment, except its influence is much harder to trace.
A student could brainstorm with AI, research through AI-powered search, rewrite an AI-generated explanation in their own words, use Grammarly to edit it, and then substantially revise the final essay themselves.
Is that human-written? AI-written? AI-assisted?
At some point, the binary starts telling us less than we think.
Once we accept that AI is being used, we can start looking at what actually matters: students' thinking, learning, and effort, instead of turning essays into a probability experiment.
Why the human-versus-AI binary is becoming less useful
When ChatGPT first became widely available, schools were forced to respond quickly. Many banned it. Then came AI detectors.
The idea made sense. If students were going to use AI when they weren't supposed to, maybe software could tell us who did it.
But cracks appeared quickly.
Research raised concerns about false positives, including the well-known Stanford study that found AI detectors disproportionately flagged writing from non-native English speakers. Students also realized something simpler: AI-generated text could be rewritten, edited, or mixed with their own writing.[1]
New detectors will continue improving. Some may become very good at identifying certain forms of AI-generated text.
But even if tomorrow's AI detectors became dramatically more accurate, we'd still have the same educational problem.
One student might use AI to understand a difficult concept, write their own outline, ask for feedback, and then revise the final essay themselves.
Another might generate the entire assignment in thirty seconds and change a few sentences.
Both students used AI.
Educationally, those situations are completely different.
That is why the human-versus-AI binary is becoming less useful.
AI Didn't Create the Problem. It Exposed It.
Most responses to AI in education seem to fall into three categories:
- Incorporate AI into learning.
- Redesign assessments to make AI harder to use.
- Create clearer rules about acceptable and unacceptable AI use.
There is value in all three, but they often share the same assumption: AI is the thing education now has to solve.
I see it differently.
AI did not necessarily create education's weak spots. It exposed them.
For a long time, we could look at a polished essay and reasonably assume the student had gone through the intellectual process required to produce it.
They probably researched. They probably struggled with the argument. They probably rewrote bad ideas before producing better ones.
The final essay acted as a reasonable proxy for what happened before it.
Generative AI broke that assumption.
Now, two students can submit similarly polished essays while having gone through completely different processes.
Maybe AI has simply forced us to confront a question that existed all along:
How do we actually know that learning happened?
The Problem Isn't AI Use. It's Outsourcing the Thinking.
AI use itself isn't inherently bad.
AI becomes dangerous when it takes the struggle out of the production of writing.
Writing is not valuable because students physically type 1,500 words.
A lot of the learning happens in the difficult parts: figuring out what you actually believe, realizing a source doesn't support your argument, reorganizing an idea, or changing your thesis after the evidence changes your mind.
AI can help someone through that struggle.
It can also remove it completely.
The same AI use can also have very different relationships with learning depending on the student. The figure below shows that the relationship between ChatGPT use and post-test performance changes across students with different levels of prior performance. That is part of why simply measuring whether a student “used AI” is too blunt: the more important question is how that use interacted with the student's learning.[2]

A strong student might use AI to challenge an argument or clarify a concept. A struggling student might use the same tool to avoid forming the argument in the first place.
Both "used AI."
One may have enhanced the learning process.
The other may have bypassed it.
So I don't think the most important question is:
Did the student use AI?
I think it is:
Did the student still do the thinking?
How to Redesign Assessment Around Thinking, Not AI
If we accept that almost everyone is influenced by AI to some degree, then we have to rethink what we measure.
From a high level, I think professors should structure their courses around two things that reinforce each other:
learning opportunities and verification checkpoints.
Take-home assignments, papers, projects, and research can function as learning opportunities. They give students time to struggle with material, use resources, ask questions, and build understanding.
Then there should be moments where students have to demonstrate that understanding more independently.
Something like an in class essay where the prompt is revealed on exam day doesn't hold the same learning elements as the learning elements but works great as a periodic indicator into students progress.
The format matters less than the relationship between the two.
If students understand that honestly completing the take-home work prepares them directly for something they will later have to demonstrate themselves, the incentive changes.
Instead of only asking:
How quickly can I finish this assignment?
they also have to ask:
Will I actually understand this later?
But this still leaves a gap.
Professors don't want to discover on exam day or during an oral defense that a student never understood the material.
Ideally, they would know earlier.
That is where the writing process becomes important.
What Happens Before the Final Essay?
With most writing assignments, professors see two things:
the prompt and the final submission.
Everything else disappears.
The student researches, struggles, rewrites, changes their mind, finds better sources, and eventually compresses all of that into one polished document.
An AI detector tries to work backward from that final document and estimate what may have happened before it.
I think we should increasingly do the opposite.
Instead of trying to infer the entire process from the final essay, we should give educators more evidence about the process itself.
That doesn't mean surveilling students or pretending that process evidence proves learning.
It means giving teachers more context.
How CreativeTrail Can Make the Writing Process More Visible
That is one of the problems we are trying to solve with CreativeTrail.
Students complete their work through stages such as planning, working with sources, drafting, and revising. Along the way, instructors can see process insights such as time spent working, revision patterns, responses to Socratic questions, and writing authenticity information showing where content was transcribed or copied and pasted.
None of these signals individually proves that a student learned.
The goal isn't to replace the AI detector with another score that claims to know exactly what happened.
The goal is to give the instructor transparency into the students working. Messy working shows thinking went into writing it. A clean process raises red flags.
If a polished essay looks excellent but the process suggests the student struggled, that doesn't automatically mean misconduct occurred.
It can start a conversation:
Walk me through how you got here.
Likewise, an imperfect final essay might hide substantial research, revision, and development that would otherwise be invisible.
CreativeTrail is intended to give professors more visibility into what happened between the assignment being given and the final work being submitted, while there is still time to support the student.
What Should Educators Measure Instead of AI Use?
As we become increasingly intertwined with AI, trying to perfectly determine whether it influenced a final piece of writing becomes less useful.
Even a perfect AI detector would leave us with the more important question:
Did the student actually learn?
Use take-home assignments and other work as learning opportunities.
Connect them directly to verification checkpoints where students have to demonstrate what they understand.
And, when possible, give educators visibility into the process between those two moments.
AI didn't create the problem of confusing polished work with learning.
It just made the gap impossible to ignore.
The goal shouldn't be to build a perfect system for detecting AI.
The goal should be to build a better system for seeing learning.
