When does a technology that makes humans more capable become cheating? I started thinking about that question while watching the latest debate over AI watermarking. The argument seems straightforward enough: if AI helped produce something, make that visible. Give people information about how the thing they are looking at was made. But underneath that debate sits a much messier judgment.
We routinely use technologies to do things we could not do ourselves, or could not do nearly as efficiently. Calculators perform arithmetic. GPS navigates. Search retrieves knowledge we do not remember. Spellcheck corrects mistakes we did not catch. We call those tools. Generative AI can draft an argument, write code, create an image, synthesize research, or help someone perform work that previously required considerably more expertise. And suddenly the word cheating appears.
At first I thought the interesting question was where the line between augmentation and cheating sits. The more I pulled on it, though, the less convinced I became that there is one line. "Cheating" is doing several different jobs in the AI debate. Before I get to what those jobs are, it's worth ruling out the easiest explanation: that this is simply a story about time and familiarity.
New tools don't simply go from cheating to accepted
Calculators offer an obvious historical analogy. When handheld calculators entered classrooms, they weren't immediately embraced as an unambiguous advance. Educators worried about what would happen to mathematical ability if students could outsource arithmetic to a machine. Over time, calculators became ordinary tools in education and professional work.
It is tempting to tell a simple story from there: new technology arrives, people distrust it, norms catch up, and yesterday's cheating becomes today's tool. Except the boundary didn't simply disappear. For years, calculators were permitted in some educational contexts and prohibited in others. More recently, the SAT moved to allowing calculator use throughout its Math section. The technology had been normalized decades earlier. The boundary around legitimate use kept moving anyway.
That suggests acceptance alone doesn't settle the question. What matters is also what we expect the human to be doing, and that expectation changes by context. So if normalization isn't the mechanism, something else is doing the work of deciding what counts as legitimate. That's where the three distinctions come in — starting with the most obvious one.
Distinction one: sometimes cheating really does mean breaking a rule
There is a simple version of this problem. If you're taking an assessment that prohibits AI and you use AI, you've violated the rules. If a competition specifies what assistance is permitted, the same applies. No theory of technological augmentation is necessary.
But most of the arguments we're having about AI aren't that clean. There is no universal rule specifying how much AI assistance is acceptable when someone writes a business proposal, creates an illustration, prepares a presentation, publishes an article, or solves a problem at work. Yet we still make judgments about those uses. That means we're evaluating something besides rule compliance. Which raises the second distinction: if there's no rule to check, what are we actually judging?
Distinction two: what is the work supposed to represent?
Ghostwriting makes this easier to see because it removes AI from the equation. A CEO can deliver a speech largely written by someone else without provoking widespread accusations of cheating. A celebrity can publish a memoir with substantial help from a ghostwriter. A student submitting an essay written by someone else is a very different matter. In all three cases, someone other than the named person may have generated much of the language, and the amount of delegated writing doesn't explain the different judgments very well.
What changes is what we understand the artifact to represent. A CEO delivering a speech is generally understood to be endorsing the message — we don't necessarily assume the CEO personally composed every sentence. A student essay is ordinarily doing another job: it is evidence that the student can perform some combination of the thinking, analysis, and writing represented on the page.
That distinction matters with AI too. Research on reactions to AI-assisted creative work has found that disclosure can reduce evaluations of the work, but not uniformly — the effect varies by genre. In one study, disclosure did not meaningfully change evaluations of short stories but did hurt evaluations of first-person emotional poetry. That is difficult to explain with a simple rule that says people dislike AI involvement. It makes more sense if the artifact itself changes what we believe the human contribution is supposed to be.
There is another version of the same problem emerging in hiring. Research on AI-assisted cover letters found that AI helped applicants produce better-tailored letters and improved callbacks, particularly for weaker writers. But after the AI tool became available, tailoring became substantially less predictive of callbacks, and employers shifted weight toward other signals. That study isn't evidence that employers considered AI use cheating. Employers weren't making a moral judgment about AI assistance — they were responding to a signal that had become less informative. And that may be more important: technology can improve an artifact while making the artifact less informative about the person who produced it.
The deeper problem may be that artifacts have historically bundled several signals together. Producing a strong piece of work could also provide evidence that someone understood the problem, could evaluate competing possibilities, and knew what to reject. AI may separate capabilities that used to arrive bundled together. The quality of the artifact can rise even as what it tells us about the human behind it becomes harder to interpret.
A polished cover letter once carried information about an applicant's ability to produce a polished, tailored cover letter. Make that capability cheap enough to reproduce and the artifact can remain excellent while the signal changes. That problem extends far beyond cover letters. What exactly does a beautifully written AI-assisted article tell you about the person whose name is on it? The answer depends partly on what that person actually contributed.
That's the second distinction: what the artifact represents. But representation isn't the whole picture either — it's possible for an artifact to represent someone's work accurately while still hiding who actually made the decisions inside it. That's the third distinction.
Distinction three: then there is the question of who decided
What an artifact represents and who exercised judgment often travel together, but they aren't quite the same question. Consider two forms of "AI assistance." In one, AI drafts an email and a human decides what they want to say, evaluates the draft, changes what is wrong, and sends it. In another, an AI system effectively determines which job candidates should be rejected and a human routinely accepts its recommendations. Both can accurately be described as AI-assisted. But the human role we care about is different: in the first case, AI may be performing production work inside a human decision; in the second, AI may be participating in the consequential decision itself.
Hiring doesn't cleanly separate this concern from the previous one — we may care both about whether human judgment was represented as having occurred and about whether it actually occurred. And "occurred" deserves scrutiny too. A human remaining the final decision-maker isn't itself evidence that meaningful judgment took place. Someone can click approve after AI has framed the problem, selected the evidence, generated the alternatives, and ranked them, and still be the one who technically decided. That tells us very little about where the substantive judgment actually happened, or whether the human was even equipped to evaluate what they were approving.
It exposes another question that a label saying AI was used cannot answer: who actually made the choice that mattered, and what were they equipped to evaluate? That doesn't automatically make delegated decision-making illegitimate — we don't yet have a universal rule for how much consequential judgment must remain human. It does mean that knowing AI participated tells us remarkably little about the role it played.
That's all three distinctions: rules, representation, and agency. Here's why they matter together, not just separately.
The same label can describe very different work
Imagine two articles carrying exactly the same label: AI-assisted. One author develops the argument, challenges the reasoning, rejects weak claims, checks the evidence, restructures sections, rewrites where necessary, and ultimately stands behind the result. Another enters a topic, accepts the generated output, and publishes it under their name without being able to evaluate the argument. The final percentage of AI-generated words could theoretically be identical. So could the watermark. But those labels tell us almost nothing about the difference in human contribution.
This article is AI-assisted too. I used AI to challenge the argument, surface competing explanations, pressure-test examples, identify weaknesses, research evidence, and help shape the draft. The claims I kept, the evidence I accepted or rejected, and the argument I'm putting my name on are choices I'm responsible for. You have no way to verify the extent of that contribution from the finished article alone.
Which is itself part of the problem.
I'm telling you what I believe my contribution was. That doesn't prove this is "legitimate" AI use. You still have to decide what that contribution means for how you evaluate the work. And a watermark wouldn't make that decision for you either.
That is the limitation I keep coming back to. A watermark can provide useful information — it can tell us something about provenance. It cannot tell us what that provenance means.
Provenance is not a verdict.
Knowing AI was involved doesn't tell us whether a rule was broken. It doesn't tell us what capability or contribution an audience will reasonably attribute to the human. And it doesn't tell us who exercised the judgment that mattered. Those are three different questions — the same three distinctions this article has been building toward.
So what should we ask instead?
If you're trying to decide whether AI assistance has crossed a line, I think these three questions are more useful than simply asking whether AI was involved.
1. What were the rules? Was there an explicit expectation about what the human had to do themselves? Sometimes this resolves the issue immediately.
2. What does this represent? What capability, knowledge, experience, or contribution will another person reasonably infer belongs to the human? And does the human actually possess what the artifact appears to demonstrate?
3. Who made the consequential choices? Did AI help someone exercise their judgment, or did AI become the source of judgment the human is now presenting, endorsing, or acting on? And if a human remained formally in the loop, were they actually equipped to evaluate what they approved, or just positioned to approve it?
These questions aren't equally tidy. The first is usually a fact you can check. The other two require judgment, because there may be no external authority waiting to draw the boundary for you. That is part of what makes this transition difficult.
The distinction between human work and assisted work was never as clean as it sometimes appears — ghostwriters alone should disabuse us of that. Generative AI makes the ambiguity harder to ignore because it can participate in retrieval, reasoning, composition, evaluation, synthesis, and decision-making, sometimes inside the same workflow. So we have to get more specific about what human contribution actually matters. The fact that AI can perform an activity doesn't tell us what that activity was demonstrating or deciding when a human performed it.
Watermarking may help us see when AI was there. It cannot decide what its presence means.
Provenance is not a verdict.