What Does the Work Prove?

In brief: AI-assisted writing creates a problem that goes deeper than authorship. We've historically used writing as evidence of the person behind it: what they understand, how they think, and what they can do. AI makes that inference harder. But treating suspected AI involvement as a shortcut creates a different problem: provenance starts standing in for judgments about quality, contribution, and capability that it cannot answer on its own.

Anyone else feel like they're living in the Kobayashi Maru?

For the non-Trekkies, it's a test designed so you can't win. Captain Kirk famously dealt with that by reprogramming the test so he could. I'm starting to understand the impulse.

I'm applying for roles that expect me to leverage AI: automate content, build AI-enabled workflows, scale what I produce. So: use AI.

Except I'm also a writer, and apparently that complicates things.

Use AI to build something you couldn't otherwise build and you're leveraging the technology. Use AI while writing something and people may start wondering whether you really wrote it.

And being resentful about that fact is probably an understatement. Especially because I've used AI to do some pretty complex coding as well as building my own website (arguably not as complex, but I wouldn't have attempted it without AI).

I'm not a frontend developer. I cannot code from scratch. I can poke around in code, particularly with AI helping me figure out what I'm looking at, but there is a substantial amount of implementation on that site that I could not produce independently.

The code most certainly did not come from my brain. And guess what? Nobody cares.

Nobody is asking which lines I personally typed before allowing me to say I built the site. Nobody is running the source through an AI detector to determine whether I'm fraudulently taking credit for the AI's fancy CSS and JavaScript.

But use AI to help me write, in the domain where I actually do have expertise, and suddenly the provenance of individual words can become evidence that I didn't really do the work.

That seems backwards at a bare minimum. One might even say asinine. In fact, I was pretty happy calling it a very cruel double standard.

Why does AI expand what someone can create in one context but seem to replace their judgment and diminish their contribution in the other?

Then I made the "mistake" of thinking about it more.

Because if the problem were simply that AI performed work a human otherwise would have performed, we should be every bit as suspicious of my website. I would argue even more suspicious. AI is supplying a capability there that I genuinely don't possess.

So why aren't we?

Maybe AI involvement isn't actually the thing we're trying to detect.

So why does the website feel different?

When I say I built my website with AI, I'm not claiming I could reproduce the underlying implementation without it. Obviously I couldn't. That was rather the point of using AI.

But I'm not absent from the process either. I know what I want the site to do and how I want the information organized. I can see when the result is wrong, ask for changes, compare alternatives, notice when something breaks, and decide whether what I'm looking at actually accomplishes what I intended.

My ability to judge the result also has boundaries. Something could look and behave perfectly well from my perspective while the underlying code is an abomination. I could miss an accessibility problem, a performance issue, a security vulnerability, or a maintenance nightmare because I don't have the expertise to recognize it.

AI hasn't made me a frontend developer. But we seem perfectly capable of understanding that I can supply the intent, requirements, decisions, information architecture, and evaluation while AI supplies implementation capability I don't independently possess.

And you don't need the code to prove something about me before deciding whether the website is any good. You can use it. You can see whether the navigation works, whether you can find what you're looking for, whether links go where they should, and whether the site does what it appears to have been designed to do.

The AI-built artifact is least controversial precisely where it reveals next to nothing about my ability to produce its underlying implementation.

Nobody looks at my website and concludes that I must be a frontend developer. But what if I were asking them to?

If I put that same website in a portfolio and presented it as evidence that I should be hired as a frontend developer, suddenly I suspect people would care considerably more about which capabilities were mine and which came from AI.

Maybe the difference was whether we were evaluating the artifact for what it does or using it as evidence of what the person who made it can do. That seemed like the distinction.

Writing has been doing two jobs at once

Writing can be evaluated independently of its author. You can decide whether an explanation is clear, whether an argument holds up, whether the evidence supports the claim, or whether something is enjoyable to read without knowing anything about the person who wrote it.

But we've also been asking writing to do another job.

If I read a clear explanation of a complicated technical concept under someone's name, I'm not only evaluating the explanation. I'm learning something about the writer, or at least I think I am.

I assume they understand the subject. I may infer that they can identify what matters, structure an argument, distinguish a useful qualification from a distracting one, anticipate confusion, and communicate all of that clearly.

That inference has never been foolproof. Ghostwriters, plagiarism, heavy editing, and people producing wonderfully polished nonsense all predate AI. Still, ten years ago, the fact that someone had produced a really good explanation of a difficult technical concept was reasonably useful evidence that they possessed at least some of the capabilities visible in it.

Not proof. But not a terrible assumption either.

AI weakens an inference we used to make almost automatically: that the capabilities visible in the writing belonged to the person whose name was on it.

That makes the portfolio explanation pretty satisfying. An ordinary visitor to my website is judging the site, not using it to judge my ability to build one. A hiring manager looking at the same site as evidence that I'm a frontend developer would be doing both.

Except even that isn't quite enough.

An engineer's manager is also judging capability and may be perfectly comfortable with extensive AI assistance. One feature doesn't have to prove everything about that engineer. The manager has design discussions, code reviews, debugging, technical decisions, previous work, and repeated experience to draw on.

If you're reading this and you don't know me, you may have none of that. You may have read other things I've written, seen me argue about something in the comments, or looked through the rest of this site. Or this may be the first thing of mine you've ever encountered.

The less independent evidence you have about my capability, the more evidentiary weight the work itself has to carry. At the extreme, you have the artifact.

And we ask artifacts like this to carry an astonishing amount of information about the person behind them. Does the author understand the subject? Can she reason through a complication? Recognize when her own explanation fails? Communicate the result clearly?

Historically, the writing itself gave you at least some evidence for those things. AI makes that evidence harder to interpret because it can now produce some of the same visible signals.

That doesn't mean AI involvement proves those capabilities aren't mine. It means the artifact alone may no longer be enough to establish that they are.

If an engineer were taking a coding assessment, where the artifact had to carry much more of that evidentiary load, I suspect we'd care considerably more about what AI contributed.

AI may be most destabilizing when a single artifact is doing too much evidentiary work about the person behind it.

Diagram titled 'Independent Evidence and Evidentiary Load,' showing a continuum from more independent evidence (repeated work, discussions, decisions, explanations, history) to less independent evidence (résumé, portfolio, writing sample, single encounter), with evidentiary load on the artifact increasing as independent evidence decreases.

That seems like it should solve the problem.

It doesn't.

Suppose I've worked with a writer for years. I've watched her reason through problems, challenge bad assumptions, explain her decisions, respond to feedback, and consistently produce good work. I have plenty of independent evidence that she knows what she's doing.

But I still have to care whether other people think her writing sounds like AI.

Because the reader doesn't have the evidence I do. They encounter the writing, notice patterns they associate with AI, and infer something about how it was produced.

There's a strange feedback loop hiding in that process. AI learned statistical patterns from human writing. It reproduces those patterns. We start recognizing some of them as signs of AI. Then human writers start avoiding them because they've become signs of AI. And the standard moves again.

Em dashes. 50 cent words. Parallel constructions. Not-X-but-Y contrasts. Tried-and-true rhetorical techniques. We've turned perceived similarity to AI into a writing standard and made centuries-old techniques suspicious.

Underneath it all seems to be the idea that real writers shouldn't need AI.

I don't buy that.

Expertise isn't defined by refusing tools. It should make you better at using them. A good writer knows when a sentence is technically fine but wrong, when an argument has flattened, when the voice has disappeared, when something sounds polished but says almost nothing.

The skill isn't proving you produced every word unassisted. It's knowing what should be there when you're done.

None of that means the provenance inference is necessarily wrong. Sometimes the writing really was produced with AI assistance. Sometimes the patterns people think they recognize really are there because AI was involved. But they don't tell us how much the writer understood, contributed, evaluated, rejected, corrected, or changed.

And even when the provenance inference is correct, something else can happen.

Once people believe AI was involved, that belief can change how they evaluate both the work and the person behind it. In a 2026 experiment involving 547 participants, AI-assisted workplace communication was judged less trustworthy and authentic than human-only communication, even when AI had been described as helping the writer refine and present their own thoughts. [Sahebi, Formosa & Bankins, 2026]

Once we think we've recognized the provenance, it can change what happens next.

"This sounds AI-generated" can become a proxy verdict before the work has been independently evaluated. Sometimes it even stops the evaluation entirely. The reader stops reading.

Which means there really is a problem here — just, unfortunately, not quite the one I started with.

What we're actually trying to recover

We're not only policing how writing gets made. We're trying, however imperfectly, to recover information about the writer that the finished artifact no longer gives us as reliably.

That's a legitimate problem. And it has consequences beyond whether a reader gives an essay a fair hearing.

Sometimes we use the artifact to make decisions about the person behind it. Hiring is an obvious example. A hiring manager may have a résumé, a portfolio, a writing sample, and perhaps an hour or two of conversation. That's not much independent evidence of how someone actually thinks and works, which means those artifacts have to carry a lot of evidentiary weight.

But this doesn't stop once someone gets the job. Most knowledge workers write, whether or not anyone calls them writers. Emails, presentations, proposals, strategy documents, analyses, and project plans all become evidence other people use to judge what their colleagues understand and can do.

Research summarized by The Wall Street Journal has also found that perceived AI involvement can affect how people judge their colleagues' competence and how genuine they find their communications, with some findings suggesting AI use can even trigger judgments about laziness and morality — the same inference problem, just with a different person doing the judging.

And the questions matter:

Is the artifact any good?

What did the person contribute to making it?

What does it tell me about what that person understands or can do?

I've written elsewhere about why production provenance alone can't answer the second question. But the problem here goes further than that. Once writing triggers the suspicion that AI produced it, perceived provenance can start answering all three before they've actually been asked.

Diagram titled 'From Perceived Provenance to Proxy Verdict,' showing perceived provenance ('This sounds AI-generated') leading to a proxy verdict that branches into three separate judgments: quality (is the artifact any good?), contribution (what did the person contribute?), and capability (what does it tell me about what the person understands or can do?). Caption reads: one inference is standing in for three separate judgments.

The writing is slop.

The person didn't really write it.

The capabilities visible in it aren't really theirs.

Those are three different conclusions. They don't follow from one another.

That's why I keep getting hung up on the phrase "AI slop." "This was made with AI" tells me something about the process. "This is slop" is a judgment about the result.

Sometimes those absolutely coincide. AI is great at producing terrible writing, and people can publish it without understanding, evaluating, or apparently even reading what came back.

But detecting AI involvement doesn't establish any of those things. And believing we've detected it from the writing itself certainly doesn't establish them.

Those are exactly the judgments we were supposed to be making independently.

"Who wrote it?" is the wrong first question

This is irritating because I started with a much simpler complaint. I wanted the double standard to be the thing to call out and address. I don't think it is.

What I missed was how much evidentiary work we've historically expected writing to do for us, and how easily suspected provenance can stop us from independently evaluating that work at all.

"Who wrote it?" isn't necessarily a bad question. The problem is what happens when we let the answer stand in for all the questions that were supposed to come after it.

This matters precisely because the underlying judgment hasn't gone away. A hiring manager still needs to know whether a candidate understands the subject and can do the work. A colleague still has to decide what someone's work tells them about their contribution and capability. A reader still has to decide whether an author's judgment is worth trusting. AI has made some of the evidence behind those judgments harder to interpret.

Sometimes the consequence of the proxy verdict isn't that someone stops reading your essay. It's that they stop considering you.

Treating suspected AI provenance as a shortcut doesn't recover the missing evidence. It just substitutes one uncertain inference for another.

I don't want to wave the underlying problem away. If two people can produce similarly impressive writing while bringing radically different levels of understanding and expertise to it, we do need some way to distinguish between them. But the shortcut still isn't.

"This sounds like AI" doesn't establish that the writing is bad. It doesn't establish what the human contributed. And it doesn't establish what the person understands or can do.

Maybe the writing is terrible. Maybe the person contributed almost nothing. Maybe they couldn't explain or defend a single idea in it if challenged. Those are all possibilities worth caring about — the shortcut just still isn't how you find out which one is true.

But if we want to know what the work proves, we have to stop treating suspected provenance as proof.