The revolt of the reader

Reading is important to me. While I’m not a quick reader nor an especially voracious one, I have found that long-form reading has had a profound influence on me over my life.

And because reading is important to me, writing is too: writing not only allows us to convey our ideas, but the very act forces us to test and distill them — at once making our ideas more robust while providing the vehicle by which to share them.

But I am writing this now because — speaking as a reader — we are exasperated: too many people — people we otherwise like and respect! — are writing (or otherwise putting their name to) pieces that are clearly LLM-authored. We readers are left with pointed questions for those promoting LLM-authored pieces: do you think readers can’t tell? Or do you think readers don’t care?

To answer the first question, readers can absolutely tell. To those who read broadly, the hand of the LLM is so clear it’s as if the writer’s intellectual fly is open. In fact, it’s so jarring that I have to believe that those writing with LLMs are either not reading enough to see the LLM’s obvious structural tells — or (and?) they aren’t even reading their own content. (A confession: with particularly egregious pieces, I have fantasized about sentencing the author to read them aloud, certain that they themselves will be unable to endure the slop that they are foisting upon the rest of us.)

As to the second question, readers emphatically care. In fact, the tells of LLM writing are so grating ("and here’s why that framing matters!") that our brains pull an LLM-triggered ejection handle, bailing us out mid-sentence in an act of self-preservation. And the first person plural there is deliberate: as Cynthia Dunlop writes, of the 668 developers that replied to her survey, 78% "stop reading immediately" when they detect an LLM. And there are consequences that outlast the piece: 71% of the respondents in Cynthia’s survey also "avoid the author in the future" (!!). Revealingly, the respondents are not after linguistic perfection, but rather authenticity: 98% reported preferring an author’s own (imperfectly) written piece over an LLM-polished one. Finally, be wary of dismissing Cynthia’s respondents as a self-selecting group: active readers on social media are exactly the folks most likely to repost or otherwise promote writing they like — the early adopters and the tastemakers of online prose.

Why do people have this reaction? Beyond having to endure aggravating stylistic tics, when reading a piece that has had substantial LLM assistance, we — the readers — don’t know what is real and what isn’t. As I wrote in RFD 576, to use an LLM to write is to void the social contract between writer and reader: we readers shouldn’t be expected to labor to understand a sentence that the writer themselves didn’t work to create.

Does the revolt of the reader matter? If it needs to be said, when you are using an LLM to author a public piece, you are no longer writing for yourself or to otherwise pressure-test your own ideas — the only purpose is to serve the reader. If readers choose to ignore you (perhaps forever!), you will have undermined yourself: instead of attracting readers you will be actively repelling them. So other judgement about its use aside, using an LLM will increasingly become simply…​ ineffective.

In this regard, I am reminded of the arc of e-mail spam. There was a time in the early 2000s when people were (reasonably!) afraid that the explosion of spam would mean the end of e-mail. This was an era largely before social networking, and e-mail was the canonical killer app of the Internet; that we were losing e-mail to spam felt deeply dispiriting. But sometime in the late 2000s, we turned a corner: spam filtering improved so much that the economics of spam were undermined. Moreover, as spam became effectively contained, the consequences of being labeled as spam became increasingly dire. Today, legitimate businesses are very careful about how they use bulk e-mail; to be labeled as spam is to effectively destroy your domain name and tarnish your brand.

The war on spam started to turn when we could identify it at scale; could something similar happen to LLM-authored writing? Like spam, an LLM’s influence is readily identifiable to humans reading it; surely this is a solvable problem?

Up until somewhat recently, the results on this problem had been decidedly mixed. I had tried to use LLMs themselves for LLM identification, but I found that their false negative rate was far too high: they were chipper in accepting stuff that I was certain was LLM-authored. Other services seemed to look for basic LLM tells, but as an avid (and unapologetic!) user of the em-dash, these superficial techniques make me shift nervously in my seat (and I found them to be so broadly unreliable that they didn’t earn regular use).

Late last year, Pangram Labs launched their Pangram 3 model. I found it to be a huge leap over other detectors, and became an avid user. Importantly, over months of use (and especially on samples that I otherwise knew the origins of), I found its false positive rate to be very low: when Pangram identified a text as being largely AI written, I could say with some certainty that an LLM was heavily involved. (I found its false negative rate to be higher than I would like, but it was a small price to pay for a low false positive rate.)

A little over a month ago, they introduced Pangram 4, which I found to be a step-function improvement over the already-impressive Pangram 3: in my experience it has an astonishingly low false positive rate and low false negative rate (which is to say: very high accuracy!). I am finding it to be so effective (and the loss of trust in voices that use LLMs to be so precipitous) that I recently extended RFD 576 to be explicit about public writing, specifically mandating that public Oxide writing be reported by Pangram as human-authored. As I explained in the RFD, the standard for our public writing is not merely that LLMs aren’t used to write, but that readers have the confidence that it’s not LLM-authored — and increasingly that means being Pangram-clean. For those organizations that value the authenticity of their institutional voice, I would encourage the adoption of a similar policy for those writing under their banner. (Looking squarely at you, Rust Foundation!)

Indeed, Pangram has become important to so many of us that I was thrilled when Pangram Labs co-founder and CEO Max Spero joined us recently on Oxide and Friends. I would recommend that discussion not only for the technical details of Pangram, but also for Adam’s drop of a pitch-perfect (if underappreciated in the moment) Simpsons reference: when I likened LLM-authored writing to presenting takeout food as one’s own cooking, Adam quipped that such writing is like "steamed hams", making reference to Principal Skinner trying to pass off Krusty Burgers as an old family recipe. Beyond the apropos reference, it was interesting to hear Max’s insight into what has made Pangram 4 especially effective — and why its efficacy seems to evince some suspicious howls of protest.

So, writers beware: readers are in revolt. You should fully expect your writing to be run through Pangram. If your position is that we should be fine with an LLM crafting prose from your prompt, spare us all the wasted cycles and just give us your prompt. Or, better yet, consider doing what generations of writers have done before you, and treating that prompt as a skeleton that you use to write your piece yourself! This doesn’t mean that an LLM can’t help you, of course — LLMs are superlative editors! — just that you probably shouldn’t let it write it for you if you actually expect the rest of us to read it.