“It was a bright cold day in April, and the clocks were striking thirteen.” This is my favourite opening sentence in any novel, which you may recognise as being from George Orwell’s 1984, which was published in 1949. This dystopian novel involves a society with mass surveillance, secret police and a political class that manipulates facts to try and control the population.
I mention this book because I used it in a little experiment. Ever since the emergence of generative AI, people have used large language models to produce everything from template emails to essays and more. Very quickly there was a debate about how difficult it might be to tell whether a student’s homework had been written by the student or was merely the product of a large language model prompt. The use of LLMs in academia is widespread. In July 2026 a professor at Brown University in Rhode Island was suspicious when his class achieved an average score of 96% in their midterm test, far above the historical norm. He switched the final exam to an in-person, closed book test, and the result plummeted to 48.6%.
An industry of tools appeared to help teachers detect AI-written text, a sector estimated to be worth $600 million in 2025 and growing at 28.8% annually. But how accurate are the detectors? There has been plenty of evidence of problems with the tools, amongst which is a tendency for them to flag certain styles of writing, such as that of non-native English speakers, as AI written. Still, a study in June 2026 that examined four leading products (Turnitin, Pangram, Copyleaks, GPTZero) found that the number of false positives in their test had improved significantly compared to earlier research, and that “fully human text was reliably classified”. This was very different from the situation a year or two ago, when multiple studies found problems with accuracy of such tools, with some even flagging the US Constitution as being AI-generated.
So, all is well then? Not so fast. I was curious to test out Pangram, which generally seems to be the best regarded of these tools, and whose vendor claims just 1 in 10,000 false positives and 99.98% accuracy. This sounds impressive, but just to be sure I copied in the opening thousand words of my favourite Orwell’s 1984 (using Project Gutenburg, a digital library of many classic texts that are out of copyright). Pangram reckoned this was 77% human-written and 23% AI-written. Given that 1984 was published in 1949, it seems pretty unlikely that George Orwell (real name Eric Blair) had access to an AI. To be fair to Pangram, I then tried some other Orwell text from another novel, and ten other similar-sized pieces of text from assorted classic books that were freely available online, and Pangram marked them all as 100% human-written. Nonetheless, the very first piece of text that I ever loaded into Pangram received a 23% AI score, despite it indisputably being written by George Orwell. Maybe this was just remarkably lucky on my part, but it suggests that Pangram is not quite as flawless as its advertising suggests.
I played around a little further with the text. When I broke the thousand-word Orwell text up into chunks and tested each separately, it turned out that an 822-word piece of text was marked as 81% AI-written. Of this 822-word piece, a 400-word chunk was marked as 31% AI but the other 422-word chunk as 100% human. If I simply switched the order of the two passages and put them back together, so still the identical 822 words but with the paragraph sequence different, it was flagged as 100% human-written by Pangram. I don’t know why this is, but you are free to try for yourself to reproduce the results. I am unaware of the exact mechanisms that Pangram uses, so I don’t want to speculate as to why this error occurred. However, the result does suggest that caution needs to be exercised when the product declares something to be AI-generated.
There is a certain irony that a novel about manufactured truth being pronounced “artificial” by a machine built to police manufactured truth. However, if Pangram can’t reliably tell the writing of George Orwell from that of a chatbot, just how reliable is it when assessing the rest of us? This is important since accusations of cheating on homework by using AI can have serious consequences for students, and indeed any author, whether of novels, magazine articles, or anything else.







