SlopShape Finds AI Writing Patterns That Survive Paraphrasing in Commercial Content
A new preprint finds that the organization of commercial web content can reveal AI-generated posts more reliably than wording alone, including after paraphrasing.
A new research preprint suggests that the structure of a web article can be a powerful signal of AI generation, even when its wording has been rewritten. SlopShape: Identifying AI-Generated Commercial Web Content reports 98.0 macro-F1 classification performance on held-out company domains using structural signals alone. For teams publishing at scale, the finding matters because it shifts attention from individual AI-sounding phrases to the repeatable ways content is organized.
The study examines commercial blog content rather than fiction or short-form prompts. Its authors, including Jochen Madler of Sitefire, built a corpus of 2,250 human-authored blog posts published before ChatGPT across 268 company domains. They then created 11,250 AI-generated versions using five frontier models, producing a 13,500-post dataset. The SlopShape preprint on arXiv describes the dataset, detection instrument, prompts, code, and aggregated artifacts made available for review and reuse.
The paper is a preprint, so its reported results should be understood as research findings rather than a settled industry benchmark. Still, its methodology raises a practical question for businesses using AI in content operations: are editorial workflows checking only wording, or are they also checking whether published pages are becoming structurally repetitive?
Why article structure can reveal AI-generated content
SlopShape defines structure more broadly than headings or paragraph count. The researchers analyze the order in which information appears, the kind and placement of evidence, and the article's voice. They developed an 11-dimension commercial template schema that generated 214 validated features, 187 of which were identified as structural signals relevant to detection.
This approach reflects a simple editorial reality. Two articles can cover the same topic and use entirely different words while still following a similar sequence: define a problem, present a neat list of benefits, insert generalized support, address a predictable objection, and end with a standard recommendation. That sequence may be useful in moderation. But when it becomes highly uniform across a site, it can make a content library feel formulaic.
The study reports that AI-generated posts occupied a comparatively tidy, formulaic structural space, while human posts appeared in rarer configurations. In other words, the signal was not a single forbidden phrase or sentence construction. It was the combined pattern created by many choices about sequence, emphasis, evidence, and voice.
| Detection approach | Primary signal | What SlopShape reports | Practical consideration |
|---|---|---|---|
| Word-focused detection | Surface wording and phrasing | The paper frames structural signals as potentially more durable when wording is changed. | Rewriting language alone may not meaningfully alter an article's underlying template. |
| Structure-based detection | Information order, evidence placement, and voice | 98.0 macro-F1 on held-out companies using structure alone. | It can help reviewers identify repeated publishing patterns, not just individual phrases. |
| Structure-based detection after paraphrasing | The same structural feature set after every AI post was paraphrased | 98.1 macro-F1, which the authors describe as essentially unchanged. | Paraphrasing is not necessarily a reliable way to remove structural regularities. |
Paraphrasing did not erase the reported signal
The paper's most consequential result for content workflows is its robustness test. The researchers paraphrased every AI-generated post with the model that created it, then reran the analysis. Performance was reported at 98.1 macro-F1, essentially unchanged from the original 98.0 result.
That finding does not mean every rewritten AI article is detectable, nor does it prove that a detection score establishes authorship in an individual case. It does show that, in this dataset and experimental setup, rewording did not remove the structural patterns the classifier learned to recognize.
SlopShape also reports 79.3 macro-F1 when assigning posts to their source among six classes: human-authored content and five AI model sources. The chance baseline was 16.7 percent. This result points to possible model-specific structural fingerprints, although it should not be interpreted as a universal attribution capability outside the study's corpus.
What content teams can do with this research
Businesses do not need to treat AI detection as a publishing gate that automatically accepts or rejects articles. A more useful application is editorial quality control. Content may be accurate, helpful, and AI-assisted while still relying too heavily on one predictable format. The problem to manage is not AI use by itself. It is whether a production process repeatedly produces thin, generic, or insufficiently differentiated pages.
A practical workflow can include the following steps:
- Audit a sample of published pages for recurring article sequences, repeated evidence patterns, and formulaic conclusions.
- Create multiple approved content structures for common page types, rather than relying on a single brief or prompt template.
- Require subject-matter input and source checks where claims, examples, or recommendations need real-world expertise.
- Review articles for information value, including whether they contain specific evidence, useful context, and a clear reason to exist beside similar pages.
- Track detection findings as a review signal, not a standalone verdict about quality, originality, or authorship.
This is particularly relevant to AI search and Generative Engine Optimization. The authors position the approach as applicable to commercial content in that setting. For website owners, the immediate lesson is not that search systems use SlopShape, because the paper does not establish that. Rather, it is that structural sameness is measurable, and content operations can monitor it before a library becomes dominated by interchangeable pages.
For companies that publish with AI assistance, a structural review can also expose process weaknesses. A team may discover that briefs lack source requirements, that writers have no reason to depart from a standard outline, or that editors are optimizing for output volume rather than distinct usefulness. Those are operational problems that can be addressed without abandoning AI tools.
If your site is producing more AI-assisted content, GEO Search Leads can help assess your AI search visibility and identify where content needs stronger differentiation. A focused review can reveal whether your pages provide distinct, source-backed answers or rely on patterns that make them easy to overlook. Use that insight to prioritize improvements to briefs, editorial review, and high-value pages, then start an AI visibility review.
Frequently Asked Questions
What is SlopShape?
SlopShape is an arXiv preprint about identifying AI-generated commercial web content through structural patterns, including information order, evidence placement, and voice.
How accurate was SlopShape's structure-based classifier?
The authors report 98.0 macro-F1 on held-out company domains using structure alone. This is a result from the study's dataset and methodology, not a guarantee for every website or detector.
Did paraphrasing AI-generated posts reduce detection performance?
No. After every AI-generated post was paraphrased with its generating model, the paper reports 98.1 macro-F1, essentially unchanged from the original result.
Can a structural detection score prove that a specific article was written by AI?
No. The research reports classifier performance across a corpus. A score should be treated as a review signal, not conclusive proof of authorship for an individual article.
What should businesses change in their content workflow?
They can review recurring article templates, vary legitimate content structures, strengthen source and expert-review requirements, and evaluate whether each page adds distinct information for readers.
Conclusion
SlopShape provides evidence that AI-generated commercial content can carry detectable patterns beyond its wording. Its reported robustness to paraphrasing makes structural repetition a useful area for editorial attention. For content teams, the practical response is not to rely on a detector alone, but to build publishing processes that reward specific evidence, genuine expertise, and meaningful variation in how useful information is presented.