Summarise this article with:
Sandbox Media original research, July 2026.
We had seven leading AI models write 175 articles for UK aesthetics clinics, judged them blind in a round-robin tournament that produced 4,341 head-to-head verdicts, and put every draft through a range of checks: compliance with UK advertising rules, factual accuracy, reading age, how well it followed the brief, and how machine-written sounding they were.
On the briefs closest to prescription-only treatments, nearly half of all drafts failed at least one check serious enough that a clinic couldn’t publish them as written. That includes drafts from the most expensive models available on the market. One popular budget model failed every single draft it wrote on those briefs.
If you run a clinic and read nothing else:
We tested seven leading AI writing tools on clinic content. On the treatments closest to prescription-only medicines, nearly half of what they wrote had a problem serious enough that a clinic could not safely publish it. The cheap tools failed most, but every price level failed at least once. The failures read fluently, so you won’t always spot them without knowing the rules. If AI writes any of your content, send us one treatment page, and we’ll check it against the full test standard from this study, free.
The question nobody is asking about AI content
Clinics everywhere are using AI to write treatment pages, blog posts, and social captions. In most industries, the worst outcome of mediocre AI content is that it doesn’t capture people’s attention or doesn’t rank. In UK aesthetics, the stakes are a little different.
Botulinum toxin, whatever brand name it is sold under, is a prescription-only medicine. Under CAP Code rule 12.12 and the Human Medicines Regulations 2012, prescription-only medicines can’t be advertised to the public. The ban itself is absolute. The grey area is what counts as advertising, and that is exactly where clinics get caught. A clinic can list a toxin treatment and its price factually, or mention it as an option to discuss at consultation.
The same brand name in a homepage banner, a special offer, or a sentence about how good the results are is promotion, and the ASA treats almost any promotional reference, direct or indirect, as a likely breach. It has been steadily closing the workarounds too: even phrases like “wrinkle-relaxing injections” have been ruled indirect promotion when used to advertise. Enforcement is active. The MHRA issued 47 enforcement notices to aesthetic businesses in 2024, and persistent breaches can be referred to professional regulators or escalated to criminal sanction.
The line between legal and illegal runs through wording, not topics. That is what makes fluent AI output genuinely dangerous here: a model can write a page that reads okay and sits on the wrong side of a line it doesn’t know exists.
So when an AI writes your treatment page, the question is not just “is it any good?” It is “is it legal to publish?”
Nobody had measured how often AI models get that wrong. So we did.
What we tested
We took seven leading AI models across three price tiers, as available in July 2026:
| Tier | Models |
|---|---|
| Budget | Claude Haiku 4.5 · GPT-5.4-mini · Gemini 3.5 Flash |
| Flagship | Claude Opus 4.8 · GPT-5.5 · Gemini 3.1 Pro |
| Frontier reference | Claude Fable 5 |
Each model wrote five drafts against each of five briefs typical of real clinic content work: four patient-facing pieces (dermal fillers, anti-wrinkle injections, skin boosters, and how to choose a safe clinic) and one trade piece aimed at practitioners evaluating a resurfacing laser. The briefs deliberately contained built-in tests, including whether a model knows that prescription-only medicines can’t be promoted to the public. That produced 175 articles.
Quality was judged blind. Drafts were anonymised and compared head to head in more than 4,000 pairwise verdicts by four AI judges from different labs, and we only counted verdicts from judges whose maker did not build the model being scored. That last step actually matters more than you might think. More on that later.
Separately, every draft was scored against a nine-point compliance and accuracy checklist written against the actual UK rules: promoting a prescription-only medicine by brand name, absolute safety claims, citing the wrong country’s regulators, misstating who can prescribe, and other failures that would make a piece unsafe for a UK clinic to publish. Alongside quality and compliance, we also measured how well each model followed the brief, the reading age of what it wrote, its stylistic tells, and how consistent it was from draft to draft.
Finding one: the cheaper the model, the riskier the content
Here is the share of each model’s 25 drafts that failed at least one critical check:
| Model | Tier | Drafts failing at least one critical check |
|---|---|---|
| Gemini 3.5 Flash | Budget | 68% |
| Gemini 3.1 Pro | Flagship | 48% |
| Claude Haiku 4.5 | Budget | 40% |
| GPT-5.4-mini | Budget | 28% |
| Claude Opus 4.8 | Flagship | 8% |
| Claude Fable 5 | Frontier | 4% |
| GPT-5.5 | Flagship | 0%* |
*GPT-5.5 also served as the checklist scorer, so this figure is self-scored; its drafts are being verified by hand.
Grouped by tier, the pattern is strong. Budget models failed on 45% of drafts. Flagship models failed on 19%. The frontier reference model failed on 4%.
The risk is not spread evenly across content, either. It concentrates exactly where the regulations are strongest. On the two briefs with no prescription-only medicine anywhere near them, the laser piece and skin boosters, failures were rare: three drafts out of seventy. On the three briefs in the same topic to prescription-only territory, anti-wrinkle injections, dermal fillers and choosing a clinic, 44% of all drafts failed. Budget models failed 69% of those drafts. Flagship models failed 31%. One budget model, Gemini 3.5 Flash, failed all fifteen of its drafts on those briefs, and one flagship, Gemini 3.1 Pro, failed twelve of fifteen.
In other words: the closer the content sits to a prescription-only medicine, the more likely the AI is to get it wrong, and that is exactly the content a clinic most needs to get right.
What did failure look like? The most common failure was a factual one. One in five drafts across the whole study told readers who can legally prescribe in the UK and got the list wrong, typically naming doctors, nurses and dentists and leaving out pharmacist independent prescribers, who can also legally prescribe these medicines. It looks like a small omission until you notice what it does: a clinic whose prescriber is a pharmacist would be publishing content that quietly tells patients its own service is not legitimate. This error made it into output from premium models as well as budget ones.
Beyond that, ten drafts promoted a prescription-only brand by name. Seven leaned on US-centric authority, citing the FDA as the relevant regulator for UK treatments. Six broke character mid-article, one finishing a patient-facing piece by offering to “adapt this into a clinic-specific version”, the AI equivalent of leaving the price tag on.
Any one of these, published on a clinic’s website, is somewhere between embarrassing and reportable. And none of them announces itself.
Finding two: quality and safety go together
The blind quality ranking, built from those 4,000-plus head-to-head verdicts, split the field into clear groups. Claude Fable 5 and GPT-5.5 finished statistically tied at the top, winning the large majority of their matchups. Claude Opus 4.8 was a clear third. The remaining models clustered in the middle, and Claude Haiku 4.5 finished clearly last.
Put the two findings side by side, and the pattern is pretty clear: the models that wrote the best content were also the cleanest on compliance. The two failures compound at the budget end. A cheap model is more likely to write a generic article, and more likely to make it unpublishable.
Before anyone concludes the top of the table is safe, though: even the frontier model failed one draft in twenty-five. On this evidence, no model earned unsupervised publishing, and in our opinion none ever should.
Averages hide something clinics should care about more: the quality of the worst drafts. Because every draft was compared dozens of times, we can score each individual draft, not just each model. The two tied leaders separate here. GPT-5.5 was the steadier performer, its weakest draft still winning half its blind matchups, while Fable 5’s weakest won 42%. The mid-tier models were the most volatile of all, capable of a draft that beat the leaders one moment and a draft that lost almost everything the next.
And the weakest model produced a draft that lost every single comparison it appeared in. A clinic does not publish a model’s average. It publishes individual drafts, and on this data you can’t know which kind you’ve been handed without evaluating it.
Finding three: content type changes the answer
There was one place the budget tier held its own. On the trade brief, written for practitioners rather than patients, GPT-5.5 won, and the budget-tier GPT-5.4-mini came second, beating flagship models that cost several times more. On the patient-facing, compliance-sensitive briefs, the frontier models pulled far ahead.
The practical rule that falls out of this: the right question is not “which AI is best” but “which AI for which job”. Technical, business-to-business content is a job a budget model can get by with. Patient-facing content about regulated treatments is where the premium models earn their price, in both quality and legal safety.
But be careful what that rule does and doesn’t give you. It picks the model to start from; it doesn’t make any single draft safe. On the briefs nearest prescription-only territory, even the flagship tier failed three drafts in ten. And safe is not the same as good: the best-routed draft in this study still arrived at the wrong reading age for its audience, frequently over length, carrying the stylistic fingerprints of machine writing, and saying nothing another clinic’s AI couldn’t say word for word. The rule has a shelf life as well: it came out of testing, and it will change the next time the labs ship an updated model, which happens several times a year. Knowing which AI for which job is necessary, but it’s nowhere near all you need.
Finding four: every single model failed the reading-age test
While there is obviously no reading-age rule for private clinic websites, NHS England’s accessibility standards and the Patient Information Forum recommend aiming patient information at a reading age of nine to eleven, because that’s roughly the average adult reading age in the UK. A private clinic writing for a broad consumer audience has more flexibility than an NHS leaflet, of course, and some will purposefully write in a more polished way to match their brand. But content about medical treatments still has one job in common with that leaflet: helping someone understand a procedure and its risks well enough to make the right decision for them.
We scored every draft with standard readability measures. All seven models wrote at a reading age of seventeen to eighteen, A-level standard, on patient-facing briefs that named the audience explicitly. Not one adjusted. The gap was really consistent: even the most readable model sat five to six school years above the health-sector benchmark, and well above any format a clinic would choose on purpose. Academic work finds the same pattern.
This is a different kind of failure from the compliance breaches, and in some ways a more important one. It’s not a rule the models don’t know. It’s a judgement about the reader that none of them made unprompted.
Smaller failures pointed at the same thing. Both OpenAI models overshot the brief’s word count on two thirds of their drafts, running 30% long on average, while every other model held the length. And the strongest writers in the test were also the heaviest users of the em dash, the single most recognisable AI fingerprint in published text, at rates no human editor would leave in place.
A finding about the judges themselves
One result from the study deserves a mention even though it’s about the method. When we let AI models judge the writing, they weren’t neutral about their own makers. The ChatGPT judge rated content from its own company substantially higher than independent judges did, a gap of roughly 18 points on our scale. The Claude judge went the opposite way and marked its own company’s content down.
This is why our headline ranking only counts verdicts from judges with no stake in the result, including two judges from labs with no model in the test at all: DeepSeek and Qwen. It is also a caution for anyone tempted to ask ChatGPT to evaluate ChatGPT’s writing: the referee has a team.
What this means if your clinic uses AI content
Start with the question of which model actually drafts your basic copy. Most AI content tools and many agencies default to budget-tier models because they cost a fraction of a penny per article to run. Based on these results, the default fails on nearly half of drafts overall and on more than two-thirds of drafts about treatments closest to prescription-only medicines, which, for most clinics, is the core of the site. It’s a fair question to ask any provider: which model, and what sits between it and your site?
Then check your existing pages for the specific failure modes this study surfaced.
- If any page tells readers who can prescribe, check the list is complete. In the UK, doctors, dentists, nurse independent prescribers and pharmacist independent prescribers can all prescribe these medicines. A page that names doctors and nurses and stops there is factually wrong, and if your own prescriber is a pharmacist, your website is quietly telling patients your service isn’t legitimate.
- Look for prescription-only brand names anywhere promotional: a homepage banner, a special offer, a sentence about how good the results are. A plain price list entry is allowed. Promotion is not, and the ASA reads “promotional” broadly.
- Look for references to the FDA or other overseas regulators. The FDA has no role in UK treatments, and its presence on a page is a reliable sign the content was never written for the UK market at all.
- Look for absolute safety language: “completely safe”, “risk-free”, “no side effects”. No treatment earns those words, and the claim is itself a breach of advertising guidance.
- And look for any sentence where the tool talks instead of the clinic: an offer to “adapt this article”, a note addressed to whoever pasted the text rather than to the patient reading it.
If AI sits anywhere in your content workflow, a compliance check, alongside human input and editing, has to sit between the model and the publish button. At every tier we tested, drafts failed, and every failing draft often fails silently. Generation and verification are different jobs.
And verification alone isn’t the end. Even a draft that passes every rule still arrived at the wrong reading age for patients, over length, and carrying the stylistic fingerprints that mark text as machine-written. Someone has to do that editorial work on every draft, or the content might be legal but still poor and of low quality to both users and search engines.
Add up what publishing AI content safely actually requires: current knowledge of which model leads this month, because the answer changes with every release; the regulatory knowledge to catch failures that read fluently; the editorial layer that fixes register, length and tells; the SEO and AEO expertise to make the content visible in the first place, both in Google rankings and in the AI answers patients increasingly read instead, because a safe page nobody finds achieves nothing; and the evaluation infrastructure to check any of this at all. That is not a tool you subscribe to. It is a capability, and it is a strange one to expect a clinic to build in-house, for the same reason clinics do not run their own advertising law departments. It is a full-time discipline that sits outside the job of treating patients.
How we use AI at Sandbox Media
For transparency, here is how we actually use AI at Sandbox, because the honest answer is not “we don’t”:
Nothing gets written just because AI makes writing cheap. What a clinic’s site needs is a strategy decision first: what patients actually search for, what questions the AI assistants are being asked, where the clinic can realistically win, and which pages deserve to exist at all. Only after that does AI draft our first version of a commodity layer: the explainers, the FAQ answers, the descriptive pages every clinic needs and no clinic is differentiated by. Which model drafts what is set by a routing table partly built from this study. And the table is re-tested when the labs ship updates, because this study is only a brief snapshot.
Every draft then passes through the same gates and checks we used here plus many more. Some of these include the compliance checklist you’ve just read and a register pass that brings the reading age down to where patient information belongs.
Generation and clean-up are the cheapest parts of what we do. The value sits either side of them: the strategy that decides what gets written, and the work that makes it get found and chosen.
And a lot of content AI never touches. The pieces that make a clinic visible in AI recommendations are precisely the ones no model can write: your prescriber explaining how they assess a first-time patient, your protocol for when something goes wrong, the questions your patients actually ask in the consultation room. We ask, we research, we write, we attribute. That layer is what differentiates a clinic, and automating it would defeat its purpose.
The part that doesn’t come off a shelf is the judgement across all of it: which model, for which job, checked against which rules, and when the choice is that AI should not write it at all. If you’d like the same standard applied to your own site, send us one treatment page, and we’ll run it through the full checklist, free.
Why not just use AI yourself?
You could, and some clinics quietly do. For internal drafts, notes, and first drafts, you probably can. Before you make it your publishing strategy, though, it’s worth being clear about what this study says about the do-it-yourself route.
Start with the tool you would actually reach for. Type into the free tier of a chatbot, or into most AI content tools, and you are usually getting a budget-tier model, because those are the cheap ones to run. On our data, that is the tier that failed two-thirds of its drafts on the treatments closest to prescription-only medicines. Picking a better model helps, but it means knowing which one is better this month, and that answer is a snapshot, not a permanent fact. The proof arrived before we could even publish: within the same month of our work, OpenAI released GPT-5.6 and Anthropic released Claude Opus 5, superseding two of the seven models we tested, including one of the two joint winners. Neither new model has been through this test. Rankings move every time the labs ship an update, and the only way to know the current answer is to keep re-running the test. That is a research habit, not a subscription.
The failures are silent. A draft that breaches the advertising rules reads just as fluently as one that doesn’t. You only see the problem if you already know what rule 12.12 prohibits, which is precisely the knowledge most clinics are hoping to outsource.
And even a perfectly chosen model needs supervision. The winner of this study overshot the requested length on two-thirds of its drafts, and like every other model, it wrote patient content five to six years above the recommended reading age. Best in test is not the same as finished.
Then there is the ceiling on the whole approach: the content that actually wins visibility is the content no model can write for you. Our audits of AI recommendations across UK aesthetics keep showing the same pattern: generic content, however clean, is interchangeable, and interchangeable content does not get recommended, normally the opposite. What moves the needle is the material only your clinic can produce: your practitioners, your protocols, your first-hand expertise. AI can only ever write what everyone else’s AI can also write.
And all of this is still only the content. A published page has to be found before it can persuade: technically sound and fast enough to rank, marked up so search engines and AI assistants can parse it, visible in local search where patients actually look, connected to a booking path that converts the visit, and measured well enough to know whether any of it is working. A chatbot writes words. It doesn’t run a marketing operation, and the words are the only part of the operation it can even attempt.
Method and limits, stated plainly
All 175 drafts were generated in July 2026 against identical briefs with identical settings, five samples per model per brief to smooth out lucky and unlucky draws. Quality judging was pairwise and blind: drafts were anonymised, compared two at a time in randomised order, and aggregated with a Bradley-Terry model with bootstrapped confidence intervals.
Four judges completed the full run: one each from Anthropic and OpenAI, plus two from labs with no model in the test, DeepSeek and Qwen. A fifth judge completed only part of the run due to provider rate limits and is excluded from headline figures. Every model’s score counts only verdicts from judges whose maker did not build it. Where confidence intervals overlap, we report a tie rather than inventing an order, which is why the top two are reported as tied.
The compliance checklist was scored by a single strong model. The checks are objective, and the flagged breaches are being verified by hand. Readability was scored with standard Flesch-Kincaid measures, and consistency comes from per-draft win rates across the same blind verdicts.
The study covers five briefs in one specialism, and results are a snapshot of the model versions available in July 2026. Models change quickly: GPT-5.6 and Claude Opus 5 both shipped within the month of our research, and neither appears in these results. We will re-run the study as the field moves.
What’s next
We are now pointing the same compliance checklist at a larger question: not what AI would write for clinics, but what UK clinics have already published. That research is underway, and the early picture suggests the compliance problem did not start with AI.
Get one treatment page checked, free
Send us a page from your site and we’ll run it through the full checklist from this study: compliance, accuracy, reading age and the tells of machine writing. You’ll get exactly what we found and what to fix. No obligation.
We’ll also tell you which model most likely wrote it.
Questions clinics ask about AI content
Is it legal for a UK clinic to advertise Botox on its website?
Is it legal to use AI to write clinic content?
Which AI writes the best clinic content?
How often does AI content break UK advertising rules?
How can I tell if my content was written by AI?
What reading age should clinic content be written at?
Sandbox Media is a UK digital agency specialising in search, AI visibility and growth for aesthetics clinics. This research was designed and run in-house; no AI company sponsored or reviewed it.













