A general-purpose AI will write you a citation for almost any claim you make. It will look perfect: a real-sounding journal, plausible authors, a tidy year and volume number. There is just one problem. A meaningful share of the time, the paper does not exist.
This is not a hypothetical risk anymore. Lawyers have been sanctioned for filing court briefs built on cases that ChatGPT invented, and a widely covered US government health report was found to cite studies that were never written. In most fields, a fabricated citation is embarrassing. In pharmaceutical promotion, it is exactly the kind of error MLR review exists to catch, and exactly the kind that should never have reached review in the first place.
Here is why AI-generated citations fail, what the failures actually look like, and a workflow to catch them before they cost you.
AI does not just get citations wrong. It invents them.
It helps to understand why this happens. A large language model does not look anything up. It predicts plausible text. Asked for a reference, it generates the most likely-looking citation, which is often a convincing blend of a real journal, real authors, and a paper that was never published. The fabrication is not a glitch; it is the default behavior of a system built to sound right rather than to be right. We dug into that in why general AI can't handle pharma compliance.
This is measurable, and getting worse. Earlier studies estimated that between 30% and 69% of the references LLMs generate in biomedical contexts are fabricated. And the problem is no longer confined to chatbots: a 2026 audit in The Lancet scanned 2.5 million biomedical papers and found fabricated references in the peer-reviewed literature rising more than twelvefold since 2023, reaching roughly one paper in 277 by early 2026, with the inflection point lining up with the spread of AI writing tools. Tellingly, the fake references were not crude. They were topically specific, correctly formatted, and attributed to real researchers, which is exactly why they survive a casual read.
In a promotional draft, that same plausibility produces a few specific failures:
- The real paper, the wrong claim. A genuine study is cited, but it does not actually support the sentence in front of it, or it measured something adjacent.
- The wrong source type. A patient-education webpage or a press release is offered as the evidence for a clinical claim, in place of the primary study.
- The wrong label. A claim is checked against the wrong region's prescribing information, and reads as fully substantiated against a label that does not apply in your market.
Each of these passes a casual glance and fails MLR.
Why it is worse in pharma
In most settings, a bad citation wastes a reader's time. In pharmaceutical promotion, an unsupported or mis-attributed claim is a regulatory and patient-safety issue, which is the whole reason the review apparatus exists. A single hallucinated reference that survives into a live promotional piece is close to a worst case.
This is also why general-purpose assistants are not fit for the job. A tool like Copilot or ChatGPT is not anchored to your approved label or your reference library; it has no idea what your sources actually say. The medical reviewers we have spoken with are blunt about it: these tools are not regulatory-ready, and the moment AI is used to "save time" on referencing, the saved time reappears, with interest, in review.
What the people doing the work actually think
The practitioners closest to MLR are not anti-AI. They are precise about where it helps. Asked where AI offers the most genuine near-term value, they pointed to assistive, narrow tasks: speeding up initial data-gathering and fact-checking (27%), and identifying inconsistencies or missing information in materials (27%). Useful, bounded jobs.
Asked the opposite question, which AI claim is most overstated, the top answer was telling: "speeding up compliance and regulatory processes" (27%). The people who do the work are optimistic about AI as an assistant and skeptical of it as a shortcut through compliance. And when an AI-assisted piece causes a breach, they hold the human reviewer accountable, not the tool.
The throughline is simple: use AI to find problems, not to have the final word.
A verification workflow for AI-generated citations
If you use AI anywhere in your drafting, treat its citations as claims to be proven, not facts to be trusted. This is a workflow you can run on any AI-assisted draft:
- Assume every AI-supplied citation is unverified. Guilty until proven. This single mindset change prevents most failures.
- Confirm the source exists. Look it up by DOI or PMID in PubMed or on the publisher's site, not by title. A fabricated reference dies here.
- Confirm it is the right kind of source. A peer-reviewed primary study, not a press release, a patient page, a secondary citation, or a review standing in for original data.
- Confirm it says what the claim says. Open it, find the exact supporting line, and check that the specifics match: patient number, comparator, endpoint, and population. A real paper attached to the wrong claim is still wrong.
- Confirm it matches the right label. Check the claim against the prescribing information that applies in your market, not a foreign one.
- Anchor it. Highlight the supporting passage in the source so your reviewer verifies in seconds instead of re-hunting.
Run honestly, this catches fabrications, wrong-source swaps, claim drift, and region mix-ups before a reviewer ever sees them. The catch is obvious: doing all six steps by hand, for every claim, on every AI-assisted draft, is exactly the kind of work nobody sustains. The workflow only survives if it is automated. That is not just our view: facing the same problem in the published literature, the Lancet audit's authors concluded that automated reference verification should run before review, and that the barrier to adopting it is institutional rather than technological.
Automating the check
This is the gap PharmaText.ai was built to close. Instead of trusting AI to produce citations, it does the opposite: you write the copy and supply the sources, and the tool checks every claim against your uploaded source PDFs, flags anything that is unsupported, mis-sourced, or contradicts the label, and anchors what passes to the exact supporting line. The AI does the evidence work, not the inventing, and a human still makes the call. And when you need to turn a verified primary source into a clean reference, our free AMA citation generator handles the formatting.
AI is a real asset in medical writing. But the version that helps is grounded in your sources and checked by a person, not the one that confidently makes things up. Until your tool can prove a citation, verify everything.
Related: see why general AI can't handle pharma compliance, what actually needs a reference, and linking and anchoring.
Sources: fabricated-reference data from Topaz M et al, "Fabricated citations: an audit across 2·5 million biomedical papers," The Lancet 2026;407(10541):1779-81. AI value and "most overstated" figures from a 2024 industry webinar by Impatient Health.
Build Compliant Content Faster
PharmaText.ai helps teams reduce MLR cycles by 40% using precision traceability.
Book a demo →