Nothing in This Literature Ever Fails
Across 252 controlled trials from four countries, not one found the treatment ineffective. We audited our own 3,385 sources and found the same machine running inside our database.
· 14 min read
He who knows only his own side of the case, knows little of that.John Stuart Mill, On Liberty
In 1775 the French Académie des Sciences stopped reading proposals for perpetual motion machines. Not because it had tested each one and found the flaw – the submissions kept arriving, each ingenious, each with its own arrangement of weights and wheels and capillary tubes, and refuting them one at a time had become a full-time occupation for people with better work to do. The Académie stopped because it had understood something more useful than any individual refutation. A class of machine that never loses energy is not a machine that has solved friction. It is a machine whose losses are going somewhere you have not looked.
This is the most useful test we know for a body of evidence, and it costs nothing to apply. Do not ask whether the results are good. Ask where the failures went.
A machine with no losses
In 1998 Andrew Vickers and colleagues reviewed 252 controlled trials of acupuncture and sorted them by country of origin. Trials from China, Japan, Hong Kong and Taiwan were positive. Not mostly positive: all of them. Of eleven trials from Russia and the USSR, ten were positive. Their broader sample of 405 trials across other interventions found the same signature – 99% positive from China, 97% from Russia, 95% from Taiwan, against 75% from England.
Their conclusion is one sentence and it is worth reading slowly: no trial published in China or Russia/USSR found a test treatment to be ineffective.
A field in which experiments can fail will produce failures at some rate. That rate is the friction – the energy the system loses to reality, and the only reason its outputs mean anything. A literature reporting no failures has not escaped that cost. It has relocated it.
We know roughly where. A 2020 systematic review tracked 1,758 acupuncture trials registered with the WHO International Clinical Trial Registry Platform between 1990 and 2018 and looked for the published results. It found 178. Ten percent of registered trials reported their outcomes in a full paper. The other ninety percent are the losses, sitting in the drawer where the machine keeps them.
What we found in our own database
It would be comfortable to write that paragraph about somebody else. AcuiQ reads treatment protocols out of the published literature at scale and matches them to the symptoms people search for. Last week we audited every source behind every protocol we hold – 3,385 distinct references – and resolved each one to its PubMed record to ask a single question: were the subjects human?
For 466 of them, they were not.
| Sources audited | 3,385 |
|---|---|
| Sourced from animal research | 466 sources · 541 protocols |
| Confirmed human clinical evidence | 2,067 sources |
| Classical clinical reference texts | 5 sources |
| Unverifiable – no retrievable metadata | 847 sources · 1,034 protocols |
The 541 were rats, mice, rabbits, dogs and cats. One was a systematic review of acupuncture for laminitis in horses. They had been sitting in a database that answers questions from people with symptoms, and they had been sitting there for as long as the database has existed.
The case that surfaced it is exact enough to be worth naming. A user searching bradycardia – a slow heart rate – was served two protocols, and two was all we had. The first came from a 2013 study of moxibustion at different temperatures on cardiac function in rats. The second came from a 2017 paper on mast cell activation in electroacupuncture against pituitrin-induced bradycardia in rabbits. Both were indexed by the National Library of Medicine with the descriptors that say so. Nobody had looked.
The papers are not the fraud
Here is the part that resists the easy telling. Neither of those studies is bad science. Both are careful mechanistic work, faithfully reported, correctly indexed by their publishers, doing exactly what animal research is for. The rabbit paper is a competent investigation of mast cell signalling. It makes no claim about treating a human being and never pretended to.
The failure was ours, and it happened at the moment of aggregation. A point prescription validated in an anaesthetised rabbit carries no dose, no depth, no retention time and no point selection that means anything for a person – and every one of those parameters silently acquired a human interpretation the instant the protocol entered a table next to human trials. Nothing was falsified. A row simply crossed a boundary without the thing that gave it meaning, and the arithmetic proceeded without complaint.
This is the mechanism worth understanding, because it does not require anyone to lie. It only requires that nobody is asked to check. Charlatanry at scale is rarely a person inventing a result. It is a pipeline in which no stage owns the question of whether the evidence supports the claim, and every stage assumes an earlier one did.
A field that has not agreed where the points are
Consistency is the second place the machine leaks, and here the field documents its own problem better than any critic could.
When the World Health Organization set out to standardise acupuncture point locations for the Western Pacific Region, it found international disagreement on the position of roughly a quarter of the 361 classical points. Ninety-two were contested. Resolving them took seven informal consultations and four task force meetings, and produced agreement on eighty-six. Six points – LI19, LI20, PC8, PC9, GB30 and GV26 – were published with alternate locations, because the delegations could not agree on where they are. That standard was issued in 2008. Those six points still have two official positions.
Our extraction work meets the downstream version of this daily. Reading 266 papers in a single pass last week, we logged source papers mis-coding their own points against the names they themselves printed:
| Point named in the paper | Code the paper printed | Code that name actually carries |
|---|---|---|
| Hegu 合谷 | LI14 | LI4 |
| Neiting 內庭 | ST4 | ST44 |
| Shenque 神闕 | GV8 | CV8 |
| Yintang 印堂 | GV24+, EX2 | EX-HN3 |
| Taichong 太衝 | LI3 | LV3 |
Hegu is the single most commonly needled point in the clinical literature. A paper wrote it as LI14, which is a real and entirely different point on the upper arm. These are not exotic edge cases; they are the field’s most familiar anatomy, mis-numbered in peer-reviewed journals, and the only reason we catch them is that we resolve every point by its printed name rather than trusting the code beside it.
The catalogue argues with itself
The third leak is not an absence. It is two fields of the same row disagreeing, in public, on the same screen – and no coverage metric will ever find it, because coverage counts what is missing and both fields are present.
XL07 is Lanwei, the Appendix point, an extra point below the knee that exists for one purpose. Its page carried a Cautions block reading avoid in acute appendicitis, set directly above an Indications block reading appendicitis, acute. The auricular Appendix point said the same of itself. So did AT16, the Epilepsy point, cautioning against seizure disorders; EL16, the Groove of Hypotension, cautioning against low blood pressure; HX06, the Anus point, advising against use in haemorrhoids.
| Point | Cautions block | Indications block |
|---|---|---|
| XL07 Appendix Point | Avoid in acute appendicitis | appendicitis, acute |
| AT16 Epilepsy | Caution with seizure disorders | epilepsy |
| EL16 Groove of Hypotension | Caution in severe hypotension | hypotension |
| AH12 Thyroid | Caution with thyroid disorders | hyperthyroidism, goitre |
| SP01 Hidden White | Avoid in patients with bleeding disorders | abnormal uterine bleeding |
| LV03 Great Surge | Caution in hypertension | hypertension |
Thirty-one point rows contradicted themselves this way. A further 147 protocols recommended a point whose own caution named a condition that protocol was treating: LI11 and LV03 in 26 hypertension protocols each, BL15 and CV17 in 45 cardiac protocols, ST41 in 15 protocols for the ankle sprain it warns you off. Every one of those cautions tracks the row’s name rather than its clinical risk, which is the signature of text generated from a label. Across 773 point rows, 759 carried a caution, and between them they drew on 164 distinct strings.
Pregnancy is the largest family and the only one with a resolution already available. LI4 Hegu and SP6 Sanyinjiao appear on every forbidden-in-pregnancy list ever printed, and both are indicated, in the same catalogue, for inducing labour. That reads as incoherent until someone says the missing sentence out loud: they are forbidden because they are held to move a pregnancy along, so at term the prohibition becomes the indication. The catalogue printed the prohibition. It printed the indication. It never printed the sentence, and neither does much of the literature the catalogue was built from.
The prohibition itself is thinner than its confidence suggests. Carr’s 2015 review in Acupuncture in Medicine gathered 15 trials in which 823 women received somewhere between 4,549 and 7,234 treatments at forbidden points, and found rates of preterm birth and stillbirth equivalent to untreated controls. The lists themselves differ from text to text. A prohibition transmitted for centuries as a list, with no mechanism attached and no trial behind it, will contradict the indications sitting beside it, and the contradiction will propagate into every database built downstream. Ours is one of those databases.
Two rows had it worse, and this failure is entirely ours. Source catalogues print indications and contraindications in a single field. Our parser split one on the wrong delimiter, and the auricular Ovary point entered the database as treating a condition named etc.contraindications: pregnancy. The Endocrine point came in treating genitourinary disorders such as prostatitis.contraindications: pregnancy. The contraindication was not dropped. It changed sign, and was served to readers as something the point was good for – while the Endocrine row, robbed of the only contraindication it had, sat there declaring No specific contraindications.
A name is not a species
In 1993 a Brussels clinic prescribing a herbal weight-loss preparation produced a cluster of young women with rapidly progressive kidney fibrosis. The formula was supposed to contain Stephania tetrandra. It contained Aristolochia fangchi, which carries aristolochic acid. Belgium recorded 128 cases of the resulting nephropathy; by 2000 the New England Journal of Medicine was reporting urothelial carcinoma in the same cohort, and China struck the herb from its Pharmacopoeia in 2004.
The substitution happened because both plants are traded as 防己, fang ji. Stephania tetrandra is Han Fang Ji. Aristolochia fangchi is Guang Fang Ji. The Chinese drug name names a role in a formula, and a role is not an organism – two plants from different families answer to it, one of them a carcinogen.
Our herb catalogue holds 5,397 rows and keys every one of them on a Latin binomial. Asked for the drug names, it returns this:
| Drug name in the catalogue | Binomial it is filed under | Safety text on that row |
|---|---|---|
| Han fangji | Stephania hancei | Limited safety data; likely contains alkaloids similar to other Stephania species |
| Guang fangji | Stephania guangxiensis | Limited safety data; likely contains alkaloids similar to other Stephania species |
| Guan mutong | Aristolochia kaempferi | Contains aristolochic acid; nephrotoxic and carcinogenic |
There is no Aristolochia fangchi row in our catalogue at all. The name that put 128 people into renal failure is filed under a Stephania, and carries the reassurance that it is probably much like the other Stephanias. The 1993 substitution is not merely recorded in our data. It is reproduced by it, with the toxicology replaced by an inference drawn from the binomial.
Guan Mu Tong is the same error with a luckier outcome. The drug is Aristolochia manshuriensis; our row files it under A. kaempferi. The species is wrong and the aristolochic acid warning is right, because the hazard tracks the genus and the genus happened to survive the mistake.
The binomial fails in the other direction too. Melia azedarach occupies six rows – melia fruit, melia bark, the tree – under one shared key. One says highly toxic; can cause vomiting, seizures, and death; avoid all internal use. Another says use only under professional supervision. Both are correct about their part of the plant, and 751 herb codes in the catalogue are duplicated like this. On 67 of them the duplicate rows disagree about safety. Our protocol sync resolves a code by building a lookup keyed on it, so where rows collide the last one loaded wins – across those 67 codes, 147 rows are shadowed by a sibling with a different safety profile.
None of this has ever reached a reader. No herb appears in any protocol, and every app query filters the herb rows out. We report it because the defect is sitting in the catalogue waiting for the day herbs are served, and because the shape of it is the point: a caution generated from a name will always be a description of the name. Of the 5,397 herb rows, 1,395 distinct warning strings cover the lot; 2,667 name pregnancy, 1,783 say the safety data is limited, and 740 assert that no specific contraindications are known.
The third we cannot check at all
Of our 3,385 sources, 847 resolve to nothing. Bare DOIs from journals no index carries, free-text citations with no retrievable record, papers that exist only as a line someone typed. They account for 1,034 protocols. We cannot confirm those were human trials. We also cannot confirm they were not.
We kept them, and we want to be exact about why, because it is the least defensible decision in this piece. They are mostly Chinese-language clinical studies, and deleting a third of the database on suspicion would be its own kind of dishonesty – the kind that mistakes what is easy to verify for what is true, and rewrites the evidence base to match the reach of an English-language index. So they stay, and they are labelled, and this paragraph exists so that nobody has to discover the caveat by reading our source code.
What we changed
All 541 animal-derived protocols are gone – removed from the database and from the file that feeds it, because deleting from one and not the other would have restored them on the next sync. Every extraction path now runs a species gate before an article reaches the pipeline, and the rule is stated in the extraction spec itself: human subjects only, and an honest empty result is correct when a paper does not support one.
Bradycardia now returns nothing. That is the right answer. We do not have human evidence for acupuncture in bradycardia, and a page that says so is worth more than a page that offers a rabbit.
Twenty-four further protocols turned out to be rows that could never render a recommendation at all – no resolvable points, and in twenty cases a corrupted citation sitting where the point codes belong. They inflated our coverage statistics and offered nothing. They have been removed too. One of them, we noticed on the way out, was a study in arthritic rats that our species audit had missed, because the audit only examined rows that had a reference and that row’s reference field was empty. Every filter has an edge, and the edge is where the next error lives.
The self-contradicting cautions are rewritten – 50 point rows, each correction written by hand. A generated replacement would have been the original defect with a fresh timestamp, since the contradiction is what a machine can see and which of the two fields is wrong is exactly what it cannot. The Appendix point now says that acute appendicitis is a surgical emergency and must never wait on a needle. SP1 says that anticoagulants raise the risk of a haematoma at the needle site, which is a fact about technique, leaving uterine bleeding where it belongs in the indications. The forbidden points carry the prohibition, the reason for it, and Carr’s finding, together, in one sentence. Where a caution was doing two jobs the second one survived: GB21 warns about pregnancy and about pneumothorax, and the tool now refuses to apply any rewrite that would drop an anatomical hazard.
The two inverted rows are unlinked, the Endocrine point’s real contraindication is restored, and the indication half that came in fused to it – genitourinary disorders, prostatitis – is relinked rather than thrown away with the malformed string. The detector now runs inside the standing database audit, so the next one surfaces on the next pass instead of the next year.
Five point rows still contradict themselves: SP6, KD21, CV2, CV3 and CV4, every one of them a gestational indication meeting a precaution about needling over a pregnant uterus. SP6 is indicated for restless foetus syndrome, a mid-pregnancy use, on a point the same row calls forbidden in pregnancy. No text edit resolves that. It needs a clinician deciding which field is wrong, and until one does, it stays in the report where we can see it.
The test is cheap and nobody applies it
The hostility that meets questions in this field is usually described as cultural, and that framing is too flattering. What we can measure is structural, and structure is more damning than temperament: ninety percent of registered trials never report a result, entire national literatures have never published a negative one, the anatomy is disputed at a quarter of its points by the standards bodies that define it, the safety lists forbid the treatments they prescribe, and the herbs are indexed by names that two different plants answer to. No individual has to be defensive for that system to behave exactly as if it were.
None of this establishes that acupuncture does not work. That is a different claim, it requires different evidence, and this audit does not touch it. What it establishes is narrower and more practical: the evidence base is not currently in a condition to answer the question either way, and a great deal of what is published as clinical evidence will not survive being asked what species it came from.
We are going to keep asking. This article is a ledger rather than an essay, and it will be extended as the audit continues – the next passes cover credibility grading, duplicate protocols, the herb catalogue’s binomial keys, and whether trials that registered an outcome published the one they registered. We will report what we find in our own data first, because a critique that exempts the critic is not worth reading.
Corrections are welcome and will be credited. Every figure here comes from our own corpus or a cited source, and if one is wrong we would rather know. Sources for this pass: Vickers et al., Controlled Clinical Trials 1998; Carr, Acupuncture in Medicine 2015;33(5); Vanherweghem et al., The Lancet 1993; Nortier et al., New England Journal of Medicine 2000;342:1686. Last updated 29 July 2026.