---
title: "Reading an acupuncture trial"
url: "https://acuiq.com/wiki/research/reading-an-acupuncture-trial"
updated: 2026-10-09
last_verified: 2026-10-09
license: CC-BY-4.0
license_url: "https://creativecommons.org/licenses/by/4.0/"
section: "research"
basis: "Modern research"
audience: "Anyone"
published: "2026-10-09"
reading_minutes: 10
sources: 11
---

# Reading an acupuncture trial

_What the STRICTA reporting standard asks a trial report to contain, how the control, blinding, outcome and sample size shape a result, and how RoB 2 and GRADE judge it._

A report of an acupuncture trial answers a narrower question than its title usually suggests: this treatment, given this way by these practitioners, compared with this control, in these patients, measured on this scale. Three published tools help a reader work out which question it was. STRICTA, the reporting standard for acupuncture trials, says what the report must describe. Cochrane’s risk-of-bias tool, RoB 2, judges how far the result could have been distorted by the way the trial was run. GRADE rates how certain a whole body of trials is. This page takes each in turn, along with the design choices they look at: the control, blinding, the outcome and the number of patients.

## What STRICTA asks a report to say

STRICTA, the Standards for Reporting Interventions in Clinical Trials of Acupuncture, was [revised in 2010](https://doi.org/10.1371/journal.pmed.1000261) by Hugh MacPherson and colleagues with the CONSORT Group, which maintains the general standard for reporting randomised trials, and the Chinese Cochrane Centre. A draft went to 47 experts in 15 countries and was revised at a workshop of 21 people in Freiburg in 2008. The revised STRICTA is an official extension of CONSORT: in an acupuncture trial it replaces the CONSORT item that describes the intervention. It has six items and 17 sub-items.

| Item | What the report should state |
| --- | --- |
| 1. Acupuncture rationale | The style of acupuncture (for example Traditional Chinese Medicine, Japanese, Korean, Western medical, Five Element, ear); the reasoning behind the treatment, with sources; how far treatment was individualised |
| 2. Details of needling | Number of needles per session; the points, and whether one side or both; depth of insertion; the response sought, such as de qi or a muscle twitch; manual or electrical stimulation; how long needles stayed in; the needle type |
| 3. Treatment regimen | Number of sessions; how often and how long they were |
| 4. Other components | Anything else the acupuncture group received, such as moxibustion, cupping, herbs, exercise or advice; the setting, and what patients were told |
| 5. Practitioner background | Each acupuncturist’s qualifications, years in practice and other relevant experience |
| 6. Control or comparator | Why that control was chosen, with sources; a precise description of it, and for a sham, the same detail as items 1 to 3 |

A reporting standard does not make a trial sound; it makes it possible to tell whether it was. How far reports follow STRICTA varies. A [2016 analysis](https://doi.org/10.1371/journal.pone.0147244) by Ma and colleagues of 1,978 randomised trials of acupuncture and moxibustion in Chinese journals, published from 1978 to 2012, found reporting had improved after CONSORT and STRICTA were introduced in China. Even in the latest period, though, 1.2% of reports described how their sample size was calculated, 26.3% how the random sequence was generated, 4.9% how allocation was concealed and 9.1% any blinding. None described the setting of treatment or the acupuncturists. [Where the trials come from](/wiki/research/where-the-trials-come-from) covers the geography of the literature.

## The control decides the question

STRICTA lists the controls a trial may use: sham acupuncture, usual care, another active treatment, a waiting list or no treatment. Each answers a different question.

| Control | The question the comparison answers |
| --- | --- |
| No treatment or a waiting list | Does the whole package of acupuncture, ritual included, do better than leaving the condition alone? |
| Usual care, or acupuncture added to usual care | Does adding acupuncture to what patients would otherwise receive change the outcome? |
| Another active treatment, such as a drug | Is acupuncture better than, or as good as, that treatment? |
| Sham acupuncture | Does the needling itself, at the chosen points, add anything to a procedure that looks the same to the patient? |

The choice changes the size of the result. In a [2014 reanalysis](https://doi.org/10.1371/journal.pone.0093739) of 29 chronic pain trials by the Acupuncture Trialists’ Collaboration, acupuncture beat every kind of control, but by less against a sham whose needles went through the skin than against one whose needles did not, and by more against unspecified routine care than against care that followed a protocol, though that last difference was not statistically significant. The authors pointed out that this matters when planning how many patients a trial needs. [Sham acupuncture](/wiki/research/sham-acupuncture) sets out the kinds of sham and why none is inert.

Other features of a trial move the result too. The [FAMOUS study](https://doi.org/10.1136/bmjopen-2021-060237), a 2022 analysis by Gang and colleagues of 584 acupuncture trials with more than 100 participants each, published from 2015 to 2019, found larger effects in trials run at a single centre than at several, in trials of needling than of non-penetrating methods, in trials with more frequent sessions, and in trials that did not report their funding.

## Who knew which treatment was given

Blinding means keeping someone from knowing which arm a patient is in. In acupuncture there are three groups to consider, and they differ in how far blinding is possible.

- **Patients.** Sham needles and devices are designed to keep patients from knowing, and the validation studies on [Sham acupuncture](/wiki/research/sham-acupuncture) report how often patients guessed. A trial against no treatment, usual care or a drug cannot blind patients at all.
- **Practitioners.** The person inserting a needle knows whether it is real. A [2009 review](https://doi.org/10.1136/bmj.a3115) of 13 trials with acupuncture, sham and no-acupuncture arms by Madsen and colleagues found practitioners blinded in none of them. A [double-blind needle](https://doi.org/10.1186/1472-6882-7-31) designed by Takakura and Yajima in 2007 is one attempt to change that; ten acupuncturists using it identified real and non-penetrating needles about as often wrongly as rightly.
- **Outcome assessors.** Where an outcome is a test result or an examiner’s rating, the assessor can be kept unaware. Where the outcome is pain, sleep or mood, the patient is the assessor.

The [Cochrane Handbook](https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-08) explains how this is judged. A self-reported outcome such as pain is rated at least “some concerns” for bias whenever knowing the assigned treatment could have influenced the report, and “high risk” if that influence is likely, even when a blinded interviewer collects the answers.

## The outcome, and how much change matters

A trial should name one primary outcome in advance, with the time at which it will be measured, so that it cannot choose afterwards the measure that came out best. Registration in a public trial registry is how that is checked; [Where the trials come from](/wiki/research/where-the-trials-come-from) reports how often acupuncture trials changed their primary outcome between registration and publication.

The size of a change matters as much as whether it is statistically significant. The Initiative on Methods, Measurement, and Pain Assessment in Clinical Trials (IMMPACT), a consensus group that sets methods for chronic pain trials, [wrote in 2009](https://doi.org/10.1016/j.pain.2009.08.019) that a fall in a patient’s pain of 30% or more has been treated as moderately important and 50% or more as substantial. The same paper warns that a clinically important change in one patient and a clinically important difference between two groups’ averages are often confused. The distinction matters because a trial in which many patients in both arms improve can still show a small difference between the arms. The Madsen review is an example of the second kind of number. It found acupuncture beat sham by about 4 mm on a 100 mm pain scale and judged that too small to matter clinically. The IMMPACT paper recommends reporting responder rates, the share of patients whose pain fell by a set amount, alongside group averages, as the GERAC trials described on [Sham acupuncture](/wiki/research/sham-acupuncture) did.

## Sample size

A trial needs enough patients to detect the difference it is looking for, and the report should say how that number was calculated, which is one of the CONSORT items the Ma analysis checked. The calculation depends on the control: a trial against sham expects a smaller difference than a trial against no treatment and so needs more patients, which was the point the Acupuncture Trialists’ Collaboration made in its 2014 reanalysis. In the Ma analysis of Chinese-journal trials, almost none reported a sample size calculation. The FAMOUS study excluded trials of 100 participants or fewer, so it cannot say whether small trials report larger effects; it found that country, sample size and blinding of participants were associated with effect size only when each was looked at alone, not once other factors were taken into account.

## Risk of bias: RoB 2

[RoB 2](https://doi.org/10.1136/bmj.l4898), published by Jonathan Sterne and colleagues in 2019, is the tool Cochrane reviews use to judge each trial result. According to the [Cochrane Handbook](https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-08), it asks about five sources of bias: the randomisation process; deviations from the intended treatment, where blinding of patients and practitioners is assessed; missing outcome data; measurement of the outcome, where blinding of assessors is assessed; and selection of the reported result. Each is judged low risk, some concerns or high risk, and the overall judgement takes the worst of them.

The judgements predict results. A [2023 analysis](https://doi.org/10.1186/s12874-023-01904-w) by Li and colleagues of 84 Cochrane acupuncture reviews found that, in both Chinese-language and other trials, those at high or unclear risk of bias reported larger effects than those at low risk: on allocation concealment among the Chinese-language trials, and on blinding of patients and practitioners among the rest.

## Certainty of the evidence: GRADE

A review pools trials, and GRADE, set out by Guyatt and the GRADE Working Group in [2008](https://doi.org/10.1136/bmj.39489.470347.AD), rates how much confidence the pooled estimate deserves. Evidence from randomised trials starts as high certainty and is rated down for five reasons: limitations of the trials, which is where RoB 2 judgements feed in; inconsistent results between trials; indirectness, such as patients or comparisons different from the question asked; imprecision; and publication bias. The four levels mean:

- **High:** further research is very unlikely to change confidence in the estimate.
- **Moderate:** further research is likely to matter and may change the estimate.
- **Low:** further research is very likely to matter and is likely to change the estimate.
- **Very low:** any estimate of effect is very uncertain.

A review that reports “low-certainty evidence that acupuncture reduced pain” is stating both a finding and a warning that the next trials may move it. The evidence pages in this wiki, such as [Low back pain](/wiki/evidence/low-back-pain), give each review’s GRADE rating where it gave one, and [Evidence maps](/wiki/research/evidence-maps) reports how the large overviews sorted conditions by certainty.

## What these tools cannot tell you

STRICTA judges the report, RoB 2 the conduct of a trial and GRADE a body of trials. None of them judges whether the points, depth and number of sessions were well chosen, which is a clinical question, and a trial can satisfy all three while testing a treatment few practitioners would give. AcuiQ records what each study prescribed and links its citation; it does not rate the trial’s risk of bias. [Reading a protocol](/wiki/method/reading-a-protocol) explains the fields on an AcuiQ entry, and [What AcuiQ does not do](/wiki/method/what-acuiq-does-not-do) sets out its limits.

### Sources

1. MacPherson H, Altman DG, Hammerschlag R et al., Revised STandards for Reporting Interventions in Clinical Trials of Acupuncture (STRICTA): extending the CONSORT statement, PLoS Medicine 2010 — https://doi.org/10.1371/journal.pmed.1000261
2. Ma B, Chen ZM, Xu JK et al., Do the CONSORT and STRICTA checklists improve the reporting quality of acupuncture and moxibustion randomized controlled trials published in Chinese journals? A systematic review and analysis of trends, PLoS One 2016 — https://doi.org/10.1371/journal.pone.0147244
3. MacPherson H, Vertosick E, Lewith G et al., Influence of control group on effect size in trials of acupuncture for chronic pain: a secondary analysis of an individual patient data meta-analysis, PLoS One 2014 — https://doi.org/10.1371/journal.pone.0093739
4. Gang WJ, Xiu WC, Shi LJ et al., Factors Associated with the Magnitude Of acUpuncture treatment effectS (FAMOUS): a meta-epidemiological study of acupuncture randomised controlled trials, BMJ Open 2022 — https://doi.org/10.1136/bmjopen-2021-060237
5. Madsen MV, Gøtzsche PC, Hróbjartsson A, Acupuncture treatment for pain: systematic review of randomised clinical trials with acupuncture, placebo acupuncture, and no acupuncture groups, BMJ 2009 — https://doi.org/10.1136/bmj.a3115
6. Takakura N, Yajima H, A double-blind placebo needle for acupuncture research, BMC Complementary and Alternative Medicine 2007 — https://doi.org/10.1186/1472-6882-7-31
7. Higgins JPT, Savović J, Page MJ, Elbers RG, Sterne JAC, Chapter 8: Assessing risk of bias in a randomized trial, Cochrane Handbook for Systematic Reviews of Interventions — https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-08
8. Dworkin RH, Turk DC, McDermott MP et al., Interpreting the clinical importance of group differences in chronic pain clinical trials: IMMPACT recommendations, Pain 2009 — https://doi.org/10.1016/j.pain.2009.08.019
9. Sterne JAC, Savović J, Page MJ et al., RoB 2: a revised tool for assessing risk of bias in randomised trials, BMJ 2019 — https://doi.org/10.1136/bmj.l4898
10. Li J, Hui X, Yao L et al., The relationship of publication language, study population, risk of bias, and treatment effects in acupuncture related systematic reviews: a meta-epidemiologic study, BMC Medical Research Methodology 2023 — https://doi.org/10.1186/s12874-023-01904-w
11. Guyatt GH, Oxman AD, Vist GE et al., GRADE: an emerging consensus on rating quality of evidence and strength of recommendations, BMJ 2008 — https://doi.org/10.1136/bmj.39489.470347.AD
