Every assessment here ends in a score, and a score you cannot argue with is decoration. This is the whole rubric: the six dimensions, the weight each carries, and the wording behind every point on the 1-to-5 scale. The arithmetic is printed on each issue so you can redo it.
Each dimension is scored 1 to 5 by hand, against the anchors below. The composite is their weighted average.
| Dimension | What it asks | Weight |
|---|---|---|
| Source quality | How authoritative are the sources supporting this claim? | 25% |
| Data support | Is there quantitative data directly supporting this claim? | 20% |
| Reproducibility | Has this finding been confirmed by independent sources? | 20% |
| Consensus | What proportion of credible sources agree with this claim? | 15% |
| Recency | How current is the evidence? | 10% |
| Rigor | What is the methodological quality? | 10% |
Bands: strong 4.0 and above, moderate 3.0 to 3.9, mixed 2.0 to 2.9, weak below 2.0.
Every claim gets two composites, because is the effect real and how large is it are different questions and one number cannot answer both. A trial can settle the first and tell you nothing about the second — which is exactly what happened in issue one, where a double-blind 1,137-patient phase 3 crossed its prespecified threshold and released no effect size at all.
No weight is invented for the split. Each dimension is assigned to the question its own anchors actually answer, and the weights above are renormalised within each group so they still sum to 1. That is why the working on each issue is printed as a sum over a divisor — (3×.25 + 4×.20 + … ) ÷ .80 — rather than as rounded weights that would not reproduce the answer.
| Question | Dimensions it uses | Divisor |
|---|---|---|
| Is the effect real? | Source quality, reproducibility, consensus, recency, rigor | 0.80 |
| How large is it? | Data support | 0.20 |
That is a fact about this rubric and it is printed rather than hidden. Only data support's anchors speak about magnitude at all — 5 is specific statistical results, 1 is a purely qualitative assertion. Every other dimension asks about provenance, replication, currency or study design, all of which bear on whether an effect is real and none of which tell you how big it is.
The tempting move is to fold rigor in, since a phase 3 estimates an effect size more precisely than a case series. It is wrong: it would let a superb trial that published no numbers score above the floor for magnitude, which is the error the split exists to correct. A trial's quality governs how much you could learn from its numbers, not how much you have learned from numbers nobody released.
These are the words a score is measured against. They are quoted here as they are written in the scoring code, not paraphrased.
It is not a verdict on the treatment. Five of the six dimensions are about the evidence and how it reached us — who published it, whether anyone else has seen the same thing, how old it is. A trial can be excellent and score low here, because what is being scored is what has been released about it rather than the study itself. Issue one is exactly that case: rigor 5, data support 1.
It is also not precise. A weighted average of hand-assigned integers carries the precision of the integers, which is to say not much. It is published to two decimal places because that is what the arithmetic gives, not because the second decimal means anything. Two claims a tenth of a point apart are not distinguishable.
Until 3 September 2026 the issues referred to our published six-dimension rubric
and quoted its anchors, and the rubric was not published anywhere. Worse, one issue printed working whose weights were not the rubric's — it used 15% for reproducibility and recency instead of 20% and 10% — and reached 3.40 where the rubric gives 3.35. Both fall in the same band, which is why nobody noticed.
Later the same day an outside reviewer pointed out that the single composite breached our own standard: direction and magnitude are separate questions and one verdict cannot express both. That 3.35 was wrong about both halves — it understated the trial and overstated what was known about the size of what it found. Hence the two scores above. All three problems are corrected, and the rubric is here so the next such error is one a reader can catch.