The rubric

Six dimensions, and what each number means

Every assessment here ends in a score, and a score you cannot argue with is decoration. This is the whole rubric: the six dimensions, the weight each carries, and the wording behind every point on the 1-to-5 scale. The arithmetic is printed on each issue so you can redo it.

The weights

Each dimension is scored 1 to 5 by hand, against the anchors below. The composite is their weighted average.

DimensionWhat it asksWeight
Source qualityHow authoritative are the sources supporting this claim?25%
Data supportIs there quantitative data directly supporting this claim?20%
ReproducibilityHas this finding been confirmed by independent sources?20%
ConsensusWhat proportion of credible sources agree with this claim?15%
RecencyHow current is the evidence?10%
RigorWhat is the methodological quality?10%

Bands: strong 4.0 and above, moderate 3.0 to 3.9, mixed 2.0 to 2.9, weak below 2.0.

Two scores, not one

Every claim gets two composites, because is the effect real and how large is it are different questions and one number cannot answer both. A trial can settle the first and tell you nothing about the second — which is exactly what happened in issue one, where a double-blind 1,137-patient phase 3 crossed its prespecified threshold and released no effect size at all.

No weight is invented for the split. Each dimension is assigned to the question its own anchors actually answer, and the weights above are renormalised within each group so they still sum to 1. That is why the working on each issue is printed as a sum over a divisor — (3×.25 + 4×.20 + … ) ÷ .80 — rather than as rounded weights that would not reproduce the answer.

QuestionDimensions it usesDivisor
Is the effect real?Source quality, reproducibility, consensus, recency, rigor0.80
How large is it?Data support0.20
The magnitude score rests on one dimension

That is a fact about this rubric and it is printed rather than hidden. Only data support's anchors speak about magnitude at all — 5 is specific statistical results, 1 is a purely qualitative assertion. Every other dimension asks about provenance, replication, currency or study design, all of which bear on whether an effect is real and none of which tell you how big it is.

The tempting move is to fold rigor in, since a phase 3 estimates an effect size more precisely than a case series. It is wrong: it would let a superb trial that published no numbers score above the floor for magnitude, which is the error the split exists to correct. A trial's quality governs how much you could learn from its numbers, not how much you have learned from numbers nobody released.

The anchors

These are the words a score is measured against. They are quoted here as they are written in the scoring code, not paraphrased.

Source quality 25%

  1. FDA approval/label, major peer-reviewed RCT (NEJM, Lancet, JAMA)
  2. Government agency report (CMS, CDC, WHO), systematic review
  3. Peer-reviewed observational study, reputable health policy analysis (KFF, RAND)
  4. News reporting on primary source, industry press release, advocacy group
  5. Opinion, social media, non-peer-reviewed preprint

Data support 20%

  1. Specific statistical results (p-values, confidence intervals, effect sizes)
  2. Specific numbers (percentages, dollar amounts, patient counts)
  3. Quantitative direction with approximate magnitudes
  4. Qualitative assertion with indirect numeric context
  5. Purely qualitative assertion with no numeric support

Reproducibility 20%

  1. Multiple independent RCTs or large studies confirm the finding
  2. At least 2 independent sources with consistent findings
  3. One strong primary source plus corroborating secondary sources
  4. Single source, not yet independently confirmed
  5. Preliminary or contested finding

Consensus 15%

  1. Universal agreement across all relevant source types
  2. Strong majority agreement with minor variation in specifics
  3. General agreement but some credible dissent or caveats
  4. Divided — credible sources on both sides
  5. Minority position or actively debated

Recency 10%

  1. Published within the last 12 months
  2. Published 1-2 years ago
  3. Published 2-4 years ago
  4. Published 4-6 years ago
  5. Published more than 6 years ago

Rigor 10%

  1. Phase 3 RCT, large N, pre-registered, peer-reviewed
  2. Well-designed controlled study or comprehensive regulatory review
  3. Observational study with appropriate controls, or rigorous policy analysis
  4. Descriptive analysis, market report, or uncontrolled comparison
  5. Anecdotal, opinion-based, or methodologically flawed

What the score is not

It is not a verdict on the treatment. Five of the six dimensions are about the evidence and how it reached us — who published it, whether anyone else has seen the same thing, how old it is. A trial can be excellent and score low here, because what is being scored is what has been released about it rather than the study itself. Issue one is exactly that case: rigor 5, data support 1.

It is also not precise. A weighted average of hand-assigned integers carries the precision of the integers, which is to say not much. It is published to two decimal places because that is what the arithmetic gives, not because the second decimal means anything. Two claims a tenth of a point apart are not distinguishable.

Why this page exists

Until 3 September 2026 the issues referred to our published six-dimension rubric and quoted its anchors, and the rubric was not published anywhere. Worse, one issue printed working whose weights were not the rubric's — it used 15% for reproducibility and recency instead of 20% and 10% — and reached 3.40 where the rubric gives 3.35. Both fall in the same band, which is why nobody noticed.

Later the same day an outside reviewer pointed out that the single composite breached our own standard: direction and magnitude are separate questions and one verdict cannot express both. That 3.35 was wrong about both halves — it understated the trial and overstated what was known about the size of what it found. Hence the two scores above. All three problems are corrected, and the rubric is here so the next such error is one a reader can catch.