We assess evidence on contested questions. That is worth nothing unless you know who funds it, what they can influence, and how often we turn out to be wrong. All of it is below.
Three things with three different jobs, and it matters that they are not the same thing.
CivicScale sells five commercial products — Parity Health, Employer, Broker, Provider and Billing — to individuals, employers, benefits brokers, medical practices and billing companies. They analyse medical bills, benchmark health plans against public Medicare rates, and audit payer remittances. That is how the money is made. There is no outside investment. No venture capital, no strategic investors, no undisclosed backers.
Because you would find it anyway, and finding it yourself is worse.
CivicScale sells tools that help employers and providers pay less for healthcare. This publication assesses the evidence on healthcare — including drug pricing, treatment efficacy and coverage policy.
Some of the companies whose evidence we assess operate in markets where CivicScale's customers have a direct financial stake in the answer. An evidence publication with an undisclosed commercial parent is worth less than nothing, because a reader who discovers it later is right to discount everything they read before.
Separating the publication onto its own domain does not dissolve that conflict. It makes it legible. What follows is what we actually do about it.
Paid plans include the ability to request an investigation. That needs a hard boundary.
Requested topics pass an automated scope screen before entering the queue. A request that is not an answerable evidence question is rejected with a reason. Being paid for does not exempt it.
Once a topic is opened, the customer who requested it has no further involvement. They see the finished assessment when everyone else does. If the evidence goes against the position they hoped for, that is what publishes.
Including the parts that are automated, which is most of them.
Claims are extracted from public sources — peer-reviewed research, regulatory filings and government data — and each is scored from 1 to 5 on six dimensions: source quality, data support, reproducibility, consensus, recency and rigor. A published weighted formula combines them into the composite you see.
Scores are generated by AI models. The formula, the weights and the 1–5 anchor definitions for every dimension are published in full, so you can disagree with a score on the record rather than guessing how it was reached. Every topic passes a human quality review before it appears publicly.
No score is hand-adjusted to reach a preferred answer. Where we think an assessment is wrong, the fix is to correct the inputs or the methodology — both visible — never to override the output.
The section most publications do not have.
On 27 August 2026 we checked all 53 published assessments against what our own method produces today. 46 still match. 7 do not.
Here is why that happens, because the mechanism matters more than the number. Assessments are produced by an AI model, and the model behind the name we call changes over time without our code changing at all. One assessment — on GLP-1 drug pricing — had been published as “debated” since March. Re-running the identical method over the identical evidence in August returned “consensus”, and reading the underlying claims confirms August is right: they are prices, spending figures and coverage rules, with no opposing position among them.
That label was wrong for five months and nothing in our system was capable of noticing. We found it by accident, while building something else.
Those 52 assessments are now frozen as a reference set and re-checked automatically every week, and within minutes of any change to the method. A moved judgment fails a test and a human reads it. The worst case for noticing is now seven days rather than five months.
We are also rebuilding how a verdict is reached, so that each is derived two independent ways and we publish no verdict at all when the two disagree. That work is not finished, and this page will say so until it is.
The five known-stale assessments have not yet been corrected on the site, because the rebuild may change what a correction should look like. When each is corrected, the change and the date it happened will appear in that topic’s revision history, which is published alongside the assessment.
You will at some point. We would rather hear it.
Evidence assessments are claims about the world, and claims about the world can be checked. Write with the topic, the specific claim or score, and what you think is wrong with it.
Acknowledged within 48 hours. Resolved, or explained why not, within 10 business days.
Substantive challenges get a substantive answer. Where a challenge is right, the assessment changes, the change is logged with its date, and the correction says what it was responding to. Where we disagree, we say why rather than going quiet.
Responsibility for this policy sits with CivicScale as the publisher. Corrections: corrections@whatholdsup.org