Issue two Evidence review Published 28 August 2026 · Updated 29 August 2026 · guideline version 6.2026

The Category Difference

whatholdsup.org

Three drugs of the same class treat the same disease. A guideline lists all three as preferred, and grades the evidence behind one of them higher than the evidence behind the other two. That grade gets read as a ranking of the drugs — it was read that way by this page, in the version published on 28 August. It is not one. No trial has ever tested these drugs against each other, and every comparison this page examined of the two the guideline grades apart finds nothing between them.

What the guideline says

And the reason it gives.

The NCCN breast cancer guidelines, version 6.2026, cover first-line treatment of hormone-receptor-positive, HER2-negative advanced breast cancer. For an aromatase inhibitor combined with a CDK4/6 inhibitor, the guideline assigns ribociclib category 1. The same combination with abemaciclib or palbociclib is category 2A.

Category 1 means high-level evidence and uniform panel consensus. Category 2A means lower-level evidence with uniform consensus. The guideline states the basis for the difference: the overall survival benefit seen with ribociclib.

That is one of two systems NCCN runs, and it is not the one that says which drug to prefer. The Categories of Evidence and Consensus — the 1 and the 2A above — record how strong the evidence is that an intervention is appropriate. A separate set, the Categories of Preference, records the panel’s assessment of a regimen’s value — Preferred being defined as based on superior efficacy, safety, and evidence taken together. It is not an ordinal ranking and the guideline does not present it as one. On that second system the guideline lists all three aromatase-inhibitor-plus-CDK4/6-inhibitor combinations as preferred first-line options — for postmenopausal patients and for premenopausal patients on ovarian suppression alike — and explains the evidence grade elsewhere, in the categories-of-evidence discussion rather than beside the preference listing: Due to the clear OS benefit seen with ribociclib in combination with AI, it is a category 1 recommendation. Abemaciclib and palbociclib are category 2A. The preference designation does not differ across the three. Only the category of evidence does.

And there the guideline stops. We asked whether it anywhere tells a reader how to read a category 1 sitting beside a category 2A when all three regimens are equally preferred — whether it cautions, in the discussion, the principles section, a footnote or a panel comment, against taking the higher evidence grade for the better drug. It does not. The tables print the preference tier and the evidence category side by side and offer no interpretive guidance about what the difference between them means. That is not a criticism of the guideline, which is a reference document and not a reading primer. It is the answer to why this misreading is so easy: the document that draws the distinction carefully is silent on the one question a reader arrives with, and the silence is filled by whoever is reading.

The version of this assessment published on 28 August did not say that. It reported the category difference, asked what a reader should make of it, and never mentioned that the guideline itself had already answered the question elsewhere in the same document. That was a selective omission: the one fact most likely to change a reader’s conclusion, left out of a page built to criticise how the rest of the document reads. It is corrected here and recorded below. It is also, exactly, the failure this assessment is about — an evidence label read as a ranking of drugs, this time by us, after eighteen automated checks and two reviews on the question of whether it was one.

The distinction is specific to that setting, and it is worth being precise about. In combination with fulvestrant rather than an aromatase inhibitor, both ribociclib and abemaciclib are category 1, and palbociclib is not. So the claim is narrow: first-line, with an aromatase inhibitor, ribociclib alone carries category 1.

That second setting is worth pausing on, because the guideline gives the same reason for it: ribociclib and abemaciclib with fulvestrant are category 1 because those combinations showed an overall survival benefit, in MONALEESA-3 and MONARCH 2. Which means abemaciclib is a category 1 drug in one setting and a category 2A drug in another — on the same evidence standard, consistently applied. The grade is not a verdict on the molecule. It is a record of which of its trials produced a significant survival result, and the guideline says so plainly.

The guideline also records that the three drugs have not been directly compared in clinical trials.

The four trials

Each tested one drug against endocrine therapy alone. None tested one against another. Every figure is the trial publication’s own and every cell names the analysis it comes from. These trials have each reported several times; a table that says only “the latest” invites the reader to compare one trial’s first readout against another’s fourth.

TrialDrugProgression-free survivalOverall survival
MONARCH 3abemaciclibnot reached vs 14.7 mo
HR 0.54 (0.41–0.72)
primary analysis, JCO 2017
66.8 vs 53.7 mo
HR 0.80 (0.64–1.02)
final analysis, 8.1 y, Ann Oncol 2024 · two-sided P = .0664 against a final-analysis threshold of .034not significant
MONALEESA-2ribociclib
postmenopausal
25.3 vs 16.0 mo
HR 0.57 (0.46–0.70)
updated analysis, Ann Oncol 2018
63.9 vs 51.4 mo
HR 0.76 (0.63–0.93)
final analysis, 6.6 y, NEJM 2022 · p = 0.008 — significant
MONALEESA-7ribociclib
pre/perimenopausal, endocrine therapy including tamoxifen
23.8 vs 13.0 mo
HR 0.55
primary analysis, Lancet Oncol 2018
not reached vs 40.9 mo
HR 0.71 (0.54–0.95)
protocol-specified interim analysis, 34.6 mo, NEJM 2019 · one-sided P = .00973 against a prespecified stopping boundary of .01018crossed it
Later exploratory extended follow-up (53.5 mo): 58.7 vs 48.0 mo, HR 0.76 (0.61–0.96)
PALOMA-2palbociclib24.8 vs 14.5 mo
HR 0.58 (0.46–0.72)
primary analysis, NEJM 2016
53.9 vs 51.2 mo
HR 0.96 (0.78–1.18)
final analysis, 90.1 mo, JCO 2024 · one-sided P = .34 — not significant

On the endpoint every one of these trials was powered to measure, nothing separates them. Among the three postmenopausal aromatase-inhibitor trials — the three the guideline lists side by side — the progression-free hazard ratios are 0.54, 0.57 and 0.58. MONALEESA-7, in a different population, reports 0.55. Nor does the choice of readout rescue a difference. Across the investigator-assessed progression-free analyses these four trials have published between them — primary, updated and final, eight readouts, all named in the sources below — every hazard ratio falls between 0.52 and 0.58. That is a claim about the eight we checked, not about every analysis ever run on these datasets.

That comparison is across trials, and a comparison across trials is weak evidence whichever endpoint it uses. It is offered here for exactly what it is worth, which is the same as what the overall-survival comparison below is worth: these trials were never designed to be read against each other.

The guideline does not make that mistake, and it is worth being exact about what it does instead. It grades each regimen on its own trial, against that trial’s own significance threshold — a per-trial test, not a comparison of effect sizes. Then it prints the resulting grades in one table, in one column, beside each other. Nobody performed a cross-trial comparison; the format performs one on the reader’s behalf, and a category 1 sitting next to a category 2A is read as a ranking of the drugs whether or not it was meant as one.

The p-values in that table are not on one scale Each is reported as its own publication reports it, and they were not all tested against the same number. MONALEESA-2’s final analysis prints p = 0.008. We have not established the direction of that test, or the boundary it faced. An earlier version of this page called the p-value two-sided; we cannot source that, and a reader who reviewed this page after publication contends it was one-sided against an adjusted boundary. The paper’s statistical section is behind a wall we could not open, so neither reading is printed here as fact: the number stands as the publication states it, with nothing added to it. If it too spent alpha at an interim, its bar also sat below .05. PALOMA-2 states a one-sided p at an alpha of 0.025, which puts the critical value in the same place on the benefit side — a result that clears one clears the other — while differing in what the test was capable of concluding at all: a one-sided test is built to detect benefit and cannot declare harm. So on the narrow question of how hard it was to reach significance, PALOMA-2’s bar sat where the conventional one sits — while the test itself was not the conventional test, because it was built only to look for benefit. The other two faced bars that did not sit there at all. MONALEESA-7’s is an interim result against a group-sequential stopping boundary of .01018. MONARCH 3’s is two-sided and was judged against .034, not .05: its cumulative two-sided type I error of 0.05 was divided between the intention-to-treat population and a visceral-disease subgroup, and spent across interim and final analyses under an O’Brien-Fleming spending function. Whatever the order of those two operations — and an earlier version of this page asserted an order it could not source — the arithmetic that reaches the reader is the same: the final overall-survival analysis faced .034, not .05.

That does not rescue the result: .0664 misses .05 as well as .034. But it is worth seeing plainly, because this whole piece is about a grade that turns on a threshold, and the trials were not held to one threshold. Lining four p-values up in a column invites a reader to rank them, and they cannot be ranked. What each supports is only the binary its own trial reported: whether that trial cleared its own bar. MONALEESA-7’s is one-sided: the ASCO 2019 abstract of the same analysis states that “statistical comparison was made by 1-sided stratified log-rank test”. An earlier version of this page said the direction could not be determined, which was not true — it could not be determined from the journal paper we could open, and treating that as unknowable was turning a number we had not found into evidence that did not exist.

The grade turns instead on overall survival, and the three trials did not design that analysis the same way. PALOMA-2 powered for it: its final survival analysis was planned after at least 390 events, with 80% power to detect a hazard ratio of 0.74 or better at a one-sided 0.025. MONARCH 3 did not: survival there was a gated secondary endpoint, its final analysis planned at about 315 events with alpha split between the whole population and a visceral-disease subgroup. So palbociclib's null result comes from an analysis built to find an effect of that size, and abemaciclib's near-miss comes from one that was not. An earlier version of this page said none of these trials was powered for survival, which is false and flattened exactly this difference; an outside reviewer caught it.

Survival was a secondary endpoint in all three — which, as the paragraph above shows, covers both a powered analysis and an unpowered one. That is not a comment on how much survival matters — it is the outcome that matters most, and this page is not going to imply otherwise while criticising others for the reverse. It is a statement about statistical power. A trial sized to detect a difference in progression is not necessarily sized to detect one in survival, so a survival result from such a trial can fail to reach significance while the underlying benefit is real, and can reach it while the benefit is smaller than the point estimate suggests. That is the machinery the grade is resting on.

Where the intervals fall

Overall survival, first-line, postmenopausal, with an aromatase inhibitor, against the no-effect line. MONALEESA-7 is not on this chart: it enrolled pre/perimenopausal women on ovarian suppression with an aromatase inhibitor or tamoxifen, and putting it beside three postmenopausal aromatase-inhibitor trials would be the cross-trial comparison this piece spends its second half warning about.

AbemaciclibMONARCH 3 · HR 0.80
RibociclibMONALEESA-2 · HR 0.76
PalbociclibPALOMA-2 · HR 0.96
00.51.0 — no effect1.5
The distance that separates a category 1 from a category 2A Abemaciclib’s hazard ratio is 0.80, on an interval running 0.64 to 1.02. Ribociclib’s is 0.76, on 0.63 to 0.93. The point estimates are four hundredths apart — the same spread the whole class shows on progression-free survival. The intervals overlap along almost their entire length: 97% of ribociclib’s interval lies inside abemaciclib’s. That overlap is also a function of how wide abemaciclib’s interval is — itself a consequence of what that trial was powered to detect — and not only of how close the two point estimates are. The same caution this page applies to thresholds applies to it. The figure is ours, not a published one, and it is one subtraction, on the same rounded bounds shown above: ribociclib’s interval runs 0.63 to 0.93, a width of 0.30; the part of it inside abemaciclib’s runs 0.64 to 0.93, a width of 0.29. 0.29 ÷ 0.30 = 96.7%. Using abemaciclib’s unrounded lower bound of 0.637 in place of the rounded 0.64, the same subtraction gives 97.7%. We print 97%. (0.637 and 1.015 are abemaciclib’s bounds throughout; ribociclib’s are 0.63 and 0.93.)

One interval stopped at 0.93 and the other at 1.015, and that is the whole of it: ribociclib’s result reached statistical significance (p = 0.008) and abemaciclib’s did not (P = .0664, against its own boundary of .034). The guideline applied each trial’s result correctly. What separates a category 1 from a category 2A here is eight and a half hundredths in where an upper bound landed, on point estimates four hundredths apart. Those are the published bounds, 0.93 and 1.015. This page rounds to two decimal places elsewhere and prints 1.02, which is correct rounding; here, where the whole argument is about where a bound landed, the rounded figure was doing work the published one should do. A gate run made that objection eight times before we accepted it. That bound is not arbitrary and this page is not calling it that: where an interval ends is set by how many patients were enrolled, how many events occurred and how variable they were. It is a real measurement of how confidently each trial spoke. Nor is this observation ours: Tanguy and colleagues made it formally in npj Breast Cancer in 2018, before any of these trials had mature survival data, arguing that MONARCH 3 was less powered than PALOMA-2 and MONALEESA-2 to detect a survival difference and that divergent significance across the three might reflect chance rather than different drug efficacy. We found that paper through a fact-check run on this page, after making the argument, and it belongs here with their names on it. The question is whether a grade should record the confidence of one trial or the effect of a drug, when the two trials asked nearly the same question and got nearly the same answer.

Whether that should carry a category difference is the question this piece puts. It is not a question about the arithmetic, which is right, or about the classification, which follows the rules. It is a question about what a category is for — whether it should record which trial cleared a threshold, or what the evidence says about the drug.

Palbociclib is a different case, and the difference is real. A hazard ratio of 0.96 on an interval from 0.78 to 1.18 is not a near miss. It is a result compatible with a meaningful benefit at one end and with harm at the other, centred almost exactly on no effect — which is not the same as proof that the drug does nothing, and this page has just spent a paragraph saying so about the other two. What PALOMA-2 established is that it did not demonstrate a survival benefit. That is all it established, and it is less than “palbociclib does not work”. PALOMA-2 is also the trial whose investigators recorded a large and disproportionate imbalance in missing survival data between its arms — survival status was unknown for 13.3% of the palbociclib arm and 21.2% of the placebo arm — which they said limits interpretation. They then did something about it. A sensitivity analysis using recovered data cut the unknowns to 9.2% and 11.7% and moved the hazard ratio from 0.96 to 0.92 (0.76–1.12). Four hundredths — and this piece has spent several hundred words on what four hundredths does and does not establish, so it will not now call that no movement. What did not change is the conclusion: the interval still runs through 1.0 and the investigators still reported no survival benefit. The defect is real, and their own remedy points where their headline number already pointed. A different number, from a trial with a defect its own authors flagged, is still not a measurement of a different drug — but this piece is not claiming palbociclib’s result sits where abemaciclib’s does. It does not.

What this comparison cannot show Overlapping intervals do not prove two effects are equal, any more than one interval crossing the no-effect line proves two effects differ. What can honestly be said is narrower: these four trials, read against each other, do not establish a difference between the drugs — they were never designed to be read against each other, and a comparison they were not built to support is weak evidence rather than none. That is a statement about these four trials and not about the whole record — the indirect comparisons in the next section are evidence bearing on whether a difference exists, and they are weaker evidence than a randomised head-to-head trial would be.

What a hazard ratio is not

Everything above is written in hazard ratios, because that is how the trials report. It is not how anyone lives.

A hazard ratio of 0.80 does not mean 20% fewer deaths. It does not mean a 20% better chance of anything, and it is not a probability that applies to a person. It is a ratio of rates: over the course of follow-up, among patients still alive at each moment, deaths occurred in the treated group at about four-fifths the rate of the comparison group. It is an average of that ratio across years of follow-up, and if the ratio changed over time — which it usually does — the single number conceals the change.

What a person wants to know is how much longer. The trials answer that too, in months, and the answer is worth setting beside the grade.

TrialMedian overall survivalDifferenceHazard ratioCategory
MONARCH 3 abemaciclib66.8 vs 53.7 mo13.1 months0.802A
MONALEESA-2 ribociclib63.9 vs 51.4 mo12.5 months0.761
PALOMA-2 palbociclib53.9 vs 51.2 mo2.7 months0.962A
The number a patient would ask for On the scale that means something to a person — months of median survival added over endocrine therapy alone — the category 2A drug’s trial reported the larger figure. 13.1 months for abemaciclib, 12.5 for ribociclib.

That difference is not real either. Six-tenths of a month, across trials with different patients, different follow-up and different sizes, is noise. That is the point. The two drugs are indistinguishable on the absolute scale as well as the relative one, and on the absolute scale the ordering happens to run the other way from the grade.

Two warnings about that table, both of which apply to us and not only to other people. These are medians, not lifespans: a median difference of 13.1 months is the gap between the times by which half of each group had died, and no individual is promised anything by it. And comparing medians across trials is a cross-trial comparison, with everything that costs — different patients, different follow-up, different trial sizes. That is not what this page criticises the guideline for; the guideline grades each trial on its own threshold. It is a weakness of this table, owned here, and the table is offered as an illustration of scale rather than a comparison of drugs.

Palbociclib is the case where the absolute and relative pictures agree. 2.7 months of median difference, a hazard ratio of 0.96, an interval running through the no-effect line, and a sensitivity analysis that moved the estimate to 0.92 without changing the conclusion. Whatever is true of the other two, PALOMA-2 did not show a survival benefit.

What happens when you do compare them

Four studies, by three methods. On palbociclib they contradict each other. On the two drugs the guideline grades apart, they do not.

No randomised trial has tested one of these drugs against another. Several research programmes have compared them indirectly or observationally. This page examines four of them. It does not claim to have the whole literature — an outside reviewer found two more that we had missed, one of which is now the fourth below.

A network meta-analysis (Scientific Reports, February 2024) pooled seven phase III randomised trials, 4,415 patients, at a median follow-up of 73.3 months. It found no statistically significant difference in overall survival between any pair of the three. Its MONARCH 3 input is the final overall-survival result — HR 0.804 (0.637–1.015), at 97.2 months, the figure given above — not the earlier interim one.

P-VERIFY took 9,146 patients from US oncology records and weighted them by inverse probability of treatment. On overall survival it found nothing anywhere: ribociclib versus palbociclib 0.98 (0.87–1.10, P = 0.75), abemaciclib versus palbociclib 0.95 (0.84–1.08, P = 0.43), abemaciclib versus ribociclib 0.97 (0.82–1.14, P = 0.70). A second paper from the same cohort found nothing on progression either. Its arms are very uneven — 6,831 patients on palbociclib, 1,279 on ribociclib, 1,036 on abemaciclib — which is what prescribing actually looked like and is also a limit on what the smaller arms can show. For the comparison this piece turns on, abemaciclib against ribociclib, the two arms are of similar size and the imbalance is not the issue; it is their size against palbociclib’s that limits the other two comparisons.

One thing about it a reader should have. Its full name is the Palbociclib Verifying Evidence of Real-world Impact study, and its own poster states that it was funded by Pfizer Inc., which makes palbociclib. Industry-funded real-world studies are ordinary, they are held to the same methodological standards as any other, and this does not make its findings wrong. The fact is recorded here and no inference is drawn from it.

And the same disclosure has to run the other way, or it is not a disclosure but a point-scoring device. We could not open PALMARES-2’s declaration of interests. That is a statement about our access and not about its investigators, of whom we know nothing and about whom this page will not speculate. What follows from it is only this: we have read one study’s funding and not the other’s, so neither is offered here as the disinterested one.

PALMARES-2 (Annals of Oncology, April 2025) followed 1,982 patients across eighteen Italian centres and weighted them the same way. On progression-free survival it found abemaciclib ahead of palbociclib at an adjusted hazard ratio of 0.76 (0.63–0.92), p = 0.004, and ribociclib ahead of palbociclib at 0.83 (0.73–0.95), p = 0.007. It also reports overall survival, as an exploratory endpoint, pointing the same way. This page does not print those survival figures. The secondary summaries we could reach disagree with each other about which numbers are the survival ones and which are the progression ones, the paper itself is behind a wall we could not open, and a figure we cannot attribute to a primary source is a figure this page does not carry. What the authors do say about that endpoint is that the data are immature — 464 events at a median 31.3 months — and that follow-up is badly uneven between the arms: 45.7 months on palbociclib against 25.2 on ribociclib and 22.4 on abemaciclib. That imbalance cuts against their own finding, and it is their caveat, not ours.

A fourth, an indirect comparison from reconstructed patient data (Cancers, 2023), rebuilt patient-level survival curves from PALOMA-2, MONALEESA-2 and MONARCH 3 — 1,827 patients — and estimated every pairwise difference directly. On overall survival it separates nobody: abemaciclib versus ribociclib 0.933 (0.753–1.157), p = 0.528. On progression-free survival its one-stage model puts abemaciclib against ribociclib at 0.722 (0.520–1.002), p = 0.051 — the closest any study has come to separating those two, and it does not, while its two-stage model on the same data gives 0.921 (0.597–1.420), p = 0.710. That borderline figure belongs in this piece precisely because it is the one number that cuts against its conclusion.

None of these is a randomised head-to-head trial and none should be read as one. Together they are the best comparative evidence in existence, and on palbociclib they contradict each other. The larger cohort, on both survival and progression, finds nothing. The smaller one finds palbociclib behind on progression, on intervals that exclude the no-effect line — abemaciclib’s by 0.08 and ribociclib’s by 0.05, which is significant but not comfortably so, and reports an exploratory survival result pointing the same way on data its own authors call immature. An earlier version of this page reconciled them by saying they had measured different endpoints. That was wrong — both programmes reported both endpoints — and it was the tidier story rather than the true one. Whether palbociclib is behind the other two is unsettled, and this page does not settle it.

What all four agree on The studies disagree about palbociclib. On abemaciclib against ribociclib — the comparison the category difference actually rests on — they do not disagree at all. The network meta-analysis finds no significant difference. The reconstructed patient-data comparison, on survival: 0.933, on 0.753 to 1.157, p = 0.528. P-VERIFY, on overall survival: 0.97, on 0.82 to 1.14, P = 0.70. And PALMARES-2 — the study that did separate palbociclib, on both endpoints — puts this pairing, on progression-free survival, at 0.91, on 0.73 to 1.14, p = 0.425. Its palbociclib figures are the 0.76 and 0.83 above; this one is the comparison it could not separate. PALMARES-2 publishes no abemaciclib-versus-ribociclib survival figure, so on that pairing the survival evidence is P-VERIFY’s and the meta-analysis’s.

Eleven thousand patients across two cohorts, two endpoints, and a disagreement between them about a third drug. Not one of them can tell abemaciclib and ribociclib apart. The guideline grades the evidence behind them apart.

They do not contradict the guideline’s own statement, and it is worth saying so plainly. In guideline language, “directly compared in clinical trials” means a randomised head-to-head trial, and none exists. The Flatiron study says as much itself: it opens by describing its own purpose as working in the absence of randomised trials that directly compare the three. The guideline is accurate as written. What these comparisons add is not a correction to it, but the only evidence there is on the question it records as untested.

Where the three genuinely differ

Three different burdens, from the labels. None of them is what the grade is about.

An earlier version of this page said cardiac monitoring was “the one clear difference” between the drugs, and that it ran the other way from the grade. That was selective. Each of the three carries its own labelled warning, and reading the labels side by side is the only fair way to do this.

Ribociclib is the only one requiring cardiac monitoring: an ECG before treatment and again at about day 14 of the first cycle. Its label also requires serum electrolytes “prior to the initiation” and at the beginning of the first six cycles, liver function tests every two weeks for the first two cycles, and blood counts on the same schedule.

Abemaciclib carries a labelled warning for venous thromboembolism — thrombotic events in 2% to 5% of patients treated across the trials the label reports, which include the early-breast-cancer trial as well as the two metastatic ones, with monitoring for signs and symptoms — and diarrhoea in 81% to 90% of patients across four trials, grade 3 in 8% to 20%, with patients instructed to begin antidiarrhoeal treatment at the first loose stool.

Palbociclib carries neutropenia: 80% of patients in PALOMA-2, grade 3 or worse in 66%, with blood counts required on day 15 of each of the first two cycles as well as at the start of every cycle.

These are three different experiences of taking a drug, and they do not line up with the grade in any direction. The category 1 drug is the one with the cardiac requirement. One category 2A drug carries a thrombosis warning and the other a two-in-three rate of grade 3 neutropenia. The point is not that one of them is worse. It is that the category records which trial produced a significant survival result, and the burden a patient actually carries sits in a different column entirely — one the grade does not look at.

This piece is not arguing that the three drugs are identical. They are not. It is arguing that the evidence does not establish that one of them works better than another.

What is established, and what is not

As of this draft.

Established

  • All three drugs substantially improve progression-free survival against endocrine therapy alone. In the three first-line postmenopausal aromatase-inhibitor trials the hazard ratios are 0.54, 0.57 and 0.58; MONALEESA-7, in pre/perimenopausal women on a different endocrine backbone, reports 0.55. Every one of them is statistically significant, and the last of them is not comparable with the other three.
  • Ribociclib plus letrozole improved overall survival against letrozole alone: 63.9 versus 51.4 months, HR 0.76 (0.63–0.93), p = 0.008, at 6.6 years of median follow-up. Among the three first-line postmenopausal aromatase-inhibitor trials, it is the only overall-survival result that reached significance. MONALEESA-7, in pre/perimenopausal women, also reached significance for first-line overall survival.
  • Abemaciclib plus a non-steroidal aromatase inhibitor did not: 66.8 versus 53.7 months, HR 0.80 (0.64–1.02), two-sided P = .0664 against a final-analysis threshold of .034, at 8.1 years of median follow-up. The point estimate favours abemaciclib; the interval crosses the no-effect line.
  • Palbociclib plus letrozole did not: 53.9 versus 51.2 months, HR 0.96 (0.78–1.18). A sensitivity analysis recovering missing survival data moved the point estimate to 0.92 (0.76–1.12) and left the conclusion where it was.
  • The guideline assigns ribociclib category 1 first-line with an aromatase inhibitor, and abemaciclib and palbociclib category 2A, on the stated basis of that survival result.
  • No randomised trial has compared any of the three against another.
  • None of the four comparative studies examined on this page separates abemaciclib from ribociclib: not the network meta-analysis, not P-VERIFY on 9,146 patients, not PALMARES-2 on 1,982, not the reconstructed patient-data comparison — whose progression-free estimate comes closest, at p = 0.051, and still does not reach significance. That is a statement about four studies we read, not about a literature we have surveyed to the end.
  • Each of the three carries a different labelled burden: ribociclib a cardiac monitoring requirement, abemaciclib a venous thromboembolism warning and diarrhoea in 81–90% of patients, palbociclib grade 3 or worse neutropenia in 66%. None of these is what the category measures.

Not established

  • That abemaciclib and ribociclib differ in effectiveness. No randomised trial has tested it, and none of the four comparative studies examined here separates them. Indirect comparisons are weaker evidence than a randomised trial would be, and four studies are not the whole literature.
  • Whether palbociclib is behind the other two. The two real-world programmes contradict each other on exactly this, on both endpoints, and this page does not resolve it. PALMARES-2’s survival data are immature and its arms unevenly followed; P-VERIFY is larger and is funded by palbociclib’s manufacturer. What is not in dispute is that PALOMA-2 produced a null survival result where the other two trials did not.
  • That the category difference predicts any difference in outcome for a patient.
  • That it reflects anything about the drugs beyond which trial’s survival analysis crossed its own significance threshold.
  • That the divergent overall-survival results reflect a property of the molecules rather than of the trials — their populations, their follow-up, their power and, in one case, their missing data.
  • What a head-to-head trial would find. Nobody has run one. Whether anybody will is not something this page knows — an earlier version asserted that nobody would, which was a prediction about the decisions of companies and funding bodies we have not asked and cannot read.

What a reader is owed

The part of this that is not academic.

If you are taking one of these drugs, or choosing between them, the question underneath all of this is whether a category 2A drug is a worse drug. There are two answers and they point the same way. The guideline’s own answer is no: it lists all three as preferred first-line options, and the category records how strong the evidence behind each one is, not which drug works better. The evidence’s answer is that nothing has been shown to separate them: no randomised trial has compared them, and none of the four comparative studies on this page can tell abemaciclib and ribociclib apart.

Which leaves the question of why a reader would have thought otherwise. Not because the guideline says so — it does not. Because a category 1 and a category 2A printed in one column look like a ranking, and the preference designation that would correct the impression sits somewhere else in the document. This page made that mistake in its first published version, which is the most direct evidence it can offer that the mistake is easy to make.

That is a statement about evidence and not a recommendation. This assessment does not say what anyone should take. Prescribing decisions turn on toxicity, monitoring burden, comorbidities, interactions and cost, none of which a page like this can see, and all of which your oncologist can.

Sources

Every figure above traces to one of these, and each cell of the table names which: a trial publication, a drug label, or a comparative study, cited by name in the entry below. Where a correction to one of these sources is known to us, it is noted in that source’s entry.

Guideline NCCN Clinical Practice Guidelines in Oncology — Breast Cancer, version 6.2026 Source of the category assignments, the stated reason for them, and the record that the three drugs have not been directly compared in clinical trials. Read directly by a human: the licence forbids putting the document through an AI tool, so no automated check on this page has seen it. Its own trial figures are accurate and are not used here — a guideline cites registrational readouts and rounds them, which is right for a guideline and wrong for a table about where intervals end.
Primary MONARCH 3 — abemaciclib as initial therapy. J Clin Oncol, 2017 Primary progression-free survival: median not reached vs 14.7 months, HR 0.54 (95% CI 0.41–0.72), P = .000021.
Primary MONARCH 3 — final overall survival. Ann Oncol, 2024 66.8 vs 53.7 months, HR 0.804 (95% CI 0.637–1.015), two-sided P = .0664, at 8.1 years median follow-up. Not statistically significant. A corrigendum to this paper existsAnn Oncol 2025;36:1556, doi 10.1016/j.annonc.2025.07.002 — and we have not read it. It is published open access; an earlier version of this page said it sat behind a paywall, which was wrong. Every retrieval route available to us returned a block, which is a limit of our tooling and not of the licence. What we can say is that every MONARCH 3 figure here matches the original article and the SABCS 2023 presentation of the same analysis. What we cannot say is whether the corrigendum touches any of them, because that requires reading it. Anyone who can open it, please do — corrections@whatholdsup.org. The trial maintained a cumulative two-sided type I error of 0.05 by the Lan-DeMets method with an O’Brien-Fleming spending function, with alpha divided between the intention-to-treat population and a visceral-disease subgroup by a prespecified graphical testing procedure. The threshold at this final analysis was .034, which is the figure the presentation gives and the only part of the design this page relies on.
Primary MONALEESA-2 — updated results. Ann Oncol, 2018 Progression-free survival 25.3 vs 16.0 months, HR 0.568 (95% CI 0.457–0.704), which this page rounds to 0.57 (0.46–0.70). The guideline prints the same medians against HR 0.56 (0.45–0.70). The page uses the publication’s figure, as it does throughout, and the difference is one of rounding rather than of data.
Primary MONALEESA-2 — overall survival with ribociclib plus letrozole. N Engl J Med, 2022 63.9 vs 51.4 months, HR 0.76 (95% CI 0.63–0.93), P = 0.008 as the paper prints it, at 6.6 years median follow-up. The direction of that test and the boundary it was judged against are not established by us. The only overall-survival result among the three first-line postmenopausal aromatase-inhibitor trials that reached significance.
Primary MONALEESA-7 — ribociclib plus endocrine therapy in premenopausal women. Lancet Oncol, 2018 Progression-free survival 23.8 vs 13.0 months, HR 0.55, P < .0001.
Primary MONALEESA-7 — overall survival. N Engl J Med, 2019 The protocol-specified interim analysis: median not reached vs 40.9 months, HR 0.71 (95% CI 0.54–0.95), one-sided P = .00973 against a prespecified stopping boundary of P = .01018, at 34.6 months median follow-up. The direction of the test is stated in the ASCO 2019 abstract of the same analysis — “statistical comparison was made by 1-sided stratified log-rank test” — which is where this page takes it from, the journal paper being behind a wall we could not open. The paper prints 0.71; the more precise 0.712 appears in conference reporting and not in the publication this page cites. A group-sequential test with an adjusted threshold, which is why its p-value cannot be read beside MONALEESA-2’s 0.008 as though the two numbers meant the same thing.
Primary MONALEESA-7 — updated overall survival. Clin Cancer Res, 2022;28:851 The exploratory extended follow-up at 53.5 months: 58.7 vs 48.0 months, HR 0.76 (95% CI 0.61–0.96). The authors describe it as exploratory in those words. It is not the figure this page leads with for MONALEESA-7 and it is not on the interval chart; the protocol-specified result above is.
Primary PALOMA-2 — palbociclib and letrozole. N Engl J Med, 2016 Primary progression-free survival: 24.8 vs 14.5 months, HR 0.58 (95% CI 0.46–0.72), two-sided P < .001. These are the figures the paper prints, and they are what the table shows. The unrounded 0.576 (0.463–0.718) appears in the 2019 extended-follow-up paper (Breast Cancer Res Treat), which restates the primary analysis — not in NEJM 2016.
Primary PALOMA-2 — final overall survival. J Clin Oncol, 2024 53.9 vs 51.2 months, HR 0.956 (95% CI 0.777–1.177), one-sided P = .34, at 90.1 months median follow-up. Source also of the survival analysis’s design: “The final OS analysis was planned to be performed after at least 390 events, providing an 80% power to detect a hazard ratio [HR] ≤0.74, using a stratified log-rank test with a one-sided significance level of 0.025.” Source also of the missing-data imbalance (unknown survival status 13.3% vs 21.2%) and of the investigators’ recovered-data sensitivity analysis (9.2% vs 11.7%; HR 0.92, 95% CI 0.76–1.12, P = .21).
Comparison Indirect comparison from reconstructed patient data. Cancers, 2023;15(18) Patient-level survival curves rebuilt from PALOMA-2, MONALEESA-2 and MONARCH 3 by graphical reconstruction, 1,827 patients, every pair estimated directly. Overall survival: ribociclib vs palbociclib 0.903 (0.746–1.094, p = 0.297); abemaciclib vs palbociclib 0.843 (0.690–1.030, p = 0.094); abemaciclib vs ribociclib 0.933 (0.753–1.157, p = 0.528). Progression-free survival, one-stage model: abemaciclib vs ribociclib 0.722 (0.520–1.002, p = 0.051); two-stage model on the same data, 0.921 (0.597–1.420, p = 0.710). This page did not have this study until an outside reviewer supplied it on 2026-08-29.
Comparison PALMARES-2 — real-world comparison of first-line palbociclib, ribociclib and abemaciclib. Ann Oncol, April 2025 1,982 patients, eighteen Italian centres, inverse-probability weighting. Progression-free survival: abemaciclib vs palbociclib aHR 0.76 (0.63–0.92), p = 0.004; ribociclib vs palbociclib 0.83 (0.73–0.95), p = 0.007; abemaciclib vs ribociclib 0.91 (0.73–1.14), p = 0.425. Overall survival is reported as an exploratory endpoint and its figures are not used on this page: the summaries we could reach disagree about which hazard ratios belong to survival and which to progression, and the paper is not open to us. The authors describe the survival data as immature — 464 events at a median 31.3 months — with follow-up of 45.7 months on palbociclib against 25.2 on ribociclib and 22.4 on abemaciclib. This is the study that separates palbociclib from the other two, on both endpoints. Note for anyone checking: the ASCO 2024 conference presentation of the same study gave 0.91 (0.70–1.19) for abemaciclib versus ribociclib; the figures above are the published paper’s.
Label KISQALI (ribociclib) — US prescribing information Source of the cardiac requirement: ECG before starting and at about day 14 of the first cycle; serum electrolytes prior to initiation and at the beginning of the first six cycles; liver function tests and blood counts every two weeks for the first two cycles, then at the start of each subsequent four.
Label VERZENIO (abemaciclib) — US prescribing information Source of the venous thromboembolism warning — events in 2% to 5% of patients treated across the trials the label reports — the early-breast-cancer trial as well as the metastatic ones — with monitoring for signs and symptoms — and of diarrhoea in 81% to 90% of 3,691 patients across four trials, grade 3 in 8% to 20%.
Label IBRANCE (palbociclib) — warnings and precautions Source of neutropenia in 80% of PALOMA-2 patients, grade 3 or worse in 66%, and of the requirement for blood counts on day 15 of each of the first two cycles.
Comparison Network meta-analysis of CDK4/6 inhibitors. Scientific Reports, February 2024 Seven phase III trials, 4,415 patients, 73.3 months median follow-up. No significant pairwise overall-survival difference. Its Table 1 lists MONARCH 3 at HR 0.804 (0.637–1.015) with “year of updated data 2023”, and its reference list cites the SABCS 2023 final overall-survival abstract — the final analysis, not the interim. (The 73.3 months above is the pooled median across the seven trials, over a range of 48.7 to 97.2 months.) An earlier version of this page stated the opposite, on the strength of our own automated check rather than the paper; an outside reviewer opened the table and corrected us.
Comparison P-VERIFY — comparative overall survival of CDK4/6 inhibitors plus an aromatase inhibitor, US real-world setting, 2025 The Palbociclib Verifying Evidence of Real-world Impact study. 9,146 patients from the Flatiron Health electronic-health-record-derived deidentified longitudinal database — 6,831 on palbociclib, 1,279 on ribociclib, 1,036 on abemaciclib — with stabilised inverse-probability weighting. Overall survival: ribociclib vs palbociclib aHR 0.98 (0.87–1.10, P = 0.7531); abemaciclib vs palbociclib 0.95 (0.84–1.08, P = 0.4292); abemaciclib vs ribociclib 0.97 (0.82–1.14, P = 0.6956). The study’s own SABCS 2024 poster states: “This study was funded by Pfizer Inc.” Pfizer makes palbociclib. A second paper from the same cohort reports real-world progression-free survival and likewise finds no significant differences between the three. The paper describes its own purpose as working in the absence of randomised trials that directly compare the three — which is why this page does not call it a direct comparison.