There is a widely held sense that using AI is making us worse at thinking, and that there is now real evidence for it. The evidence usually offered is weak. The best evidence anyone has is real, and almost nobody outside medicine has met it with its qualifications intact: in four Polish endoscopy centres, the rate at which experienced doctors found precancerous growths without AI fell from 28.4% to 22.4% after their units adopted it. That study is retrospective, confounded, and its authors list eleven limitations. Three named academics said so publicly the day it appeared, and some of that did travel. What did not is the part that tells you how loosely to hold the number.
Six percentage points, in colonoscopies performed with no AI at all.
Four endoscopy centres in Poland adopted computer-aided polyp detection at the end of 2021. Budzyń and colleagues went back through the records and compared the three months before with the three months after, looking only at the colonoscopies performed without AI assistance in both periods — 1,443 of them, 795 before and 648 after.
The adenoma detection rate — the share of procedures in which at least one precancerous growth is found, and the standard quality measure for the procedure — fell from 28.4% (226 of 795) to 22.4% (145 of 648). That is a drop of 6.0 percentage points, 95% confidence interval −10.5 to −1.6 percentage points — both ends below zero — at P = 0.0089. The nineteen endoscopists involved were not novices: each had performed more than 2,000 colonoscopies, and they averaged 27.6 years since graduation, over a range of 8 to 39.
It is a retrospective observational study, and that governs everything you can take from it. Nobody was randomised. The authors list eleven limitations of their own, and the largest is not subtle: the total number of procedures in the second period nearly doubled, from 795 to 1,382 — while the non-AI procedures actually measured fell from 795 to 648 — because Poland suspended its colonoscopy-based screening programme on 1 January 2022 and the centres used the freed capacity to draw patients off the waiting list. The population being examined plausibly changed. Withdrawal time — how long the endoscopist spends inspecting on the way out, which strongly affects what is found — was not recorded. There was no blinding. The first 100 procedures per centre after installation were excluded to allow for familiarisation. And there were too few procedures per operator to analyse any individual endoscopist reliably.
The authors' own conclusion is hedged accordingly: continuous exposure to AI may reduce
the detection rate of standard colonoscopy, and they call for robust prospective studies such as randomized crossover trials
. Their December 2024 preprint had said reduced
, and a negative impact on endoscopist performance
; the version that came through peer review says may reduce
, and behaviour
. Both versions are public, so you can watch review do its work.
And the caveats were not hidden. On 12 August 2025, the day the study published, the Science Media Centre put three named academics on the record for journalists. Professor Venet Osmani of Queen Mary University of London named the volume problem exactly — the number of colonoscopies performed nearly doubled after the AI tool was introduced, going from 795 to 1382
— and argued that three months is too short a window for genuine skill loss in experienced clinicians. Professor Allan Tucker of Brunel noted that only one AI system was studied and that randomised crossover trials are needed for firmer claims. Dr Catherine Menon of the University of Hertfordshire took the underlying concern seriously.
Twelve studies, a consistent direction, a null we were offered that is not one, and a null we missed that is — whose authors explain why it does not cancel the first result.
Heudel and colleagues at the Centre Léon Bérard published a scoping review of exactly this question in 2026. It assessed 65 articles in full and included twelve studies. (We give the number that its own flow diagram supports: 65 assessed, 53 excluded with reasons, 12 included. The earlier steps of that diagram do not reconcile — 432 database records and 41 from grey literature, less 37 duplicates, is 436 rather than the 373 printed, and 373 less 332 excluded is 41 rather than the 65 it then reports assessing. We say so because we are about to use this review's count, and because a page that checks other people's arithmetic should show its own.) Their own summary of the state of the evidence, in the abstract, is that Evidence of clinical deskilling, though scarce, is consistent across specialties
, and, in the body, that the number of high-quality empirical studies remains limited
, with many reports resting on expert opinion or small qualitative work rather than large multicentre trials.
So, within the twelve studies that review includes: not one study, and not a settled literature either. What it assembles is a consistent direction carried by very few measurements. Among them, and each read here at source rather than through the review’s summary of it:
Radiology. Wang and colleagues ran a randomised crossover study of 40 clinicians diagnosing anterior cruciate ligament tears on MRI. AI assistance raised overall diagnostic accuracy from 87.2% to 96.4%. It also accounted for the errors that remained: 45.5% of the errors made under AI assistance were attributable to automation bias, across all levels of experience. When the authors tested selectively withholding AI outputs that were likely to mislead, automation bias fell by 41.7%.
Pathology. Rosbach and colleagues put 28 trained pathology experts through a diagnostic task with an AI decision-support system, some of whose advice was wrong. Where it was, 7% of judgments the pathologists had initially got right were overturned by it. Time pressure did not make that happen more often; it made it worse, with heavier reliance on the system's negative calls and a measurable fall in performance.
Structural rather than cognitive. England's move to HPV primary screening cut cytology workload by about 70%, on Public Health England's estimate as reported by the Royal College of Pathologists, and the number of accredited laboratories fell from 48 to 8. Nobody's skill decayed; there was simply far less of the work on which the skill is built. (The scoping review puts the cut at 80–85% and the laboratories at 45 consolidating to 8 — citing, in its body, an NHS Digital annual report and a UK National Screening Committee evidence review, and, in its table, Rebolj and colleagues in Cancer Cytopathology, which we could not reach at the publisher, on PubMed or in PMC. The two counts differ and we cannot say why: the College's report and the review's source disagree about the starting number, and NHS Digital's own annual report says 48 consolidated to 8 without mentioning any earlier closures. We give the College's figures because those are the ones we read at source.)
And the study the review offers as pointing the other way — which is not what that study says. Heudel and colleagues summarise Savardi and colleagues as showing 32 radiology residents gaining 12 percentage points of sensitivity, reading 18% faster, and holding that performance at a three-month follow-up. We went to the paper because a check on this page requires every name to be verified against the author list. It has eight radiologists in training, not 32. Its outcome measures are scoring error against a consensus standard and agreement between readers, not sensitivity or reading time. And it contains no three-month follow-up at all.
What the paper does show is worth having. Eight trainees scored 150 chest X-rays for COVID-19 severity across three conditions — no AI, AI on demand, and AI integrated into the reporting system. Agreement between readers rose from an intraclass correlation of 0.665 without AI to 0.813 with it integrated. Residents asked for the AI in 70% of cases. And when the AI's errors exceeded an acceptability threshold, they overrode it. The authors' own framing is that balancing educational benefit against deskilling risk is essential.
And there is a null result. We missed it, and an outside reviewer found it. On 3 June 2026 — nearly three months before this page was drafted — Pedersen and colleagues published a prospective, multicentre, registry-based trial in Endoscopy that did the thing this page says the field needs: it ran endoscopists through three successive phases, before computer-aided detection, during it, and after it was taken away. Thirteen endoscopists, 5,013 colonoscopies.
Among the seven inexperienced endoscopists, detection of at least one polyp of 5 mm or more rose while the AI was running, from 31.9% (216 of 678) to 39.5% (257 of 651), an odds ratio of 1.43 on a 95% confidence interval of 1.11 to 1.84. After the AI was removed it was 36.3% (173 of 476) against the 31.9% baseline — an adjusted odds ratio of 1.03, on 0.79 to 1.34, which is no change. Among the six experienced endoscopists there was no significant effect between any of the phases. The authors' conclusion, in their words: In non-CADe-assisted colonoscopy, no upskilling or deskilling was observed following a period of CADe exposure.
The two studies do not contradict each other, and the reason is the best thing in either paper. Pedersen and colleagues address the Polish result directly, and identify a difference in design that we had not seen and did not find in the coverage we read. In their words: We evaluated the effect of exposure to CADe ‘after’ CADe was removed, whereas Budzyń et al. investigated how endoscopists perform non-CADe-assisted colonoscopy ‘while’ they were exposed to CADe in the other colonoscopies they performed during the same time period.
The Polish endoscopists were still working with the AI most days and being measured on the days they were not. The Norwegian endoscopists had it taken away entirely — all but two of the thirteen, who reported running it throughout. The authors draw the obvious inference: the effect will be stronger in the Budzyń study than in our study, suggesting that the deskilling effect may be a temporal phenomenon that disappears over time after CADe is removed.
That is their reading, and it is worth saying what their design cannot separate: finding no deskilling after the tool is withdrawn is consistent with an effect that faded, and equally consistent with one that was never there. They also point at Troya and colleagues — the eye-tracking study in the section above — as the mechanism that would explain the Polish result.
What the Norwegian authors say their own study cannot do. Its endpoint is not adenoma detection but polyp detection at 5 mm and above, which they call the most important
limitation and less well-validated than ADR
. It ran from November 2021 to June 2025, and they warn that over three and a half years there is always a time effect in which endoscopists may naturally upskill or deskill over time
— their inexperienced endoscopists averaged 6.8 months of practice in the first phase and 23.8 months by the third, so the tool is not the only thing that changed. The case mix moved between phases. The multivariable analysis may be underpowered. It is not randomised. And on the question this page is about, their discussion is narrower than their stated conclusion, which covers all experience levels: the absence of a fall argues against the idea that exposure to CADe erodes skills, at least for nonexperienced endoscopists, although such skill reduction has been observed in another study involving only experienced endoscopists
— and the study they cite there is the Polish one.
The experienced group is where care is needed. It shows no significant effect in any phase — but that includes no benefit while the AI was running, and the authors suggest why: experienced endoscopists may also be less enthusiastic about a software tool intended to improve their performance
, and two endoscopists ran it throughout while the rest switched it on only during withdrawal. A group that did not visibly gain from the tool is weak evidence about what losing the tool does to them. At the level of individual endoscopists the split is a wash in both directions: among the inexperienced, three declined, two were unchanged and two improved; among the experienced, three declined and three improved.
One more thing. Yuichi Mori is an author of both papers — the senior author of the deskilling study is a co-author of the prospective study that found none. That is not a contradiction to be scored. It is what a field looks like when the people in it are still working the question out.
So the honest position is this. A retrospective comparison found a six-point fall in experienced endoscopists measured during a period of AI use. A larger prospective study found no fall in endoscopists measured after the AI was withdrawn, most clearly among the inexperienced, on a different and weaker endpoint. Those are compatible, and the second study's authors say so. Neither is the randomised crossover trial. The training study we were offered as the null result is not one. This one is — specifically on the question of what happens after the tool is withdrawn, since it found a real gain while the tool was running — with the limits its own authors put on it.
The mechanism was measured in 2022, by people who then said what worried them about it.
If AI exposure changes what an endoscopist can do unassisted, something has to have changed in how they look. Troya and colleagues at the University Hospital Würzburg measured it three years before the deskilling study was published, and are cited in that paper's own reference list.
Twenty-one participants — ten novices and eleven experienced staff, physicians and nurses — watched 29 recorded colonoscopy sequences while wearing eye-tracking glasses, pressing a button when they saw a polyp. Each sequence was viewed once with the AI overlay and once without, in a crossover design with a three-week washout between assessments, for 1,218 experiments in all.
With the AI running, eye travel distance fell from a median of 248.86 cm to 232.68 cm (P < 0.001). That is a reduction of about 6.5% across all participants. Among the experienced it was larger — 247.89 cm to 227.09 cm, about 8% — and among novices smaller: 255.66 cm to 247.89 cm, about 3%. The effect is real, statistically clear, and modest. The authors' reading of it is the sentence this section exists for:
Effort, as expressed by eye travel distance, decreased, and gaze remained more focused, presumably waiting for a bounding box to appear. This might be efficient but risks missing a polyp that is not captured by the system.
Two other findings in the same study cut against any simple story that the AI simply helps. The system detected polyps faster than people did — a median 1.16 seconds against 2.97 — but the people using it were no faster than the people working without it (2.90 seconds, indistinguishable from 2.97). And misreading got worse, not better: participants falsely identified a polyp in a median of 4 cases without the AI and 6 cases with it (P = 0.001), at both experience levels. They also inspected the AI's false-positive boxes for less than half the time those boxes were on screen — 43.8% overall, and least of all for the briefest ones.
What this is not. Twenty-one people watching short video clips is not twenty-one people performing colonoscopies, and the authors say so: their stated limitation is the experimental design using only short video sequences
. Their language about consequences is careful throughout — possible consequences of these findings might be prolonged examination time and deskilling
, and a potential risk of overreliance and deskilling
. This page uses their hedges, not firmer ones.
The same narrowing has been measured in breast imaging, and read there as a gain. Gommers and colleagues eye-tracked 12 radiologists across 150 screening mammograms: with AI support their accuracy improved, area under the curve rising from 0.93 to 0.97, while the share of the breast covered by their fixations fell from 11.1% to 9.5% and their dwell time inside lesion regions rose from 4.4 to 5.4 seconds. Their description is a more efficient search
. One team measures the narrowing with the tool present and calls it efficiency; another measures the outcome with the tool absent and worries about what was not looked at. Nobody has tested whether those are the same thing, and this page does not claim they are.
The same radiologists, the same task, and a forty-point swing.
Dratsch and colleagues showed 27 radiologists a set of mammograms alongside AI-generated BI-RADS category suggestions. The first ten cases carried correct suggestions, which established the system's credibility. Of the remaining forty, twelve were wrong.
When the AI's suggestion was correct, all three experience groups performed similarly, scoring roughly 80% — 79.7%, 81.3% and 82.3%. When the AI's suggestion was incorrect, the same readers scored 19.8% if inexperienced, 24.8% if moderately experienced, and 45.5% if very experienced, on standard deviations of 14.0, 11.6 and 9.1. The most experienced readers in the study were wrong more often than they were right, on average, whenever the machine was wrong — though 9.1 is a spread wide enough that the group reaches across the halfway line. The authors' conclusion is that radiologists at every level of experience are prone to automation bias when being supported by an AI-based system
.
This is not a study about skills decaying over months. It is a study about what happens inside a single reading, and it is the clearest measurement in this entire field of the thing people actually describe when they say AI has changed how they work: you take what it gives you.
A distinction we are borrowing, not making.
Ke and fifteen colleagues, writing in Nature Medicine in 2026 — online in May, in the June issue — separate three effects that ordinary discussion treats as one. We are using their frame and it is theirs.
And it has already been applied to this exact literature, by someone with standing to do it. On 5 June 2026, two days after the Norwegian trial appeared, Gastroenterology published an editorial by Weinberg and Bretthauer titled Upskilling, Deskilling, or Never Skilling: Who Benefits From Artificial Intelligence in Colonoscopy?
Michael Bretthauer is an author of the Polish deskilling study. So the specialty is working this frame over its own evidence, in its own journals, and has been since the week the null result was published. We found that editorial through a fact-check run on this page, after writing the section above. What this page does that it does not is carry the distinction out of the specialty to a reader deciding something about their own work — and that, rather than the frame, is the part we can claim.
Deskilling is an experienced practitioner losing a skill they had. That is what the colonoscopy study is about.
Never-skilling is a trainee who never develops the skill, because the tool was there from the beginning. In their words, trainees who rely on AI during the early formative years may fail to develop the foundational reasoning skills that safe, independent practice requires
. Troya and colleagues reached the same worry from the other direction in 2022, suggesting AI systems may have an impact on the learning curve of colonoscopy trainees by preventing them from developing the visual ‘bottom U’ pattern of high-performing endoscopists
.
Mis-skilling is internalising the machine's errors as fact. That is Dratsch's forty-point swing, and the pathologists who abandoned correct judgments, and the endoscopists who saw more polyps that were not there.
Their own summary is the closest thing this literature has to a settled conclusion: AI is not inherently harmful to learning; its educational impact depends on how and when it is introduced.
The evidence a general reader has been offered is much weaker than the evidence they have not.
The belief that AI is degrading thinking does not usually travel on any of the above. It travels on three studies, and all three are weaker than their reputations.
The brain-scan study. The MIT Media Lab preprint Your Brain on ChatGPT has been widely read as showing that AI users' brains are less active. Fifty-four participants completed its first three sessions; the fourth session, which carries the claim about lasting effects, had eighteen, split across two crossover arms. It has been a preprint since June 2025. A comment on it by four researchers at Vienna and Dresden calculates that roughly 159 participants would have been needed for adequate power, notes that some reported figures rest on two to four essays per group, and observes that the paper reports 55 completers while analysing 54, and that the criteria for dropping the one remain unclear. The paper does give a reason — it says it reports 54 to ensure data distribution (as participants were assigned in three groups)
— and the comment does not address that sentence. We read both. Their most consequential point concerns the brain measure itself: a significant connection
in that paper meant a coupling stronger in one group than another — a relative measure — so, as they put it, a group with fewer significant connections should not necessarily be interpreted as exhibiting weaker overall neural activity
. Stanković and colleagues close their comment with a judgement of that one paper and not of the field — they are describing the study they audited, and this page carries it as their verdict on it rather than as a finding about the field.
The Google-effect study. Sparrow, Liu and Wegner's 2011 finding that people offload memory to the internet is the origin of the modern version of this worry. Hesselmann's 2020 replication analysed 89 participants of 117 recruited and did not find the effect: the data were about five times more likely under the null than under the alternative, and about sixteen times more likely when tested against the original's specific predicted pattern. He places his result alongside an earlier large-scale replication project that also failed to reproduce the original finding. This page has not been able to read the 2011 paper itself, and so describes what the replications tested rather than what the original reported.
The critical-thinking correlation. Gerlich's 2025 study is the source of the figure most often quoted, a correlation of −0.68 between AI use and critical thinking. It surveyed 666 people recruited through social media in the UK. It is cross-sectional, and the author does not claim causation. The more specific problem is that AI use, cognitive offloading and critical thinking were all three measured by self-report in the same instrument, and the three correlations among them are uniformly strong — −0.68, +0.72 and −0.75. That uniformity is what common-method variance looks like. In fairness to the author, two further things: the recruitment was by convenience and purposive sampling promoted on social media, which he reports; and the paper carries a published correction of 10 September 2025, in which Table 4 had duplicated Table 3 and was replaced. The correction notice states that the scientific conclusions are unaffected, and nothing on this page turns on the corrected table.
The reassurance is in worse condition than the alarm. The most-cited meta-analysis finding that ChatGPT improves learning — cited several hundred times — its own journal's counter read 402 at one point on 30 August 2026 and 399 an hour later, which is what live metrics do, and other counts of the same paper run from 262 to over 500 — was retracted by its journal on 22 April 2026. Two researchers at the Arctic University of Norway had documented the problems publicly on 2 July 2025, nine and a half months earlier. Their audit found that the study carrying the greatest weight in the analysis, about 19% of it, had itself been retracted two months before the meta-analysis was published, and that the second heaviest, about 10%, measured self-efficacy and technology acceptance rather than learning. A table listing 41 unique entries supported an analysis reporting 44. One included study gave 36 students a single shared grade per group and was entered as though each had an individual score, producing an effect size of 4.009. Effect sizes above 1 are rare in education research; in that analysis they appeared in about 36% of included studies.
Separately, four learning scientists audited the studies behind the next-most-cited synthesis and found that of 19 comparisons, four met three basic criteria: a well-defined treatment, an adequately described control, and a valid learning measure. A learning measure was present in 10 of the 19. Their sentence about the rest is the one that matters here: If measured during treatment, differences likely reflect situational (dis)advantages due to the (non)availability of the treatment rather than durable learning.
What does exist outside medicine is narrow and good. Three randomised experiments — 1,222 people randomised, 1,060 analysed — had participants work through practice problems with or without AI, then removed the AI without warning and gave them three more. Unassisted solve rates fell, at effect sizes of −0.42, −0.19 and −0.42, and people gave up sooner. That is causal, and it held in all three experiments — which is internal replication by one team in one preprint, not confirmation by anybody else. It is also about fifteen minutes long, and it is a preprint with no pre-registration statement. Separately, three pre-registered experiments with 1,372 participants across 9,593 trials randomised whether an AI assistant's answer was right or wrong: participants consulted it on more than half of trials and, having consulted it, followed it on 79.8% of the trials where it was wrong. Access to it raised their confidence by 11.7 percentage points while about half its answers were faulty — and their confidence did not fall as they met more wrong answers. That is a working paper, not a reviewed one.
So the honest position for a reader outside medicine is this. The strongest evidence that AI use degrades a real skill, in real work, comes from an endoscopy suite — measured not after the tool was taken away but on the days it was not used, during a period when it was in use on the others. And it is a retrospective study with eleven limitations. The clearest evidence about what happens when the machine is wrong comes from a reading room. The studies being quoted in general coverage are not the strong ones, in either direction.
The caveats existed. Some travelled. The ones that would make a reader hold the number loosely did not.
We expected to find that this evidence had reached the public without its qualifications, and the first thing we looked for was whether that was true. It is not, at the source. The colonoscopy authors' eleven limitations are in their paper. The Science Media Centre published three named academics on the day of release, one of whom identified the volume confounder precisely. A critical appraisal for the American College of Gastroenterology laid out the design problems for gastroenterologists and proposed the gaze mechanism.
Some of it crossed, and the split is not random. Two things travelled: that the study was observational, and that something other than AI might account for the fall. Time carried both on 13 August 2025, quoting Professor Osmani on the rise in procedure volume and the fatigue it could cause, and Professor Tucker on automation bias not being peculiar to AI, and calling the study observational throughout. The ASCO Post reported the same day that the authors had acknowledged its observational nature
and that factors other than the implementation of AI use may have influenced the findings
, and quoted the senior author on what the result implies for the randomised trials that came before it. Healthgrades called it a retrospective observational study
and carried a named expert's judgement that the deskilling concern remains hypothetical rather than proven
.
What did not travel is the specific. Professor Osmani had named the rise itself at the Science Media Centre on the day of publication, and Time carried him saying so. What did not travel is why it rose: not one piece of coverage we read named the suspension of the Polish screening programme as the reason, or the absence of withdrawal-time data, or the eleven limitations as a body. The authors' call for randomised crossover trials did travel, but only inside the specialty: Professor Tucker made it at the Science Media Centre on the day of publication, and the American College of Gastroenterology's appraisal asked for prospective before-and-after trials. It did not reach a general reader. And the further the finding got from the specialty, the less arrived with it. Fortune carried the six-point fall to a business readership on 26 August 2025, connected it to an aviation analogy and to a Microsoft and Carnegie Mellon study, and reported none of the eleven limitations. Its headline put the same result a different way: doctors became 20% worse at spotting abnormalities on their own
. Six points is the absolute fall and about a fifth — 21% — is the relative one, and both are true. The headline was accurate. The sentence under it was not. Fortune told its readers that the endoscopists went from detecting potential polyps at a rate of 28.4% with the technology to 22.4% after they no longer had access to the AI tools they were introduced to
. Neither half of that describes the study. The 28.4% is the rate before the AI arrived, not with it; the 22.4% was measured while the AI was in routine use on other days, not after it was taken away. That is the same design error this page is about, in a magazine read by more people than the journal — and a draft of this section said nothing false had been printed here. That was wrong. We caught it by opening the article again rather than trusting our note of it. A research highlight in Nature Reviews Gastroenterology & Hepatology labelled the study observational and retrospective, as it should, and then reported the figures with no limitations and no quotes. A radiologist writing for trainees paired the colonoscopy result with the mammography experiment — the closest thing to this page in general-audience writing — and reported the limitations of neither. Communications of the ACM took the finding to knowledge workers, observing that similar issues pop up in law, education, journalism, software development, and other fields
; it carries caveats of its own, about governance and about deskilling having good outcomes as well as bad ones, but none of this study's caveats.
And the sharpest instance is not general coverage at all. The 2026 scoping review that assembles this literature for doctors describes the colonoscopy study in its own abstract as a multicenter randomized trial
. It is a retrospective observational study; the review's body says so. The same sentence carries a second error, and it is the one this page is about: it says the rate dropped when endoscopists reverted to non-AI procedures after repeated AI use
, which implies the tool had been withdrawn. It had not. They were still using it on other days of the same period. Its abstract also reports that over 30% of participants reversed correct initial diagnoses
in a pathology experiment under time pressure, where its body puts the same study at 7% of judgments. Both figures can be true of one dataset — they have different denominators — but the larger one is in the abstract, and the abstract is the part that gets read. Its radiology section then says that a controlled reader study published in 2023 evaluated 27 radiologists interpreting 720 mammograms with and without AI support
and that error rates rose when the AI was wrong — and attaches that sentence to its reference 16, which is a 2025 paper by different authors, 12 radiologists reading examinations from 150 women, that found AI improved accuracy. The 2023 study the description otherwise fits is not in the review's reference list at all. We read the review's full text at source on 29 August 2026, and noticed only because we had read both underlying papers that evening.
None of this makes the review's substance wrong, and we are relying on it. It makes the point that carrying a finding accurately across a boundary is difficult enough that the careful people fail at it too.
This assessment is living. These are the questions, and the date above is when somebody last looked.
The randomised crossover trial. The colonoscopy authors asked for it, and Professor Tucker asked for it independently on the day of publication. Randomise which endoscopists work with AI, cross them over, and measure unassisted detection on both sides. It settles the confounding that the volume doubling introduced, and it is the single study that would move this from an association to a causal finding.
A longer horizon. The Polish comparison is three months either side. The prospective trial runs three phases and is the longest sequence on this page, and it found no decay. Neither reaches the horizon people actually ask about, which is years.
The mechanism in live procedures. The eye-tracking evidence is 21 people watching video clips. Eye tracking during real colonoscopy, linked to what was actually found, would move the mechanism from plausible to shown.
Peer review. Three of the studies this page relies on are preprints or working papers, and we have said so each time. If they are reviewed, that changes what they carry.
Corrections and retractions. One meta-analysis in this literature has already been retracted after nine and a half months of public criticism, and a scoping review in it mis-cites a paper. We check the sources on this page against their journals' notices, and record when we last did.
Whether anyone else does this carefully. If a general-audience account carries this evidence with its limitations attached, we will say so and link to it.
The full watch list, with the searches that would answer each question, is kept in the repository that produces this page, alongside a record of every check — including the ones that found nothing.
The limits of this page.
This is a statement about evidence and not advice. It does not say whether any clinician, department or reader should use these tools, and nothing here bears on whether AI-assisted colonoscopy is good for patients — the randomised evidence that AI raises detection during the procedure is not in dispute and is not what this page examined.
Every figure above comes from the study that produced it, read directly. Where we could not read something, we say so rather than passing on somebody else's account of it. Three things on this page fall into that category: a commentary published alongside the colonoscopy study, which we could not open; the 2011 memory study, which is paywalled; and one meta-analysis whose scope we describe as its auditors reported it rather than as we read it. And one thing was not unread but unfound: the prospective trial above had been in print for nearly three months when we drafted this page, and our own standing search for exactly that study had not been run. An outside reviewer found it. We had recorded that the search was outstanding and wrote the piece anyway, which is the failure — not the missing paper, but proceeding past our own note that we had not looked.
This page carries a “last reviewed” date rather than a “last updated” one, and the difference is deliberate. It is the date somebody last looked, not the date something last changed. A review that finds nothing is still a review and is recorded as one. If that date is old, read this page accordingly.
Every figure above, and the document it came from. Where we read a study through somebody else’s account of it rather than at source, this says so.
had been mistakenly omitted from the multivariable analysis presented in supplementary table 6 in the appendix of this Article, that the appendix has been updated, that
the only affected findings are for the variables ‘After AI introduction (using AI in colonoscopy)’ and ‘Age >60 years’, and that
these changes do not affect the interpretation of the data. Two things a reader should take from that. The figures on this page — 28.4%, 22.4%, the six-point difference and its interval — are the unadjusted comparison and are not touched by it. But one of the two findings that was affected is the exposure variable itself, AI introduction, in the adjusted model. The journal says interpretation is unchanged and we have no basis to disagree; we note it because a correction to the adjusted estimate for the thing a study is about is worth a reader knowing, whichever way it went. We found the correction on 30 August 2026 through a page-gate search, could not open it through seven routes, and read it only when the operator obtained the notice. The December 2024 preprint of the same work is on SSRN and its conclusion is worded more strongly than the published one.
a narrative review. We use the title’s term. The proof we hold prints “Volume 11 ■ Issue C” in its running footer, above “Available online xxx”; the publisher’s registered metadata for the version of record gives volume 12 and no issue, which is what we cite. This page relies on the review for that count and for nothing else; each paper named above was read at source. The review itself was read in full, in its PDF, on 29 and 30 August 2026.