For about a decade the most exciting idea in GLP-1 research was not weight. It was the brain. Animal work showed these drugs calming inflammation and protecting neurons; database studies kept finding less dementia in people prescribed them; and by 2023 the question had moved from whether GLP-1 drugs might treat Alzheimer's disease to how soon we would know.
We now know, for the population where it was actually tested. The hypothesis was taken to phase 3 in people who already have early Alzheimer's disease, by the company with the strongest commercial reason to want it to work, and it did not survive there. This page sets out what the trials found, why the database studies still say the opposite and contradict each other in a way that gives the game away, and one under-reported harm that is genuinely about the brain and has nothing to do with neuroprotection.
This page reports published results. It does not tell anyone to start, stop or change a medicine; that belongs with a clinician, as our editorial standards explain.
The short version
| Question | What the randomised evidence says |
|---|---|
| Do GLP-1 drugs slow Alzheimer's disease? | No. Two phase 3 trials, 3,808 people, difference near zero. Both discontinued. |
| Do they slow Parkinson's disease? | No. One phase 3 trial, 194 people, 96 weeks, no difference. |
| Do they prevent dementia in healthy people? | Never randomised. Nobody knows. |
| Do they improve cognition in people without dementia? | Not established. The pooled analysis most often cited for this was retracted in April 2026; the large trials show nothing. |
| Do they change appetite and intrusive food thoughts? | Yes — that part is consistently reported. |
| Can they harm the brain? | Rarely, and indirectly: rapid weight loss plus vomiting can deplete thiamine. |
evoke and evoke+: the test that settled it
Novo Nordisk ran two phase 3 trials of oral semaglutide 14 mg in early Alzheimer's disease. Between 2021-05-18 and 2023-09-08 they screened 9,981 people and randomised 3,808 across 566 sites in 40 countries: 1,855 in evoke (928 semaglutide, 927 placebo) and 1,953 in evoke+ (976 and 977). Participants were 55 to 85, with mild cognitive impairment or mild dementia, and — this matters — every one had the disease confirmed by amyloid testing rather than by symptoms alone. Mean age 72.2; mean baseline CDR-SB 3.7.
The primary endpoint was change at week 104 in CDR-SB, the Clinical Dementia Rating Sum of Boxes, on which a clinician scores memory, orientation, judgement, community affairs, home life and personal care, and higher means worse.
| evoke | evoke+ | |
|---|---|---|
| Randomised | 1,855 | 1,953 |
| CDR-SB change to week 104, semaglutide | 2.3 (SE 0.1) | 2.2 (SE 0.1) |
| CDR-SB change to week 104, placebo | 2.3 (SE 0.1) | 2.1 (SE 0.1) |
| Estimated difference (95% CI) | -0.08 (-0.35 to 0.20) | 0.10 (-0.17 to 0.38) |
| p value | 0.57 | 0.46 |
Source: Cummings et al., The Lancet, published online 2026-03-19, trials NCT04777396 and NCT04777409, funded by Novo Nordisk.
Read the two estimates together. In one trial semaglutide was a fraction better than placebo; in the other a fraction worse, both intervals sitting almost symmetrically around zero. This is not a treatment that missed narrowly, nor an underpowered trial that could not see a real effect: with 3,808 people the intervals rule out anything close to the hoped-for benefit. Novo Nordisk announced the topline on 2025-11-24, discontinued the planned extension, and the paper followed in March 2026.
Two details, because they are the commonest ways a null result gets talked around. The endpoint was clinical, on purpose. Biomarker movement is the usual salvage offered for a null trial, and we are quoting none here: the indexed abstract we worked from reports the clinical outcome and the safety data, not biomarkers, and we do not quote figures out of a full text we could not open. The reason CDR-SB rather than a biomarker was the primary endpoint is the same reason a biomarker readout would not rescue this result: biomarker movement without clinical benefit is a pattern this field has seen many times. Tolerability was not the problem. Treatment-emergent adverse events were reported in 91.2% of the semaglutide group against 84.8% on placebo, the usual gastrointestinal profile, and of the five deaths investigators judged treatment-related, four were on placebo. The trial was not sunk by dropouts. It simply did not work.
The other randomised trials
ELAD (liraglutide, Alzheimer's disease). A phase 2b trial in 204 people with mild to moderate Alzheimer's disease and no diabetes, 52 weeks. The primary outcome, cerebral glucose metabolic rate, showed no difference (-0.17; 95% CI -0.39 to 0.06; P=0.14). Nor did activities of daily living (-0.58; -3.13 to 1.97) or CDR-SoB (-0.06; -0.57 to 0.44). One secondary measure, the ADAS executive-function domain, favoured liraglutide (0.15; 0.03 to 0.28; unadjusted P=0.01) — an unadjusted secondary in a trial that missed its primary, which is the definition of a finding needing replication rather than a headline. (Edison et al., Nature Medicine, online 2025-12-01.)
Exenatide-PD3 (Parkinson's disease). The strongest earlier signal came from Parkinson's, where a small phase 2 had been positive. The replication randomised 194 people at six UK hospitals to weekly exenatide or placebo for 96 weeks. Off-medication MDS-UPDRS part III scores worsened by 5.7 points on exenatide and 4.5 on placebo — adjusted coefficient 0.92 (95% CI -1.56 to 3.39), p=0.47; serious adverse events 9 versus 11. (Vijiaratnam et al., The Lancet, 2025-02-04, funded by NIHR and Cure Parkinson's.)
Three diseases, three drugs, one pattern: promising phase 2 or observational signal, null phase 3.
The pooled analysis that pointed the other way has been retracted
The usual counterweight to the null trials is a 2025 meta-analysis of randomised trials in people with type 2 diabetes, which reported small improvements in Mini-Mental State Examination and Montreal Cognitive Assessment scores on GLP-1 receptor agonists. It is still being quoted. We do not cite it, because it has been withdrawn: PubMed carries Wan et al., Diabetes, Obesity and Metabolism (PMID 41104525) as a Retracted Publication, with the retraction notice published in the same journal in April 2026 (Diabetes Obes Metab 2026;28(4):3451, doi 10.1111/dom.70552). We are not restating its effect estimates here, in either direction. Withdrawn evidence is not weak evidence; it is absent evidence, and repeating the numbers while labelling them retracted still puts them into circulation.
That leaves no pooled randomised estimate of GLP-1 effects on cognition in people with diabetes. Note which way that cuts: had the analysis stood, it would have been the strongest published argument against the reading on this page, and it is gone.
What does not depend on it is the structural point. Cognitive impairment in poorly controlled diabetes has causes — hypoglycaemia, vascular disease, glycaemic control — that any effective diabetes drug might improve without touching Alzheimer's biology, so a positive result in that population would not have been a neuroprotection result. The screening instruments used in those trials are ceilinged: MMSE is a 30-point screen on which an intact adult scores 28 to 30, and a point or two of movement on it in a mostly non-demented group is not the same finding as slowing a neurodegenerative disease. Small trials of soft endpoints in mixed populations produce small positive numbers; the enormous trial with a hard endpoint in an amyloid-confirmed population produced zero. When those disagree, the second wins.
The observational data: large effects that contradict each other
| Study | Design and data | Comparison | Result |
|---|---|---|---|
| Wang et al., Alzheimer's & Dementia, 2024-10-24 | Target trial emulation, 1,094,761 people with T2D, US EHR network; outcome was first-time Alzheimer's diagnosis, not dementia generally | Semaglutide vs insulin | HR 0.33 (0.21-0.51) |
| Same study | Semaglutide vs other GLP-1 agonists | HR 0.59 (0.37-0.95) | |
| Zhang et al., Alzheimer's & Dementia, 2025-09 | Two real-world databases, Cox models | GLP-1 agonists vs DPP-4 inhibitors | HR ≤0.69 |
| Same study | SGLT2 inhibitors vs DPP-4 inhibitors | HR ≤0.67 | |
| Zhou, Tang et al., Diabetes, Obesity and Metabolism, online 2025-12-22 | Target trial emulation, Penn Medicine EHR | GLP-1 agonists vs DPP-4 inhibitors, 6,677 pairs | HR 0.76 (0.59-0.97) |
| Same study | GLP-1 agonists vs SGLT2 inhibitors, 8,434 pairs | HR 1.53 (1.13-2.07) — worse | |
| Inoue et al., Annals of Internal Medicine, online 2025-07-22 | Target trial emulation, Medicare, 2,418 vs 4,836 matched | GLP-1 agonists vs DPP-4 inhibitors | Risk ratio 0.83 (0.61-1.05) |
| da Silva et al., J Diabetes Complications, online 2026-02-25 | TriNetX, 44,470 matched pairs | Tirzepatide vs semaglutide, mild cognitive impairment | RR 0.12 (0.06-0.22) |
| Lin et al., JAMA Network Open, 2025-07 | TriNetX, 30,430 matched pairs, 7-year follow-up | Semaglutide or tirzepatide vs other antidiabetics, dementia | HR 0.63 (0.50-0.81) |
| Same study | Same comparison, all-cause mortality | HR 0.70 (0.63-0.78) |
Watch what happens when the comparison drug changes. In the Zhou analysis the same GLP-1 users are protected against DPP-4 inhibitors (0.76) and harmed against SGLT2 inhibitors (1.53) — one dataset, one method, one paper. A drug does not have opposite effects on the brain depending on what somebody else was prescribed. What changes between those rows is who the comparison patients are.
The tirzepatide-versus-semaglutide row is the clearest tell. A risk ratio of 0.12 is an 88% reduction achieved by swapping one weekly injection for another in the same class — larger than any effect any dementia treatment has shown in a randomised trial. The authors say as much themselves: descriptive, hypothesis-generating, very low absolute risks.
The last two rows are the second tell, and the more instructive one. In the same 60,860-patient cohort that reports a 37% reduction in dementia, GLP-1 users also show a 30% reduction in all-cause mortality — dying of anything at all, over seven years. No trial of these drugs has produced a mortality effect near that. Whatever inflates the mortality number — healthier patients, better-attended patients, patients who stay on an expensive injectable — is also inflating the dementia number, because it is the same people in the same comparison. Read the two together and the parsimonious explanation is not neuroprotection. It is who gets prescribed what.
Four mechanisms produce numbers like these with no drug effect at all. Confounding by indication: insulin goes to people whose diabetes has been severe longest, so comparing a GLP-1 to insulin largely compares healthier patients to sicker ones — which is why the biggest effect in the table, 0.33, is the insulin row. Detection bias: dementia is diagnosed when somebody brings it to a doctor, and contact with the health system differs by drug group. The healthy-adherer effect: people who stay on an expensive injectable differ from those who do not, in ways that predict dementia and are invisible in a claims database. Short follow-up against a slow disease: median follow-up here is two to three years, while Alzheimer's pathology develops over decades, so a diagnosis recorded 18 months after a prescription began before the prescription did.
The most carefully constructed of these studies has the least dramatic answer. Inoue and colleagues reported a risk ratio of 0.83, interval 0.61 to 1.05 — compatible with anything from a 39% reduction to a 5% increase. It is not a flat null everywhere in the paper, and the part that cuts against the reading here should be on the page rather than left out of it: the same analysis reported a risk ratio of 0.64 (0.46 to 0.93), a statistically significant reduction, in patients younger than 75, alongside 1.22 (0.74 to 1.66) in those 75 or older, and the authors put that age difference in their own conclusion. Two subgroups from one observational study, split at one cut point, is the kind of finding that needs replication rather than promotion — but it is a real result, and the authors' overall conclusion is the same either way: no clear evidence of a difference, and randomised trials are needed. We now have those trials, and they are null.
The real brain signal: thiamine
While the neuroprotection story was failing, a different neurological story was accumulating in the safety databases with a fraction of the attention.
Wernicke encephalopathy is what happens when the brain runs out of thiamine (vitamin B1). Classically it presents as confusion, unsteady gait and abnormal eye movements, though the full triad is often absent. Treated late, it can leave permanent memory damage — Korsakoff syndrome, in which a person cannot form new memories and fills the gaps without knowing they are gaps. Thiamine stores last weeks, and very rapid weight loss with persistent vomiting or near-absent eating is a recognised way to exhaust them, which is why the condition was already known after bariatric surgery and in hyperemesis of pregnancy. GLP-1 drugs produce the same combination in some people.
What is in the literature as of 2026-09-05:
- A French pharmacovigilance analysis described one case and identified 18 others in the literature and in VigiBase, the WHO global safety database. Nausea, vomiting or reduced intake featured in 68% of cases, with weight loss of 3.5 to 13.3 kg per month over three to six months, and Wernicke encephalopathy was disproportionately reported for semaglutide, for tirzepatide and for the class (Gras et al., European Journal of Clinical Nutrition, online 2025-09-04). That paper's abstract states the disproportionality without putting a number on it.
- The numbers usually quoted alongside it come from a different piece of work, and are often misattributed. A separate VigiBase review of 18 reports plus three published cases, presented by Manar Kadhem of Mansoura University at AACE on 2026-04-23 and written up by the Cleveland Clinic Journal of Medicine, reported reporting odds ratios of 10.2 (95% CI 5.8 to 17.9) for semaglutide and 11.4 (2.8 to 45.8) for tirzepatide. Those are the CCJM write-up's figures, read on 2026-09-05; the underlying conference abstract is not indexed in PubMed.
- An Israeli group asked the same question of the FDA's adverse-event system and found 15 cases (8 semaglutide, 6 tirzepatide), 14 reported in 2023-2024, with gastrointestinal features in 13 of 15 and long-term neurological sequelae in 7 of the 11 patients with follow-up. Their disproportionality figure is far smaller: a reporting odds ratio of 2.35 (1.38 to 4.01) (Lev et al., Clinical Nutrition, online 2026-01-02).
- A PRISMA review of published case reports found six semaglutide cases in obesity, all with prolonged gastrointestinal symptoms and substantial weight loss before neurological deterioration, and outcomes including progression to Korsakoff syndrome and death (Bidesie and Oudman, Obesity, online 2026-07-03).
Three cautions. A reporting odds ratio is not a risk: it measures how often something appears in a spontaneous-report database relative to other drugs, and it inflates with publicity, litigation and sheer prescription volume — which is why two analyses of the same phenomenon return 10.2 and 2.35. The absolute count is small: a few dozen reports worldwide against tens of millions of prescriptions. And case reports get published precisely because they are unusual.
What survives those cautions is a plausible mechanism, a specific risk profile, and a treatment that works when given early. In the AACE review's cases, hospitals gave high-dose intravenous thiamine — the Cleveland Clinic Journal of Medicine write-up puts the range at 500 to 1,500 mg per day — and the physical signs reversed quickly, while approximately 20% of those patients were discharged with cognitive or memory deficits. The FAERS series is a different and smaller cohort with a worse-looking figure for the same thing: 7 of the 11 patients who had any follow-up recorded had long-term neurological sequelae. The two are not comparable — one counts everyone at discharge, the other counts only the minority of spontaneous reports that carried follow-up at all, which selects for the cases that went badly. Neither is an incidence rate. That is an intravenous hospital treatment, not a supplement decision, and nothing here is a reason to take anything on your own. The usable point is recognition: confusion, unsteady walking or double vision in someone who has been vomiting or barely eating on one of these drugs is an emergency-department problem.
Food noise is real, and it is not cognition
The one brain-adjacent change people report constantly is that the running commentary about food goes quiet. It has a name — "food noise" — and now a measurement. The INFORM survey asked 550 US adults on injectable semaglutide to complete a five-item Food Noise Questionnaire scored out of 20. The median score recalled from before treatment was 13 (IQR 10 to 16); after, 6 (IQR 3 to 10). Agreement with statements indicating intrusive food thoughts fell from 47-63% to 15-20%, consistently across treatment duration and BMI, and 83% reported satisfaction (Arnaut et al., Advances in Therapy, online 2026-05-30).
The limitations are on the label: respondents recalled their pre-treatment state after months of treatment, the weakest form of before-and-after there is; 86% were women and 79% white; and four of the seven authors are Novo Nordisk employees at a survey Novo Nordisk funded. The authors themselves ask for prospective work.
Even at face value this is a change in appetite and preoccupation, not in intelligence, memory or processing speed. People do describe feeling mentally clearer once the food chatter stops, and that experience deserves to be taken seriously as an experience. It is not evidence that a cognitive ability improved, and no trial has tested whether one did.
Can you measure this on yourself? No, and here is the arithmetic
Blunt version: a consumer IQ test — or any home cognitive test — cannot tell you whether a drug improved your thinking. Not a good one, not a bad one, not one taken carefully at the same hour in a quiet room. The reason is not that such tests are frauds. It is that the noise is bigger than the signal, by a lot. Three published quantities, then the subtraction.
One: taking a test twice raises your score, by itself. In 54 healthy young adults — mean age 21, mean 15 years of education — given the WAIS-IV twice, Full Scale IQ improved by about 7 points at both 3- and 6-month intervals, and the authors state plainly that these gains reflect practice rather than genuine intellectual change (Estevis, Basso and Combs, The Clinical Neuropsychologist, 2012). That is a student sample, and a practice effect measured in 20-year-olds is not automatically the practice effect in a 55-year-old on a GLP-1, so treat 7 points as the top of the range rather than the expected value. In employment testing, a meta-analysis across 107 samples and 134,436 people found an adjusted overall retest effect of d = 0.26 — about 4 IQ points — larger when identical forms were used and when practice came with coaching (Hausknecht et al., Journal of Applied Psychology, 2007). A second meta-analysis, of nearly 1,600 individual effect sizes across neuropsychological tests, found the same direction and reported that alternate forms, participant age, clinical diagnosis and the retest interval all moderated the size of the gain (Calamia, Markon and Tranel, 2012); its per-test effect sizes are in the paper's tables rather than its abstract, and we do not quote numbers from tables we have not read.
Two: a professionally administered IQ score carries a measurement error of about 2 to 3 points. WAIS-IV Full Scale IQ has an internal reliability around 0.98 and a test-retest reliability around 0.96. The standard error of measurement is 15 × √(1 − reliability): 2.1 points at r = 0.98, 3.0 at r = 0.96. Comparing two testings compounds it — the standard error of the difference is that figure times √2, so 3.0 to 4.2 points. Multiply by 1.96 for a two-sided 95% band and you get 5.9 to 8.3, so a reliable-change band of roughly ±6 to ±8 points. Anything smaller is not a change; it is the instrument breathing. (Do the multiplication yourself — that is the point of showing the inputs. Confidence intervals quoted around a single IQ score are narrower than this, because they compound nothing; the wider figure is the one that applies to a before-and-after comparison.)
Three: the strongest cognitive enhancers ever measured are smaller than that. Meta-analyses of 47 studies in healthy, non-sleep-deprived adults found an overall effect of modafinil of 0.12 standard deviations and of methylphenidate 0.21, with no effect at all for d-amphetamine (Roberts et al., European Neuropsychopharmacology, 2020). On the IQ scale, 0.12 to 0.21 SD is about 2 to 3 points.
Now the subtraction. Practice effect from sitting the same test again: about 4 points in the large employment meta-analysis, up to about 7 in the WAIS-IV student study. Error band you must clear before a change means anything: 6 to 8 points. Largest pharmacological cognitive effect anyone has demonstrated: 2 to 3 points. The noise is not merely comparable to the signal — it is several times larger, and the practice component points upward, so a home retest is biased toward telling you the drug worked.
One more figure, labelled as our calculation rather than a citation because we could not find it published as a single number: take the standard error of difference at 4.2 points and the practice gain at 4, treat change as normally distributed, and roughly a third of people — 33% — will move more than 6 points between two sittings with nothing done to them at all. The inputs are above so the arithmetic can be checked or corrected; a published source for that proportion is welcome at our contact page.
None of this gets better on a test taken at home. An unproctored online test has no examiner, no controlled conditions and, as a class, no published reliability coefficient or norming sample of its own — and an error band that has never been quantified cannot be shown to be narrower than the WAIS's, which is already too wide to resolve a 2-to-3-point effect. That is a statement about the whole category, not about any one product.
IQ Revealed publishes a methodology page for its 40-item online test. Read on 2026-09-05, it states that a raw score is converted to the conventional mean-100, SD-15 scale; that "every score is an estimate with a margin of error" moved by "sleep, stress, distraction and familiarity with the format"; that an unproctored 40-item test of this kind has "no published reliability coefficient"; and that it is "not a clinical evaluation", formal diagnoses requiring a psychologist-administered instrument such as the WAIS.
If there is a real question about someone's thinking, the answer comes from a clinician and a proper neuropsychological assessment, where practice effects are anticipated, alternate forms exist, and change is scored against published reliable-change indices.
What people buy instead
A predictable consequence of "GLP-1s might protect the brain" was a market, and two things about it are worth knowing.
The drugs themselves are widely sold online without prescription, and the product is frequently not the product. A market-surveillance study bought semaglutide from illegal online pharmacies and analysed it: measured purity ran from 7.7% to 14.37% against the 99% claimed on the labels, semaglutide content exceeded the labelled amount by 28.56% to 38.69%, endotoxin was found in every sample, and three of six purchases were simply non-delivery scams. The 30 top domains involved had drawn over 4.7 million visits in a single quarter (Ashraf et al., Journal of Medical Internet Research, 2024-11-07).
Alongside them sits a set of "brain peptides" — Semax, Selank, Dihexa, Cerebrolysin, DSIP — sold as research chemicals and marketed with exactly the neuroprotection vocabulary the GLP-1 trials just spent. We do not tell anyone to use them, and we will say this much: none has anything resembling the evidence oral semaglutide had going into evoke, and evoke still failed.
Medibact publishes a library of written guides to individual compounds, among them Selank and Dihexa; the Dihexa entry is titled "Dihexa (PNB-0408): Critical Evidence Review". Read on 2026-09-05, the library states that it is for "educational use only — not medical advice… not a recommendation to use any compound", and that any doses, schedules or combinations shown "are examples of what has been reported, not instructions for you".
What this page does not decide
It does not say whether anyone should take a GLP-1 medicine. These trials tested one hypothesis — that the drugs slow neurodegenerative disease — and answered no. They say nothing about weight or diabetes, where the evidence is strong and separate; our cost page and our comparison of retatrutide and tirzepatide cover that ground. It does not say these drugs are bad for the brain: Wernicke encephalopathy is rare, understandable and treatable when caught. And it does not close the question. What is ruled out is a large disease-modifying effect in people who already have early Alzheimer's disease, and one in established Parkinson's. Prevention in healthy people has never been randomised and probably never will be, because the trial would need decades. If that changes, this page changes and says so.
Sources and dates
All sources opened or re-checked on 2026-09-05, including a check of each PubMed record's publication type for retraction notices. Trials: Cummings et al., The Lancet (evoke and evoke+), online 2026-03-19; Vijiaratnam et al., The Lancet (Exenatide-PD3), 2025-02-04; Edison et al., Nature Medicine (ELAD), online 2025-12-01; Novo Nordisk's topline announcement of 2025-11-24. Observational: Wang et al. (2024-10-24) and Zhang et al. (September 2025), Alzheimer's & Dementia; Zhou, Tang et al., Diabetes, Obesity and Metabolism, online 2025-12-22; Inoue et al., Annals of Internal Medicine, online 2025-07-22; da Silva et al., Journal of Diabetes and its Complications, online 2026-02-25; Lin et al., JAMA Network Open, July 2025. Thiamine: Gras et al., European Journal of Clinical Nutrition, online 2025-09-04; Lev et al., Clinical Nutrition, online 2026-01-02; Bidesie and Oudman, Obesity, online 2026-07-03; the Kadhem AACE 2026 review as written up in the Cleveland Clinic Journal of Medicine. Food noise: Arnaut et al., Advances in Therapy, online 2026-05-30. Measurement: Estevis, Basso and Combs (2012) and Calamia, Markon and Tranel (2012), The Clinical Neuropsychologist; Hausknecht et al., Journal of Applied Psychology (2007); Roberts et al., European Neuropsychopharmacology (2020). Online product quality: Ashraf et al., Journal of Medical Internet Research, 2024-11-07. Pages read directly: iqrevealed.com's methodology page, medibact.com's guide library, the Cleveland Clinic Journal of Medicine's AACE 2026 write-up.
Retracted and therefore not used: Wan et al., Diabetes, Obesity and Metabolism, online 2025-10-17 (PMID 41104525), retracted April 2026. It supported a section of this page in draft; that section now records the retraction instead of the finding.
Sources we could not open and therefore do not quote directly: the full texts at thelancet.com and clinicalnutritionjournal.com, which returned 403 to our requests — the trial and pharmacovigilance figures above come from the indexed abstracts of the same papers, which we link. Figures not verified against a primary table are flagged where they appear.
Corrections go to the contact page. The about page says who writes this and what we do not do.