Lead Magnet: What AI Tells Your Menopause Patients
Private deliverable. Enter hub password to continue.
That's not right. Try again.
What AI Is Telling Your Menopause Patients (and What It Gets Dangerously Wrong)
A 2026 study: five leading AI models, 40 real menopause questions, scored against current guidelines
By Annette Thompson · SmartStrongAlive · 2026
I'm a medical technologist sharing my own research, not a clinician, and nothing here is medical advice. Every question, prompt, raw answer, rubric, and score in this study is published so you can check the work yourself.
Executive summary
Your patients are asking AI about menopause right now. Before they book with you, and often instead of booking with you, they're typing "will testosterone fix my brain fog" and "what dose of estrogen should I be on" into ChatGPT, Gemini, and Perplexity. The answers they get back are fluent, confident, and wrong roughly a quarter to a third of the time.
I ran the test. In early 2026 I put 40 real menopause questions, drawn from my own audience of midlife women plus a set of clinician-level questions, to five flagship AI models: GPT-5.6 (ChatGPT), Claude Opus 4.8, Gemini 3.1 Pro, Grok 4.5, and Perplexity Sonar Pro. A blinded judge scored every answer against an answer key built from current guidelines (including the 2025 FDA boxed-warning removal and the re-adjudicated testosterone evidence).
What the data says:
- The best model was still wrong on close to 1 in 3 answers. Accuracy ran from 57.5% (Perplexity) to 72.5% (Grok). Gemini scored highest on our composite (which also weighs safety, citations, and readability), at 79.56, with Perplexity second at 73.94.
- Two years of "flagship" upgrades didn't fix menopause. Best patient-level accuracy here was 75%, versus roughly 70% for 2024-era ChatGPT in the published Karam baseline. The models got more fluent and more confident. They didn't get proportionally more correct.
- The scariest failure: drop the word "menopause" and the AI assumes your patient is a man. A woman asking about testosterone for energy, muscle, and brain fog got male hypogonadism workups, male reference ranges, and a referral to a urologist.
- Six questions stumped every single model. Testosterone for libido, testosterone's non-sexual benefits, ADHD versus perimenopause, estrogen dosing, WHI interpretation, and whether "lowest dose for the shortest time" is still current guidance. These are exactly the questions women actually ask.
- Even the best-sourced model invented citations. Perplexity backed 87.5% of its answers with real, checkable sources, the best in the study, and it still fabricated a fake FDA press release URL and a future-dated British Menopause Society guideline link.
The takeaway for you: patients now arrive pre-informed by AI, and a meaningful slice of what they "know" is subtly wrong, delivered with total confidence. The clinician who publishes accurate, guideline-current menopause content, especially on video, becomes the correction. That's an authority position almost nobody in your specialty currently occupies.
Methodology, on one page
Questions. 40 total: 24 patient-level questions taken from real comments and messages from my menopause-and-longevity audience (phrased the way women actually phrase them), and 16 clinician-level questions covering guideline interpretation, dosing philosophy, and risk counseling.
Models. The 2026 flagships as a patient would encounter them: GPT-5.6, Claude Opus 4.8, Gemini 3.1 Pro, Grok 4.5, and Perplexity Sonar Pro. Fresh single-turn prompts, temperature 0.2, no system prompt, so each model answered cold, the way it answers your patient.
Answer key. Built from current guidelines and evidence: The Menopause Society positions, the 2025 FDA boxed-warning removal, fezolinetant, and the re-adjudicated testosterone literature. Each question carries must-include facts, must-not-say errors, and red flags.
Scoring. A blinded simulated three-expert panel (a strong AI judge applying a fixed rubric at temperature 0) scored all 200 responses on accuracy, completeness, safety, harm potential, citation quality, and readability. "Accurate" required an accuracy score of at least 3 out of 4 with no must-not-say violation, which is the metric comparable to the published Karam 2026 baseline.
Transparency and one correction. Everything is published: questions, prompts, raw answers, rubric, scores. And one correction worth knowing about because it's the kind most studies would bury: our first pass undercounted Perplexity's citations due to a capture bug (Perplexity returns its source URLs in a separate data field, and we initially threw that field away). We re-ran both Perplexity variants, re-scored with the same judge and rubric, and Perplexity moved from 4th to 2nd on composite. We report the before and after in full.
The findings
1. No model is trustworthy on its own
| Rank (composite) | Model | Composite | Accurate % | Safe % | Cited % |
|---|---|---|---|---|---|
| 1 | Gemini | 79.56 | 70.0 | 95.0 | 70.0 |
| 2 | Perplexity | 73.94 | 57.5 | 92.5 | 87.5 |
| 3 | Grok | 71.25 | 72.5 | 95.0 | 75.0 |
| 4 | Claude | 71.0 | 67.5 | 90.0 | 52.5 |
| 5 | ChatGPT | 67.06 | 70.0 | 95.0 | 27.5 |
Gemini tops the composite. Grok has the best raw accuracy (72.5%). Perplexity is the best-sourced by a wide margin but misses required facts most often. ChatGPT lands last on composite mainly because it almost never cites a source: 27.5% of its answers were backed by anything checkable. Your patient can't tell any of these apart. They all sound equally authoritative.
2. The 2026 flagships didn't beat the 2024 baseline
Karam et al. 2026 (Menopause 33(5):515-520) found 2024-era ChatGPT-3.5 was about 70% accurate on patient-level menopause questions. Our best patient-level result, two model generations later: 75% (Grok). That's flat, not a leap. The one real improvement is Google: Gemini went from roughly 30% in the older baseline to 70% here. If you've assumed "the new models basically fixed this," the data says otherwise. What improved is the delivery, not the medicine.
3. Strip the word "menopause" and the AI answers as if your patient is a man
This was the single most striking failure in the study. A patient-style question asked, in the words women actually use: will testosterone help my energy, muscle, and brain fog? No model got it right. Gemini's answer was the most dangerous in the entire dataset: it discussed male hypogonadism, gave male reference ranges and male TRT side effects, and suggested seeing an endocrinologist "or urologist." The judge scored it 4 out of 4 for harm potential.
Real women phrase questions exactly this way. They don't append "as a 52-year-old perimenopausal woman" to every query. The AI's default patient is male, and the woman on the other end has no way to know the answer she just read wasn't written for her.
4. The models confuse menopause with unrelated hormone medicine
Asked "what dose of estrogen should I be on," Perplexity answered as if the question came from a transgender patient on gender-affirming hormone therapy, with dosing guidance that could genuinely harm a menopausal woman. Same blind spot as finding 3: without an explicit menopause anchor, the model guesses at the context, and sometimes it guesses badly. On that same question it also recommended the "lowest effective dose for the shortest time" paradigm, which current guidance dropped in 2025 and which the models keep repeating anyway.
5. Even the best-sourced model fabricates citations
After the capture-bug correction, Perplexity has the best citation record in the study: 87.5% of answers backed by real, resolvable guideline and literature URLs. And it still invented two: a fake FDA press release URL claiming HRT warnings had been removed, and a British Menopause Society guideline link stamped with a future 2026 date. Neither exists. A fabricated citation is arguably worse than no citation, because it teaches the reader to trust the exact answer that deserves scrutiny. If the best-sourced model does this, assume they all can.
6. On unsettled science, the models fake certainty
Where the honest answer is "the evidence is mixed," the models tended to pick a side and state it as settled. Perplexity told a woman, as fact, that perimenopause definitively worsens ADHD. Multiple models dismissed testosterone's non-sexual endpoints as settled nulls, ignoring the methodological problems in the older trials. That's how a patient walks into your office already convinced testosterone will fix her brain fog, or already convinced she has adult-onset ADHD, on evidence that doesn't support either claim yet.
7. Six questions stumped every model
Zero out of five models answered these accurately:
- Testosterone for libido
- Testosterone's non-sexual benefits (energy, muscle, cognition)
- ADHD versus perimenopause
- Correct estrogen dosing and whether to chase blood levels
- How to interpret the WHI for a newly menopausal woman
- Whether "lowest dose, shortest time" is still current guidance (it isn't)
Look at that list. It isn't obscure edge-case medicine. It's the exact conversation you have in the exam room every week.
8. The one thing the models do well, and why it's not enough
Credit where due: safety language was strong. Every model said "discuss this with your clinician" and flagged red flags appropriately 90 to 95% of the time. They almost never told anyone to self-prescribe.
But that's precisely what makes the errors dangerous. The failure mode isn't reckless advice a patient would recognize as sketchy. It's confident, fluent, subtly wrong information wrapped in a responsible-sounding disclaimer. The patient hears "talk to your doctor" and thinks the rest of the answer must be trustworthy too.
What this means for you
Your patients arrive pre-misinformed, and they don't know it. AI answers don't read like a sketchy forum post. They read like a knowledgeable friend who did the research. When a quarter to a third of that research is wrong, you inherit the cleanup: the patient who's afraid of HRT because of a 2002-era framing, the one demanding testosterone for brain fog, the one who "read" that she should be on the lowest dose for the shortest time.
The AI blind spots map exactly onto your expertise. The six questions no model could answer are the ones a menopause-informed clinician answers well every day. That's not a coincidence. The models fail where the evidence is nuanced, recently updated, or requires knowing the patient is a midlife woman. That's your entire specialty.
The correction has to exist where patients actually look, and AI is now one of the places they look for YOU. This cuts two ways. First, patients are getting menopause medicine from AI, and a quarter to a third of it is wrong. Second, and this is the newer shift, patients are asking those same AI tools who can treat them: "who prescribes HRT in Colorado," "menopause telehealth doctor near me." The AI engines answer by citing and summarizing web and video content from credible named experts. A clinician who consistently publishes accurate, guideline-current, fully sourced menopause content becomes the source patients find, the source other doctors point to, and the provider the AI cites when a woman asks for a doctor. Given how badly AI answers the medicine itself, that position is wide open. Silence cedes it to whoever publishes, accurate or not.
Demand is not the problem. Behind this study sits a larger analysis of roughly 45,000 patient comments on menopause YouTube. Women are asking these questions in enormous volume, and mostly not getting clinician answers. The gap between what patients are asking and what credible doctors are publishing is the opportunity.
Honest limitations
I'd rather you trust this study because it's straight with you than because it's loud.
- Different question set than the baseline. Our 40 questions come from my real audience, not Karam's list, so the comparison with the 2024 baseline shows direction and rough size, not an identical rematch.
- Different model generations. These are 2026 flagships against Karam's 2024-era models. Any comparison spans two very different generations.
- A post-2025 answer key. Our key reflects current positions (the FDA boxed-warning removal, fezolinetant, re-adjudicated testosterone evidence). A model citing last year's consensus can look "wrong" partly because the ground moved.
- A simulated expert panel. The judge is a strong AI model applying a fixed rubric against a validated answer key, blinded to which model wrote each answer. It's reproducible and transparent, and it's the right first pass. Real-clinician validation is the natural next step.
- One correction, reported in full. A capture bug initially undercounted Perplexity's citations; fixing it moved Perplexity from 4th to 2nd. The before-and-after numbers are published. If we'd hidden that, you'd be right not to trust the rest.
About this study, and a quiet next step
I'm Annette Thompson. I'm a medical technologist, and I write SmartStrongAlive, a menopause-and-longevity publication with a real audience of midlife women. This study exists because those women kept telling me what AI was telling them, and some of it worried me enough to test it properly. The full dataset behind this report (all 40 questions, all 200 raw answers, the rubric, and every score) is published alongside it.
If you're a clinician who treats menopause, here's the practical version of everything above: the demand is real, the misinformation is real, and patients are increasingly asking Google, Maps, and AI tools to find them a doctor. The clinician who gets found there, and whose website and videos prove she knows menopause, wins those patients. I help menopause clinicians build that whole presence: Google Business Profile and AI-search (GEO) optimization so you get found, sourced script drafts and videos (yours to change or ignore) that get you believed, and the publishing system that gets you shared, grounded in the patient-question research you just read.
If that's worth a conversation, you can book a short discovery call here: [BOOKING LINK]. No pitch deck, just a look at what your patients are already asking and whether video is worth your time.
Either way, take the study. Share it with colleagues. And when a patient tells you what ChatGPT said, you'll know exactly how much weight to give it.
Not medical advice. Discuss individual care decisions with a qualified menopause-informed clinician.
Capture-flow spec (internal, not part of the PDF)
Asset. This document, built to PDF via Typst (per ~/memory/topics/typst-pdf.md), 6-10 pages with the SmartStrongAlive-adjacent branding. The Typst build is the follow-on step; this markdown is the master.
Gate. Hosted on the funnel landing page as a gated download. Email required, first name optional. Form posts to a Cloudflare Pages Function backed by the project D1 database ({project-slug}-db per the D1 default rule). Hard rule: the backend must be built, deployed, and tested end-to-end before the form ships. If the backend isn't ready at launch, ship the page without the form rather than with a broken one (or bridge via Formspree temporarily).
Flow.
- Visitor lands on the funnel page (the study is the hook; the page teases the 3-4 most alarming findings without giving the full report).
- Email form → D1 insert (email, first name, timestamp, source tag
llm-study) → success state. - Welcome email (email #1 of the nurture series, see
PLAN-email-nurture-series) delivers the PDF download link immediately. The link points to the hosted PDF, not an attachment. - Nurture sequence follows, ending in the discovery-call CTA that mirrors the report's closing section.
Booking link. The [BOOKING LINK] placeholder in the CTA must be replaced with the live scheduler URL before the PDF is built. Same URL in the welcome email footer.
Measurement. Track downloads (D1 row count by source tag) and call bookings attributed to llm-study. The report itself carries no tracking; the email link can carry a UTM (utm_source=leadmagnet&utm_campaign=llm-study).
Follow-on steps (flagged, not done).
- Typst PDF build with cover page and branded styling.
- Landing page copy + form (blocked on backend per above).
- Swap in the live booking URL.