Issue #012 — Relapse becomes the frontier: two independent 2026 reviews map AI for predicting psychiatric relapse and land at modest AUCs, while a 95-study scoping review calls the whole LLM-in-mental-health field 'nascent and exploratory.'
Weekly Intelligence · Week 12 · 24 July 2026 · Issue #012
The strict 7-day window (17–24 July) stayed quiet for the second week running, so this issue again runs on the catch-up track — but the three reviews it surfaces line up on a subject the newsletter has not yet covered directly: not detecting a first episode, but predicting relapse and deterioration in people already diagnosed. Two independent 2026 reviews map that frontier and reach the same modest numbers; a third, broader scoping review maps the LLM landscape around it.
Executive Summary
For the second consecutive week, no new wearable, speech, multimodal, facial-CV, NLP, or digital-phenotyping primary detection result — and no new regulation clearing the bar — landed inside the strict 7-day window (17–24 July). The freshest industry signal, a NeuroLexIQ–Canary Speech voice-intake partnership announced 15 July, falls just outside the window and outside the backfill track's high-leverage-type gate, so it is logged rather than padded in. Instead this issue leans on three catch-up reviews that, read together, move the newsletter's measurement-discipline thesis onto a new axis. Since Issue #008 the through-line has been about detection — external validation, single-marker insufficiency, unstable leaderboards, implementation and measurement heterogeneity. The three reviews here are about the harder, more clinically valuable task the field is now turning to: predicting relapse and deterioration in people already diagnosed. First, a BMC Psychiatry systematic review (Dormechele et al., May 2026) of passive-sensing approaches to relapse prediction across psychiatric disorders — prospective and retrospective observational studies using passively collected smartphone and wearable data, searched to January 2026, whose verdict is that behavioral digital-phenotype shifts are plausible early-warning signals but the evidence base is not yet dependable. Second, a JMIR Mental Health scoping review (Ghelfi et al., 16 June 2026) narrowing to psychosis relapse, which reports AI-model AUCs spanning a modest 0.63–0.78 and concludes that personalized, individual-level modeling shows promise but needs far larger samples and newer methods before clinical use. Third, a JMIR Mental Health scoping review (Jin et al., 5 May 2025) of 95 studies applying large language models across mental health — 67 of them (71%) on screening and detection — whose bottom line is that despite explosive growth the field remains "nascent and exploratory." The honest read: the in-window frontier was static again, but the catch-up shelf shows the field's ambitions climbing from detection toward prediction — and carrying the same generalization and sample-size problems up the ladder with it.
Key Metrics
| Metric | Value | Source |
|---|---|---|
| Psychosis-relapse AI scoping review: reported model AUC range | 0.63–0.78 | Ghelfi et al. · JMIR Ment Health · 16 Jun 2026 |
| LLM-in-mental-health scoping review: studies / screening-detection share | 95 / 67 (71%) | Jin et al. · JMIR Ment Health · 5 May 2025 |
| Relapse-prediction systematic review: search window / data modality | to Jan 2026 / passive smartphone + wearable | Dormechele et al. · BMC Psychiatry · May 2026 |
Wearable Biosensors & Digital Phenotyping
A systematic review moves the goalpost from detection to relapse — and finds the evidence thin
Wisdom Dormechele, Isaac Yeboah Addo, Caleb Boadi, and colleagues published a systematic review in BMC Psychiatry of passive-sensing approaches to predicting relapse in psychiatric disorders — a deliberate step past the detection-and-screening question this newsletter has tracked all quarter. Searching PubMed, PsycINFO, and IEEE Xplore for studies through January 2026, the review gathered prospective and retrospective observational studies that used passively collected smartphone or wearable data to forecast relapse or clinical deterioration in people with a diagnosed disorder, and that reported quantitative model performance. The framing is the field's most clinically valuable premise: because behavioral digital-phenotype shifts — in mobility, sleep, communication, and social rhythm — may register before a person notices a downturn, passive monitoring opens a window for timely, preventive intervention rather than after-the-fact diagnosis. The review's sober contribution is that this premise is still mostly a promissory note: the underlying studies are observational, heterogeneous in devices and outcome definitions, and small, which keeps relapse prediction short of the reliability routine care would demand. For this newsletter the paper is the natural next chapter after the detection audits of Issues #008–#011: the same passive-sensing apparatus is now being pointed at a harder target, and the same constraints — sample size, heterogeneity, un-validated models — follow it there. Predicting a relapse three days out is a far more useful clinical act than scoring a cross-sectional screen, which is precisely why the gap between the promise and the evidence matters more here, not less.
Source: Dormechele W, Addo IY, Boadi C, et al. · BMC Psychiatry · May 2026 · 10.1186/s12888-026-08157-z
📅 Catch-up — published May 2026, outside the weekly window
AI/ML for Mental Health Detection
A psychosis-relapse scoping review puts a number on it: AI AUCs of 0.63–0.78
Luca Ghelfi, John Healy, Federico Piacenza, and a large multi-site author group — spanning Irish, Dutch, Norwegian, Turkish, Swiss, and Spanish centers and ending with senior authors Mary Cannon and John Lyne — published a scoping review in JMIR Mental Health narrowing the relapse question to its most-studied condition: psychosis. Searching PubMed, PsycINFO, and Embase from inception to January 2026 for any method with an AI component used to detect psychotic relapse, the review's headline finding is a quantitative one this newsletter can carry forward: across the studies that reported it, AI-model discrimination landed at a modest AUC of 0.63–0.78 — meaningfully better than chance, but well short of the 0.9-plus figures that decorate in-sample detection papers, and a useful reality check on how much harder prediction is than classification. The authors report that passive digital-phenotyping research on psychosis relapse has genuinely progressed and that personalized, individual-level modeling — learning each patient's own behavioral baseline rather than a population average — is the most promising direction, but that the field's studies still need substantially larger participant numbers and should begin incorporating newer methods, including large language models, before the approach is clinic-ready. Read against the Dormechele systematic review above, the two form an unusually clean convergence: one review synthesizes relapse prediction broadly and finds the evidence thin; the other zooms into psychosis and attaches the actual number — 0.63–0.78 — to why. The AUC range and the "needs larger samples" caveat belong in the same sentence, and together they place relapse prediction exactly where detection sat in Issue #008: a real, personalized signal whose generalization is not yet earned.
Source: Ghelfi L, Healy J, Piacenza F, et al. · JMIR Mental Health · 16 June 2026 · 10.2196/92192
📅 Catch-up — published 16 June 2026, outside the weekly window
NLP & Large Language Models
A 95-study scoping review: LLMs are everywhere in mental health, and the field is still "nascent"
Yu Jin, Jiayi Liu, Pan Li, and colleagues published a scoping review in JMIR Mental Health mapping the full landscape of large-language-model applications across mental health — the breadth-map counterpart to the depth audits this newsletter has run on LLM safety (MHSafeEval, Issue #011) and LLM benchmarks (Ishikawa & Duke, Issue #009). Across 95 included studies, the authors sorted the work into three uses: screening or detection of mental disorders, by far the largest at 67 of 95 (71%); supporting clinical treatment and intervention (31/95, 33%); and assisting mental-health counseling and education (11/95, 12%). The through-line is that the application surface has already sprawled far ahead of the evidence: the review's own summary judgment is that despite the rapid growth and diversity of LLM use, the field remains "nascent and exploratory," dominated by short-horizon, single-session, small-sample evaluations rather than the longitudinal, externally-validated studies clinical adoption would require. For this newsletter the value is contextual — it quantifies just how top-heavy the field is toward detection (the same 71% skew the newsletter keeps encountering) and frames why the safety and benchmark audits of the last two issues matter: a technology this widely applied and this thinly evaluated is precisely the kind that needs disciplined measurement before it touches care. It also sharpens Ghelfi et al.'s prescription above — "incorporate large language models" — with a caution: the LLM layer the relapse field is being urged to adopt is itself, on this review's own reading, not yet a mature instrument.
Source: Jin Y, Liu J, Li P, et al. · JMIR Mental Health · 5 May 2025 · 10.2196/69284
📅 Catch-up — published 5 May 2025, outside the weekly window
Forward Outlook
- Near-term: The two relapse reviews now travel as a pair — Dormechele et al.'s "promising but thin" synthesis and Ghelfi et al.'s concrete AUC 0.63–0.78 for psychosis — and together they give a reviewer a citable anchor for any claim that passive sensing can forecast deterioration. The next result worth flagging is the first relapse-prediction model reporting prospective, externally validated, individual-level performance at a stable AUC, rather than a retrospective in-cohort figure — the same artifact the detection track has been waiting on, now one rung harder.
- Mid-term: Both reviews point at personalization — modeling each patient's own behavioral baseline — as the direction of travel, and at larger, longer cohorts as the missing ingredient. If that holds, the deployment shape for relapse prediction converges on the same triage-grade, human-in-the-loop use the detection literature (Issues #009–#010) and the governance track (WHA79, Issue #007) have been circling: an early-warning nudge routed to a clinician, not an autonomous alarm. Ghelfi et al.'s call to fold in LLMs, read against Jin et al.'s "nascent and exploratory" verdict, suggests the field will graft an immature instrument onto an immature task — worth watching closely.
- Long-term: With this issue the newsletter's binding-constraint thesis extends from detection to prediction: across first-episode screening (Issues #008–#011) and now relapse forecasting, the limiting factor is not model capacity but the trustworthiness of the evidence — sample size, external validation, and prospective design. Accuracy and ambition will keep rising; whether the reviews, benchmarks, and safety evaluations measuring them are disciplined enough to believe is still the open question, and moving up from detecting illness to predicting its return raises the stakes on that question rather than settling it.
Sources used: 3 (0 in-window · 3 catch-up) · Week 12 · Next issue: 31 July 2026