Issue #014 — The week the field measured across the boundary: a cross-linguistic benchmark where word count beat 50 rival metrics, a suicide-attempt model validated on seven held-out sites at AUC 0.75, and a scoping review that found only ten studies exist.
Seven in-window results, and for the first time the honest ones dominate: every finding that tested transfer across a language, a site, or a population landed in a 0.72–0.83 band, one week after single-corpus papers reported 94.7%.
Issue #013 — The window reopens: seven in-window results land at once, and the field starts building for generalization — a domain-adversarial detector, a wearable-plus-MRI fusion at AUROC 0.867, and a model-card framework — while five more states go their own way on AI therapy.
After two catch-up-only issues the 7-day window produced seven results, and they show a field that has adopted the vocabulary of generalization faster than the practice: domain-adversarial training measured within a single corpus, a 94.7% voice accuracy on one dataset, a scoping review proposing model cards as the remedy, and a governance viewpoint calling for a federal floor on the same day a fifth state legislated past it.
Issue #012 — Relapse becomes the frontier: two independent 2026 reviews map AI for predicting psychiatric relapse and land at modest AUCs, while a 95-study scoping review calls the whole LLM-in-mental-health field 'nascent and exploratory.'
A second consecutive quiet in-window week; three out-of-window catch-up reviews shift the newsletter's measurement-discipline thesis onto a new axis — relapse and deterioration prediction — where a BMC Psychiatry systematic review and a JMIR psychosis-relapse scoping review both report promise undercut by small samples, and a 95-study LLM scoping review maps a field still nascent.
Issue #011 — A quiet in-window week, so three catch-up audits — a 105-study speech meta-analysis, a 42-study passive-sensing scoping review, and an LLM mental-health safety benchmark — widen the measurement-discipline thesis across three modalities at once.
No new primary detection result cleared the 7-day window; three out-of-window catch-up findings — an automatic-speech-analysis depression meta-analysis, a passive-sensing scoping review, and a role-aware LLM safety benchmark — independently report strong aggregate numbers undercut by heterogeneity, small samples, and unmeasured multi-turn failure.
Issue #010 — The in-window window finally produces its own story: two digital-phenotyping papers land six days apart — a Nature Mental Health comment celebrating the promise of early adolescent depression detection, and a 47-study Frontiers systematic review finding that methodological heterogeneity still blocks its translation.
After four catch-up weeks, two genuinely in-window digital-phenotyping papers frame the field's core tension — promise versus implementation heterogeneity — reinforced by two catch-up findings on speech-biomarker parsimony and just-in-time prediction.
Issue #009 — A fourth quiet in-window week, so three 2026 audits of the wearable-biomarker literature converge on one uncomfortable verdict: no single signal is diagnostic, passive sensing is population-level not clinical, and the leaderboards ranking detection models are unstable.
No new primary detection result cleared the 7-day window; three out-of-window catch-up findings — a 132-study depression digital-biomarker meta-analysis, a 10-month wearable brain-health study, and a five-dataset benchmark audit — independently land on the same measurement-discipline conclusion.
Issue #008 — A quiet in-window week, so three catch-up results converge on one theme: the modality, the validation gap, and the LLM that decide whether detection survives contact with real patients.
No new primary detection result cleared the 7-day window; three out-of-window catch-up findings — a visual-psychophysiology screener that catches 'silent' patients, a multimodal-MDD review exposing the external-validation gap, and a RAG-LLM depression/suicide-risk benchmark — anchor the issue.
Baseline — the state of human behavioral analysis for early identification of mental health conditions
Foundational state-of-the-field report. The dedup baseline against which every weekly issue is measured.
Every weekly issue is deduplicated against the Baseline — the foundational state-of-the-field report. Start here if you're new to the series.