The hard question for AI mental health is whether the app is built to move you forward or to keep you subscribed.
A user opens an AI mental health app every night for two months, then stops. Not because the AI was unkind. Because the conversation kept letting them stay stuck. The split is clean: products built to comfort, and products built to move you somewhere.
That comment surfaced in a recent podcast interview with Flourish Science founder Zhao Xuan, a Brown psychology PhD and former Stanford SPARQ researcher. Flourish is the small end of a category that ships companion AI into schools and into insurer benefit plans. The product, Sunnie, takes the talk-therapy methods cognitive behavioral therapy (CBT) and dialectical behavior therapy (DBT), breaks them into small daily exercises, and remembers what worked. The company's about page and press materials present Sunnie as the bottom rung of a stepped-care ladder: not therapy, but a step toward it.
Whether the design works is harder to read than the marketing suggests. Zhao said on the same podcast that 50% of paying users are still active 90 days in, versus 18% of free users. He cast the gap as evidence that the app pushes when you let it. The number is a single guest's claim on a podcast show-notes page, not a third-party retention audit. Read it as a direction, not a verdict.
The harder counterweight is safety. Open-ended AI conversation is not neutral for users with obsessive-compulsive or delusional thinking patterns. An LLM that agrees with a spiral is not a therapist. It is a feedback loop. Zhao flagged this in the same interview, calling out cases where Flourish routes users to human care. The category has not produced a public incident database the way the FDA's adverse-event system has for drugs, and no independent clinical voice has audited any consumer AI mental health product for this failure mode. That gap is load-bearing.
Does the app remember what worked last time. Does it name the next step. Does it refuse to be a yes-machine. Flourish founder notes on Substack describe motivational interviewing as a "midwife" practice: pulling the next move out of the user instead of pushing it on them. That is the mechanism. A comforting AI makes the user feel heard. A progressive AI makes the user act, and remembers which act landed.
The buyer question is where the comfort/progress split becomes a market problem. Zhao told the podcast that school buyers care about retention-to-graduation and insurance buyers care about ROI per claim. The two are not the same metric. A product that helps a student finish a semester is not the same as a product that reduces a hospital bill. The buyer whose spreadsheet measures "did the user move" is a different buyer from the one whose spreadsheet measures "did the user stay subscribed."
The academic claim is the part to watch. Flourish has posted about what it calls a multi-institutional longitudinal RCT, the arXiv preprint 2601.11530 titled "AI for Proactive Mental Health." A preprint is not peer review. The version, author list, and effect sizes have not been independently verified. If the study holds up, it is the first peer-trackable RCT in this category. If it does not, the marketing claim is the part that ages badly.
The user who stopped opening the app is the part of the story the marketing will not print. Comfort is cheap. Progress is a design choice, a memory system, and a willingness to refuse the user a little. The category will split on which one the buyer is paying for.