Using ChatGPT as a journal: what the evidence actually supports
Millions of people write their lives into a chat window. Here is what randomized studies support, what OpenAI has documented about its own system, and what is design.
Across 146 randomized studies, expressive writing showed a small overall effect — and a larger one when people were asked directed questions instead of facing a blank page.
Chatterji and colleagues (2025) sampled 1.1 million ChatGPT conversations for the National Bureau of Economic Research and classified 1.9% of messages as relationships and personal reflection. A thin slice of an enormous surface, and the one where people write what they tell nobody. There is a material reason, too: the World Health Organization’s Mental Health Atlas 2024 puts the global median at 13.5 specialized mental health workers per 100,000 people. By income group: 1.1 in low-income countries, 2.4 in lower-middle-income ones, 67.2 in high-income ones.
The short answer is awkward for both camps: writing in there works, and it works less than its defenders promise or its critics fear.
Frequently asked questions
Does typing into a chat count as journaling? By the only measure the research offers, yes. Frattaroli (2006) found that the mode of disclosure — handwritten, typed or spoken — did not moderate the effect on any outcome category. That is also the honest answer to the paper-versus-screen argument.
Is a purpose-built app better than ChatGPT for this? Not demonstrated, in either direction. Across the trials pooled by Sohn and colleagues (2026), the way a chatbot generates its replies — a generative model or retrieval from a script — was not a significant moderator. That subgroup compares two ways of building a mental health chatbot, not a specialized product against ChatGPT, and the authors warn that the generative side, eight trials against thirty-one, may lack the power to show a difference. The closest thing to that comparison is a single randomized pilot, and it ended in a tie.
If I turn off model training, is the conversation private? More private, not privileged. OpenAI’s consumer products use content to train models by default and the opt-out is not retroactive: new conversations stop being used, earlier ones are not withdrawn. Rating a reply can make that conversation usable for training even with the setting off. No hosted service holds privilege, which applies to any journaling app you might use instead, mine included.
Should I use this instead of talking to a person? No, and that holds for every AI product, mine included. This literature measures symptom scores against a control condition, a narrower question than the one you are asking. In the non-clinical samples pooled by Sohn and colleagues (2026) the effect was 0.07, with a confidence interval crossing zero. A number that small is not an argument for replacing anything.
The ingredients that make writing work have been measured
Four of them, and a chat window supplies exactly one by design — not the strongest of the four. Frattaroli (2006) pooled 146 randomized studies of experimental disclosure and found a small overall effect, r = .075, then measured what raises it: repetition, minutes on the page, a subject that is recent and unresolved, and directed questions instead of a blank page.
Three or more sessions produced r = .082, against .040 for fewer, a marginal contrast. Sessions of at least fifteen minutes reached r = .080 against −.007 for shorter ones, and that difference was significant, as was the pull of time: the further back the event, the weaker the effect (r = −.283, one-tailed p = .013). Directed questions or concrete examples reached r = .090 against .052, and .094 against .011 on psychological health — marginal on the overall effect (one-tailed p = .055), significant on psychological health outcomes (p = .0035). Reinhold, Bürkner and Holling (2018) found the same two moderators independently: a specific topic (b = 0.09, p = .020) and more sessions (b = 0.03, p = .038). They also report the least flattering result in this literature. Among physically healthy adults the effect right after the last session was significant but very small (g = −0.09, p = .006), and by the first follow-up it no longer differed from zero (g = −0.03, p = .296).
A chat window asks by default, by design rather than by the merit of any brand. The other three ingredients are still yours to arrange.
How large the chatbot effect actually is
Modest overall, and close to zero for the person most likely to be reading this. Sohn and colleagues (2026) reviewed 39 randomized trials in npj Digital Medicine; the 38 analyzed for depressive symptoms (n = 7,401) gave g = 0.31 (95% CI 0.17–0.46), and anxiety gave g = 0.28 (0.05–0.51). Different literature, different scale: those numbers do not line up against Frattaroli’s r. Split by population, the effect was 0.64 in clinical samples, 0.34 in subclinical ones, and 0.07 — with a confidence interval crossing zero — in non-clinical ones.
The closest thing to a head-to-head is a feasibility pilot. Kuta and colleagues (2026), in JMIR Mental Health, randomized English-speaking adults into three arms for three weeks: a purpose-built mental health chatbot (n = 44), ChatGPT (n = 60), and an assessment-only control that filled in questionnaires and nothing else (n = 43), 147 people in all. Against that control, both chatbots lowered PHQ-9 depression scores (d = −0.47 and d = −0.44); on anxiety and well-being neither did, and the purpose-built app did not significantly outperform ChatGPT on any outcome. Adherence ran the other way: 62% of the ChatGPT group completed all nine conversations, against 39% in the app. Its authors call the sample and the three weeks too small for clinical claims, a caveat that cuts against my product too.
The pull is older than the models. Lucas, Gratch, King and Morency (2014) told some participants that a machine, not a person, was running the virtual interviewer: that group reported less fear of self-disclosure and expressed sadness more intensely. Yin, Jia and Wakslak (2024) found AI-generated messages made people feel more heard than human-written ones, an effect an AI label reduces without erasing.
What the research on writing supports, what it does not, and the limits left in rather than edited out.
What OpenAI has documented about its own system
Sycophancy, in writing, as a property of the training rather than an accident. In April 2025 OpenAI rolled back an update to GPT-4o that had turned excessively eager to please, days after shipping it: the company described behavior that validated doubts, fueled anger, urged impulsive actions and reinforced negative emotions, and warned of safety concerns including mental health, emotional over-reliance and risky behavior. Its follow-up post named the root cause: a reward signal built on user thumbs-up and thumbs-down data, and user feedback that sometimes favors agreeable answers. Any system trained on human preference carries that gradient.
The same post holds the sentence to read before romanticizing memory in any product: in some cases, OpenAI wrote, user memory contributes to exacerbating the effects of sycophancy, though the company adds it has no evidence that memory broadly increases it.
Peer review describes the same failure. Moore and colleagues (2025) found language models expressing stigma toward mental health conditions and responding inappropriately to common scenarios; they note that models encourage clients’ delusional thinking, likely because of sycophancy, and that this held for the newer and larger ones. The American Psychological Association described the same loop in its November 2025 health advisory.
Fabrication is the other half. OpenAI published in September 2025 that models hallucinate because standard training and evaluation reward guessing over acknowledging uncertainty. Walters and Wilder (2023) counted 55% fabricated bibliographic citations in GPT-3.5 and 18% in GPT-4, the models of 2023. Dahl, Magesh, Suzgun and Ho (2024) measured legal queries: the models they tested hallucinated in more than half of the cases and “often uncritically accept users’ incorrect legal assumptions”. That is a finding about law, not about biographies. Nobody has measured how often a model accepts a false premise about your own life.
What is design, and not the quality of the model
First the credit due: ChatGPT’s history is literal, dated, searchable and exportable, which is worth saying plainly against the claim that a chat leaves no trace. What follows are conditions of the medium that no update fixes and no brand escapes, mine included.
- There is no privilege. Sam Altman said so in July 2025: if someone discusses their most sensitive matters with ChatGPT and a lawsuit follows, the company could be compelled to produce them. OpenAI is asking for that protection to cover AI conversations; until it exists, this is a condition of where you keep a journal, not a difference between ChatGPT and a smaller app.
- Someone else’s lawsuit can freeze the record. Under the order in the case brought by The New York Times, OpenAI preserved consumer ChatGPT content, deleted chats included. The obligation ended on 26 September 2025; what was preserved between April and September stays preserved.
- Memory is not enumerable. OpenAI documents that its memory summary does not include everything the system remembers, and that removing something completely means deleting every source it appears in. You can reread your history; you cannot list what was concluded from it.
None of the three is an advantage I can claim: they hold for the product I work on too.
What nobody has measured yet
There is no randomized trial of keeping a journal with ChatGPT, and no controlled study isolating memory as a variable and showing better outcomes: what exists is the warning from the manufacturer, hedged by the manufacturer. None of the 39 trials Sohn reviewed ran in a Spanish-speaking country, and Kuta recruited English-speaking adults, so most of the planet is reading conclusions drawn somewhere else.
The version of this question that ages badly is which window to type into: which model, which default, which interface. The one that does not is how much of your life goes into that box, and whether what answered you was checking or agreeing.
References
- Frattaroli, J. (2006). Experimental disclosure and its moderators: A meta-analysis. Psychological Bulletin, 132(6), 823–865. https://doi.org/10.1037/0033-2909.132.6.823
- Reinhold, M., Bürkner, P.-C., & Holling, H. (2018). Effects of expressive writing on depressive symptoms—A meta-analysis. Clinical Psychology: Science and Practice, 25(1), e12224. https://doi.org/10.1111/cpsp.12224
- Sohn, J.-S., Ha, B.-G., Park, S., Kim, J.-J., Lee, E., Oh, H., Lee, S., & Kim, E. (2026). Systematic review and meta analysis of chatbots in the management of depressive and anxiety symptoms. npj Digital Medicine, 9(1), 377. https://doi.org/10.1038/s41746-026-02566-w
- Kuta, B., Novak, L., Zidkova, R., Furstova, J., Malinakova, K., De Winter, A., & Husek, V. (2026). Effectiveness of a Fully Automated Mobile Therapeutic Versus a General Chatbot in Reducing Depression and Anxiety and Improving Well-Being: Feasibility Randomized Controlled Trial. JMIR Mental Health, 13, e82642. https://doi.org/10.2196/82642
- Moore, J., Grabb, D., Agnew, W., Klyman, K., Chancellor, S., Ong, D. C., & Haber, N. (2025). Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers. Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, 599–627. https://doi.org/10.1145/3715275.3732039
- Lucas, G. M., Gratch, J., King, A., & Morency, L.-P. (2014). It's only a computer: Virtual humans increase willingness to disclose. Computers in Human Behavior, 37, 94–100. https://doi.org/10.1016/j.chb.2014.04.043
- Yin, Y., Jia, N., & Wakslak, C. J. (2024). AI can help people feel heard, but an AI label diminishes this impact. Proceedings of the National Academy of Sciences, 121(14), e2319112121. https://doi.org/10.1073/pnas.2319112121
- Chatterji, A., Cunningham, T., Deming, D., Hitzig, Z., Ong, C., Shan, C. Y., & Wadman, K. (2025). How people use ChatGPT. National Bureau of Economic Research Working Paper 34255. https://doi.org/10.3386/w34255
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13, 14045. https://doi.org/10.1038/s41598-023-41032-5
- Dahl, M., Magesh, V., Suzgun, M., & Ho, D. E. (2024). Large legal fictions: Profiling legal hallucinations in large language models. Journal of Legal Analysis, 16(1), 64–93. https://doi.org/10.1093/jla/laae003
Double confirmation: you get an email and decide. One-click unsubscribe, anytime.