What Companion Apps Store About You
Companion apps collect five kinds of data: the content of your messages, short profile facts extracted from those messages, when and how often you open the app, device and payment identifiers, and the persona you configured. Only the first two are inherent to the thing working at all. What data AI chat apps collect is rarely exotic, and the real question is which of it the service needs and which it just keeps.
Key Takeaways:
- Message content and extracted profile facts are the categories the product cannot run without. Engagement telemetry, advertising identifiers and indefinite retention are decisions somebody made
- The memory store is a short, structured account of your life, which makes it easier to read and easier to move than the thousands of messages it was distilled from
- A persona configuration records preferences you never typed as a sentence, entered deliberately rather than inferred from behavior
- The Mozilla Foundation reviewed romantic AI chatbot apps for its Privacy Not Included guide in 2024 and rated the category badly, pointing to policies that were unclear about data sharing and gave users little real control (Mozilla Foundation, 2024)
- Deleting the app removes nothing. Account deletion is the action that reaches the stored data, and only as far as the retention terms allow
What Data Do AI Chat Apps Collect
Five categories cover almost everything, and separating them by whether the service actually needs them is more useful than counting them.
| Category | What it is concretely | Needed for the app to work |
|---|---|---|
| Message content | Every message you send and every reply, kept as a transcript | Partly. The current conversation has to exist to answer it. Keeping 3 years of them is a retention decision |
| Extracted profile facts | Short standalone notes pulled from your chats: your job, family names, a health detail, what upsets you | Yes, if memory is the feature you came for |
| Usage timing and frequency | Session start times, message counts, hour of day, streak state, how long you stay after a notification | Rarely. Rate limiting and billing need a little of this. Retention analytics use the rest |
| Device and payment identifiers | Device model, operating system version, IP address, an advertising identifier, a payment processor token | Mostly yes, with the advertising identifier as the exception |
| Persona configuration | The character you built and every edit you have made to her since | Yes as configuration, and it doubles as preference data |
Nothing on that list is unusual. A grocery app holds four of the five. The category difference is not the shape of the data, it is the content: a fitness tracker knows your resting heart rate, and a companion holds the paragraph you wrote at 2 a.m. about whether your marriage is over. Whether a person at the company can read your conversations is its own question, and it is answered in the policy rather than the app.
Which makes “do they collect data” an unanswerable question and the wrong one. Pick your actual worry, find the row it lands in, and the rest of this gets specific fast.
The Extracted Profile Is the Sensitive Part
The memory store is the sensitive object, not the transcript, and that ordering surprises people because the transcript is a hundred times bigger. The store exists at all because the model forgets everything past its context window, so the app distills what it cannot afford to lose.
Consider what each one is to read. Four thousand messages are noisy, repetitive, full of small talk and abandoned threads, and expensive for anybody to go through. The extracted profile is perhaps 40 clean sentences that somebody already did the distilling work on: father died in 2019. In recovery, 14 months. Manager is named Priya, and leaving has been on the table since spring. Does not want children.
That is a dossier. It was built to be retrieved by a machine, which means it is structured, deduplicated, and readable by a person in about 3 minutes. It also travels well in a way a chat log does not, because a tidy list of facts is exactly the kind of thing that survives an export, a database migration, or a company being bought.
Read your own memory list once as though a stranger were reading it over your shoulder. The transcript is raw material and the profile is the finished product, and you will learn more about a company’s privacy posture from one screen of that list than from a page of policy.
What the Service Needs and What It Merely Keeps
The split is cleaner than the marketing on either side suggests. To answer you at all, an app needs the message you just sent, enough of the current conversation for the reply to make sense, a stored profile if you want continuity between chats, a payment token if you pay, and an account identifier so somebody else cannot log in as you.
To answer you, an app does not need years of retained conversation, an advertising identifier shared with a third-party analytics vendor, engagement telemetry fine-grained enough to time a push notification for the hour you are most likely to be alone, or your conversations included in a training set for future models. The one arrangement that removes the storage question rather than managing it is running the model on your own hardware, paid for in conversation quality.
That last item is the one people care about most and check least. Training on companion chats is a category-specific problem rather than a generic one, because the material is not product feedback about a checkout flow. It is disclosure, produced by someone who was writing to a character rather than to a company.
The Mozilla Foundation looked at a batch of romantic AI chatbot apps for its Privacy Not Included guide in 2024 and came away with a poor assessment of the category, noting policies that left data sharing vague and offered users very little practical control (Mozilla Foundation, 2024). Nothing about that finding names any one company, and it does not have to. It describes what happens when a product category grows faster than anyone’s willingness to read its terms.
One sentence in a policy decides the training question for you. Find that sentence before you decide how you feel about the rest.
The Persona You Built Is Data Too
A persona configuration is a disclosure, and almost nobody counts it as one. It sits in the settings menu looking like a preference pane, next to notification toggles and a theme picker.
Look at what it actually records. Age, build, temperament, how the character responds when you are short with her, what she is permitted to bring up unprompted, whether she initiates. Nobody would fill in a form headed “describe what you want in a partner,” and the character sheet collects the same information more precisely, because you were paying attention while you built it. Ad networks spend years inferring this class of preference from clicks. Here it arrives typed in, by hand, and carefully.
Edits over time carry their own signal. Someone who moves a character from playful toward steady across 4 months has described a change in their life without writing a word about it, and a timestamped edit history holds that just as well as a diary would. Even a character pulled ready-made from an ai character catalog records every edit you make to it the same way.
There is a version of this that is genuinely fine. An app that stores configuration under your account and never treats it as analytics is doing nothing objectionable with it. Just do not assume it, because a persona is the one piece of stored data users hand over voluntarily and then forget they created.
What You Can Check in 10 Minutes
Five checks, all of them findable tonight, none requiring a lawyer.
- Open the memory screen, read the saved list end to end, and delete anything you would not want restated back to you by a stranger.
- Search the privacy policy for “retain” and look for a number of days after account deletion. No number is itself an answer.
- Search for “train”, “improve our services”, “model improvement”, “machine learning”. Note whether an opt-out exists and whether it is on by default.
- Search for “service providers” and “third parties” and read which categories of vendor appear, particularly analytics and advertising.
- Find out what deletion covers: whether closing the account clears the memory store and backups, or only disables the login.
What does not work is the app store privacy label. It is self-reported by the developer, it lists categories rather than retention periods or sharing terms, and a product can be honest on that label while keeping every conversation indefinitely. The in-app “clear chat” button fails the same way for a different reason: it usually clears what you see, and the extracted facts stay exactly where they were.
Do the first check before the other four. Everything else is a document about the data, and the memory list is the data.
FAQ
Do AI chat apps save your conversations? Yes, by default, on company servers rather than on your phone. Conversations have to persist somewhere for yesterday’s chat to load and for a memory feature to have raw material. Clearing a chat inside the app typically removes it from your screen without touching the stored copy, which lives behind a different control in a different menu.
Can I delete what an AI chat app knows about me? Usually yes, in two separate places, and most people find only one of them. Individual entries can be edited or removed on the memory screen. Transcripts and account records go with account deletion, which is not the same as uninstalling the app, and the policy should say how long backups hold a copy after that.
Is my chat data used to train the AI? Sometimes, and the privacy policy is the only reliable place to find out. Look for language about improving models or services, then check whether an opt-out exists and whether it is enabled by default. The stakes are higher here than in most apps, since the material is personal disclosure rather than feedback about a feature.
What is the most sensitive thing a companion app holds about me? The extracted profile rather than the transcript. It is a short list of clean statements about your family, health, work and moods, readable in a couple of minutes, and it was deliberately built to be easy to retrieve. Read your own before deciding how you feel about anything else on the list.
Privacy arguments in this category fixate on the transcript, because a transcript feels like the intimate object. It is the raw material. The finished object is a compact profile the app wrote about you out of your own messages, and it exists because you wanted the thing to remember you, which is a trade plenty of adults will make with open eyes. It stops being that trade the moment nobody looks at what got written down. Open the list. Whatever sits on it is what the company holds, what a support engineer would see, and what would move with the business if it were ever sold.
Before you let Lona hold the things you would only type to it, run the first check from this page: open the memory screen, read the short profile it wrote from your own messages, and delete whatever you would not want said back to you by a stranger.
Sources
- Mozilla Foundation, *Privacy Not Included, review of romantic AI chatbots, 2024
