Somewhere in the last few months, a stranger may have read the conversation you had with ChatGPT at two in the morning, and the reason is almost funny once you see it. According to a 404 Media investigation, OpenAI has been paying hundreds of contractors over $50 an hour, under an internal initiative called “Project Lily,” to read, summarize and rate real user conversations, and the main goal is to teach ChatGPT to sound less like a machine and more like a natural conversation partner. So humans read your most human moments in order to help the machine imitate them better, which is the kind of irony you couldn’t sell as fiction because nobody would believe it.
The same Monday, Apple joined the club. A new section in the iOS 27 privacy policy reverses Apple’s earlier position that it had no need for customer data to make Siri’s AI work, and the policy now allows Apple to store entire Siri and dictation interactions, including audio and transcripts, with some of that data reviewed by unspecified “review personnel”. It is unclear whether those reviewers are Apple employees or a third-party company contracted to read transcripts, as has happened historically. This is the company that spent years telling us that what happens on your iPhone stays on your iPhone, and now the fine print says that some of it may travel to a desk you will never see.
You can read more here: https://appleinsider.com/articles/26/09/14/apple-has-altered-course-on-using-customer-data-to-train-its-ai
Most people imagine a chatbot as a diary that talks back, private by nature because nobody else is in the room. The reality looks more like an office building, with staffing agencies, hourly rates, shift schedules and leaked training guides. The Project Lily contractors are reportedly paid through third-party staffing firms, and internal documents show that reviewers see whole conversations, and that sensitive personal details routinely get past the company’s automated safeguards. That gap between the diary people imagine and the office that exists is where the whole problem lives.
OpenAI’s defense is anonymization, and on paper it sounds reasonable. The company removes account names and runs chat logs through an automated Privacy Filter, but it acknowledges the filter can fail on rare identifiers or ambiguous phrasing, and contractors can still see a “user memories summary” showing a person’s past interests and approximate location. Removing the account name while leaving the memory summary in place is like blurring the license plate of a car and leaving the owner’s home address taped to the windshield. You don’t need a name when the conversation itself says “my manager at the only veterinary clinic in a town of four thousand people just found out about my divorce.”
I spend part of my professional life as a court-appointed expert in digital forensics, and I can tell you that re-identifying someone from context is routine work. Text is a fingerprint, because people write about their jobs, their neighborhoods, their kids’ schools and their bosses, and after a few paragraphs the “anonymous” user has a profile sharp enough to recognize at a family barbecue. :D
And the audience for these chats doesn’t stop at contractors. A Washington Post review found chatbot conversations mentioned in 12 public court cases over the last two years, a number that is probably higher because much of the evidence gathered in investigations never becomes public. In one employment dispute, a salesman had asked ChatGPT whether deleted Yahoo emails could still be retrieved, including through a subpoena, and his employer argued the question showed he was trying to hide evidence. None of these conversations carry the legal confidentiality that protects what you say to a lawyer, a doctor or a therapist, so every heartfelt prompt is, legally speaking, a document sitting on someone else’s server.
Then there is the design of consent, which deserves its own paragraph because it is so revealing. When a user enables Siri AI, the setup screen offers two options: “Share Audio and Text” or “Not Now”, and heise described that second button as visually inconspicuous. “Not now” is a deferral, a polite way of saying the question will come back, and anyone who has worked on product design knows that the second, third and fourth prompts exist precisely to wear people down. On the OpenAI side, the “Improve the model for everyone” setting does not automatically remove chats that were already eligible for model improvement when a user later opts out. Consent that only works going forward, while the data already collected keeps its travel plans, is a strange kind of control.
This is where companies are fooling themselves. The industry treats the toggle as the end of the privacy conversation, as if an opt-in screen answered every question, when the questions that matter are about the pipeline behind the screen: who reads the data, which staffing firm employs them, in which country, under what contract, with what retention, and with what memory summary pinned above the prompt. A toggle without that transparency is decoration.
If you run a company or work as a DPO, pay attention to one detail. OpenAI says consumer chats may help train models unless users opt out, while Business, Enterprise and Edu accounts default to no training. That means the protection your company pays for does nothing for the employee who pastes a client contract into a personal account on their phone during lunch. Anyone who has built a record of processing activities knows the uncomfortable question about sub-processors, and for consumer chatbots used informally by staff, the honest answer in most organizations is that nobody knows who ends up reading that data. Your AI policy should say plainly which tools are approved, and your vendor map should treat the chatbot, and the invisible workforce behind it, as part of the data flow.
For everyone else, the practical steps are simple and a little boring, which is usually a good sign. Turn off the training setting in ChatGPT, remembering that it only protects what comes after, use temporary chats for anything sensitive, and on the iPhone read the Siri AI setup screen before tapping whatever button looks biggest. And think about the people in your conversations, because when you ask a chatbot how to support your sister through her diagnosis, you are sharing her health data too, and she never got a setup screen at all.
I don’t think human review is evil. Training these systems requires human judgment, and some of the reviewers’ instructions are sensible, since they also have to make sure ChatGPT never pretends to have real human lived experiences. My problem is with the story being told to users. People confide in chatbots because they feel like a room with no one else in it, and the companies know this, because that feeling of intimacy is exactly what Project Lily is paid to improve. When the product works by making you feel alone with it, the least you are owed is a clear answer about who else is in the room.



