CourionAI
EN
Newsletter
← All news
privacy 3 min read

Real People Are Reading Real ChatGPT Conversations, and Most Users Have No Idea

404 Media obtained internal documents describing Project Lily, OpenAI's programme of hiring hundreds of contractors to read genuine user prompts and grade the answers. Personal details slip through the filters, and the setting that controls it is on by default.

A stack of sealed envelopes with the top one open and a magnifying glass above the letter

404 Media published an investigation into an OpenAI programme internally called Project Lily, based on leaked documents and on real prompts the outlet was shown. The short version: hundreds of contractors are paid to read genuine conversations between people and ChatGPT and grade the chatbot’s replies on a scale from one to seven. Not anonymised summaries, not synthetic test questions. Actual chats, including, by OpenAI’s own admission, ones that still contain personal details.

The contractors are recruited by a firm called Crossing Hurdles and paid through the hiring platform Mercor, with one worker reporting over 50 dollars an hour. They do not see usernames, and OpenAI says it tries to strip personal information before prompts reach a reviewer, but the company acknowledges that sensitive details get through anyway. Which is not surprising once you picture what people actually type: the medical worry, the message to a difficult relative, the contract they pasted in to have explained. ChatGPT has more than 900 million users, and on consumer plans the setting that allows conversations to be used for improving the model is switched on unless you turn it off. This is not an OpenAI peculiarity, to be fair. Anthropic and Google run comparable human review programmes, and so does essentially every company that has had to teach a model what a good answer looks like.

What is behind this

Models do not learn taste from more data alone. Someone has to look at two possible answers and say which is better, and that someone, so far, is a person. The technique has a name, reinforcement learning from human feedback, and it is a large part of why modern chatbots feel helpful rather than merely fluent. The awkward part is that real conversations are the most valuable training material precisely because they are real: messy, specific, occasionally desperate. A model tuned only on polite artificial questions gets good at polite artificial questions. So the industry uses the real thing, discloses it in the privacy policy, and relies on the fact that almost nobody reads privacy policies.

What this means for you: Assume anything you type into a general purpose chatbot could be read by a human being, and decide what to type on that basis. That is not paranoia, it is just the accurate mental model. If you want to reduce it, the control exists: in ChatGPT, open Settings, then Data Controls, and turn off the option about improving the model for everyone, which stops your future chats being used this way. Similar switches exist in Claude and Gemini, and business or enterprise plans generally exclude your data by default. For the genuinely sensitive things, health details, client documents, anything under an obligation of confidentiality, the better habit is to strip names and numbers before pasting, or use a model running on your own machine. And if you have a company: this is the reason free chatbot accounts and customer data do not belong in the same sentence.

Sources

Source: https://www.404media.co/inside-project-lily-the-humans-reading-your-chatgpt-chats/

Next story

Nvidia's New Rack Is 10x, 30x or 67x Better, Depending Entirely on Who Is Counting

Fresh numbers for the Vera Rubin NVL72 landed this week from Nvidia and from SemiAnalysis, and the multipliers range wildly. The spread is not dishonesty, it is a lesson in how efficiency claims are built.

Three rulers of different lengths and scales propped against the same tall block