← News

OpenAI’s Project Lily: contractors read real ChatGPT chats to train replies

404 Media, citing internal documents, Slack channels, real prompts, and a worker, reports that OpenAI’s Project Lily has hundreds of contractors reading anonymized ChatGPT user conversations to summarize intent and score model replies — OpenAI says a Privacy Filter strips personal data first but can miss details, the “improve the model for everyone” setting is on by default for free, Plus, and Pro, and Anthropic told 404 it also uses human review when users opt in.

People treat ChatGPT like a private therapist or work assistant. Paid humans score real chats, the Privacy Filter can miss details, and the default for free, Plus, and Pro is on — turn “improve the model for everyone” off if you do not want that.

On Monday 14 Sep 2026, 404 Media’s Joseph Cox reported that OpenAI is hiring hundreds of contractors who read a stream of real ChatGPT user prompts — sometimes including sensitive personal information — to improve model responses. Cox cites internal documents, Slack channels, real prompts, and a worker. That 404 Media investigation is the filing event. This is original reporting with company comments to the reporter. It is not an OpenAI newsroom announcement, and it is not a claim that every free-user chat is read by a human.

The program is codenamed Project Lily in internal material 404 Media reviewed. Project Lily, in plain English, is OpenAI’s internal name for this human-rating work. Contractors work in three stages, the report says: read the prompt, summarize what the user is asking, then rate and critique ChatGPT-generated replies. File the name and the three-stage loop as 404 Media’s. This desk did not see the dashboard.

Reviewers do not see ChatGPT usernames, 404 Media says. Some tasks include a user memories summary — a short note of what that person has used ChatGPT for, which can carry personal context such as where they may live. OpenAI told 404 it runs conversations through a Privacy Filter model first. Privacy Filter, in plain English, is OpenAI’s tool meant to scrub personal details before a human sees a chat. OpenAI also told 404 that Privacy Filter can miss uncommon identifiers. File the no-username line as 404 Media’s and the filter caveat as OpenAI’s, via 404. This desk is not saying the filter always fails.

OpenAI told 404 that chats are not used to improve models if users turn off “improve the model for everyone.” That setting, in plain English, is the ChatGPT switch that lets OpenAI use your chats for training and review. It is on by default for free, Plus, and Pro plans and applies to new conversations, not retroactively, OpenAI told 404. Enterprise, Business, and Edu have it off by default. File the defaults as OpenAI’s, via 404. This is not a claim that turning it off erases old chats already sent for review.

After 404 contacted OpenAI, the company updated help-page detail on the opt-out, the report says. 404 says the page still does not clearly say humans may read prompts for model improvement. File that as 404 Media’s read of the page after the update. This desk did not treat a help-page edit as a newsroom primary.

A North America-based worker told 404 they are paid more than $50 an hour via Crossing Hurdles and Mercor — two staffing firms that place people on AI-training gigs. The work felt very rote, the worker said, and guidelines change often. File the pay and the rote line as one worker, via 404. This is not a headcount beyond “hundreds,” and it is not a claim that every reviewer is paid that rate.

Anthropic told 404 it also uses human review to improve Claude when users enable “Help improve our AI models,” and that it de-identifies conversations first. Claude is Anthropic’s chatbot. File that as Anthropic’s comment to 404. This is not a second originating investigation, and it is not a claim that Anthropic runs Project Lily.

REPORTED here: 404 Media’s 14 Sep 2026 investigation by Joseph Cox, plus OpenAI and Anthropic comments to that reporter. No OpenAI newsroom primary the same day. NOT claimed: that OpenAI sells chats, that every free-user chat is read by a human, a contractor headcount beyond hundreds, that Privacy Filter always fails, or that TIME’s Kenya labeling story is today’s news. Distinct from the already-filed nvidia-palantir-curb-anthropic-fable, mai-code-of-conduct, altman-pacing-not-stopping, and anthropic-openai-google-standards-body-talks — those are buyer retention, a Microsoft rulebook, and the public pacing track; this filing is who reads consumer ChatGPT chats.

RELATED

ONLINE

article thread

guidelines

warming…

warming…

Sources