On September 14, 404 Media named an internal codename: Project Lily. Under it, OpenAI has hired hundreds of contractors to read real conversations between real people and ChatGPT.
Per internal documents the reporters obtained, the task is two steps. Summarize what the user was trying to ask. Then score four generated answers on a scale of one to seven.
How the work gets routed
A firm called Crossing Hurdles does the recruiting. Payment runs through Mercor, a talent marketplace. One person who worked the queue puts the rate above $50 an hour.
OpenAI draws a line on what reviewers see. Usernames are stripped. The company says it tries to remove personal information before a conversation reaches a rater, and concedes that sensitive detail still gets through. Asked whether users understand their chats are being read, a source said no, adding that nobody pictures a contractor somewhere analyzing these conversations.
The scale numbers are already public. ChatGPT has more than 900M users, and training is on by default for consumer plans.
OpenAI is not alone in this. Anthropic confirmed to 404 Media that it also uses human review to improve its models, and Google runs comparable programs.
Same job, and the price moved an order of magnitude
The role is not new. The price is.
On April 20 we covered Sama shutting its Kenya operation and putting 1,108 people out of work at once. That workforce labeled, classified and moderated AI output, and Nairobi paid it a few dollars an hour. On August 28 we covered EXL buying iMerit’s 68,000 evaluation seats for $310M, because evaluation capacity had become expensive enough to acquire wholesale rather than hire.
Project Lily is the newest point on that curve and it sits a long way up. The described activity is the same — look at model output, assign a score — but Nairobi paid single-digit dollars and Mercor is paying north of $50.
What changed is what gets judged. Content moderation asks whether a passage crossed a line, the rule is written in a handbook, and any careful reader can apply it. Ranking four answers is a different job. To know which answer is better you have to understand the thing the user was doing: redlining a contract, wiring up a service, working out a tax position. That moves the bar from literacy and patience to domain fluency.
Put plainly, Project Lily is not hiring annotators. It is hiring practitioners, on contractor terms.
This is what a job AI created looks like
The recurring question of the past two years is what work AI is producing to offset what it is removing.
Project Lily is a specific enough answer to hold up against others. The role is contract, has no level, has no ladder, is paid through a third-party marketplace, and opens and closes with the project. It is also the best-paid version of this work anyone has disclosed. At $50 an hour, it beats the hourly equivalent of plenty of salaried white-collar jobs in the U.S.
Both halves are true, which is the point. This is not a bad job. It is a good job with no shape. You are paid like a domain expert and contracted like a gig worker, and the project can close the moment a model version stops needing human scores in your field.
The thing to watch is verticals. On September 10 we covered OpenAI packaging junior financial-analyst work into a product. For a model to hold up in finance, law or medicine, somebody who knows those fields has to rank its output first. So expect the same structure to replicate industry by industry: domain professionals, hourly, contract, marketplace-paid.
For anyone carrying real industry experience, there is a concrete question underneath the rate. Ranking a model’s answers either rents that experience out or sells it outright. The hourly number looks identical either way.
Sources