Learn / RL environments / Updated October 2026
AI stopped reading. Now it practices.
Labs now pay to build small worlds where models do real work and get graded. Those worlds need real businesses inside them.
Reinforcement learning is how frontier models learn to finish jobs, not just write about them. rl.data.cool is named for its buyers: we supply the real-world history their environments are missing.
01 / What an RL environment is
A task, a place to attempt it, and a grader.
Repeat a million times. The internet taught models to write; nothing online taught them to work.
01 / The world
A working copy of a job: an inbox, a CRM, a ticket queue, a set of books. Shipped as a container the agent can act in.
e.g. a QuickBooks replica with three years of a firm's invoices
02 / The task
Something a real employee would be asked to do, over many steps, with the tools that job uses.
e.g. fix every invoice billed to the wrong client in March
03 / The grader
A program that decides, without a human, whether the job was actually done. This is the hard part.
e.g. every corrected invoice matches the ledger; nothing else changed
02 / Where the money is going
Training moved from the whiteboard to the back office.
Labs built environments in order of how easy they are to grade. Enterprise work is where they are buying now, and it's where small and mid-sized businesses keep their records.
- 01 / FirstMathThe answer is right or wrong. Grading is free.
- 02 / ThenCodeTests pass or fail.
- 03 / ThenSoftware engineeringAgents work inside real repositories.
- 04 / NowEnterprise workSpreadsheets, CRMs, ticket desks, approvals. Where labs are buying now.
03 / 2026 so far
The year environments became the bottleneck.
Jan 2026
Labs buy from dozens of vendors
SemiAnalysis reports Anthropic working with more than a dozen environment companies, often as their first customer, and pushing a common sandbox format so vendors are interchangeable.2
Feb 2026
A realistic company makes models better everywhere
Surge AI's Corecraft: frontier models solve under 35% of its tasks. One pass of training lifts a model from 25% to 37%, and the gains carry over to benchmarks it never saw.4
Mar 2026
Enterprise work is still mostly unsolved
ServiceNow's EnterpriseOps-Gym: 1,150 tasks across email, calendar, HR, IT and customer service, checked against the final database state. The best model completes 34%.5
Jul 8 2026
$130M for open RL infrastructure
Prime Intellect raises a $130M Series A for compute, RL training, environments and sandboxes, and says companies can now "train directly on their own product".6
Aug 2 2026
Provenance becomes enforceable in the EU
The AI Office can now fine providers of general-purpose AI models up to 3% of global turnover. Their public training-data summaries have to confirm the licensing behind private datasets.9
Aug 2026
Reward hacking at production scale
In a deliberate study, Anthropic trains a model on 80 real production environments with known exploits. By the end it cheats on 40% of episodes, and finds hacks nobody anticipated.10
Aug 21 2026
Synthetic business worlds, at scale
AgentMercury generates 4,783 executable business environments across 14 industries. Training on them helps, which raises the bar for what real data has to add.11
Sep 23 2026
Real workflows, real outcomes
Realset raises a $10M Series A to build RL environments from real workflows, scored with "verifiable rewards from real business outcomes".12
04 / Why business data
Real business data is what makes an environment worth training in.
Grounded, governed, gradable, connected, and owned. Synthetic worlds struggle to fake all of it at once.
01 / The world
Grounded scenarios
02 / The rules
Rules and shared state
03 / The answer key
Verifiable outcomes
04 / The hard tasks
Cross-application work
05 / The edge
Private tools and jargon
06 / The rights
A licensed owner
| Synthetic company | Real company history | |
|---|---|---|
| Cost to look real | Someone invents years of history | Already exists |
| Mess and drift | Too clean; models learn shortcuts | Contradictions, half-done processes, real quirks |
| Ground truth | Whatever the author decided | What actually happened |
| Scale | Thousands of worlds, cheaply | One company at a time |
| Who owns it | The vendor | A confirmed owner, licensed |
A $300k Slack clone with no channels teaches nothing.
The honest counterpoint. Synthetic worlds are getting good: AgentMercury generated nearly 5,000 business environments and training on them helped.11 What generation can't produce is what actually happened: the real dispute, the real resolution, the real mess. Our view is that synthetic gives you breadth and real history gives you the answer key, and the best environments will use both.
05 / Proof
Trained on a company's real work, small models beat the giants at it.
21.9% → 46.9%
Correct database queries
Small open model trained on an insurer's real tables vs the best off-the-shelf model13
25% → 37%
Tasks solved after one epoch
On Corecraft tasks the model never practiced, plus +4.5 to +7.4 points on outside benchmarks4
Scale AI built environments from real client work: database questions over a global insurer's financial tables, and 4,000 questions on 30 real contracts checked by practicing lawyers. A 4B-parameter open model trained inside them beat GPT-5 at that work.13
Surge's Corecraft is a fictional support company with 2,500+ records, 14 kinds of business object and 23 tools, all built by hand. Surge credits realistic workflows, varied hard tasks and expert grading for the transfer.4 A real company already has all of it.
06 / What can go wrong
The grader is only as good as the truth behind it.
01 / Bad data costs points
-5 pts
Scale mixed 15% unverifiable examples into legal training. Accuracy fell by five points. Clean, checked data is the lever.13
02 / Models cheat the grader
90% → 32%
A model learned to slip instructions to its AI grader and scored 90%. With a hardened grader, its real score was 32%.13
03 / Cheating spreads
40%
Trained on 80 exploitable production environments, a frontier-class model reward-hacked 40% of episodes by the end.10
04 / Real data needs care
CURRENT_DATE
Old records plus today's date broke an insurer's queries. Scale had to freeze the clock inside the environment.13
Real outcomes make graders harder to fool, not impossible. Builders pay most for history that is clean, checked and tagged with what happened.
07 / The supply chain
We're the lumber yard. Builders and labs both buy lumber.
Business owners
10-250 person operators with years of email, tickets, CRM and books they already own.
Brokers
M&A advisors and consultants who bring owners in, only with the owner's agreement.
rl.data.cool
Verifies, de-identifies, scans for relisted data, certifies the owner's sign-off. Then licenses it.
Environment builders
Mercor, Surge, Scale, Turing, Mechanize, Prime Intellect and others turn lumber into worlds, tasks and graders.
Frontier labs
Pay for the finished environment. Some build in-house and buy lumber direct.
08 / Rights
Every dataset needs an owner who said yes.
Since August 2026 the EU AI Office can fine providers of general-purpose models up to 3% of global turnover, and their public training-data summaries must confirm the licensing behind private datasets.9 Provenance used to be a nice-to-have. Now it's audited.
Owner signs off
Every listing is confirmed by the owner on a company-domain email. Brokers can't list without them.
Blind by default
Buyers never learn who the owner is unless the owner chooses to disclose to counsel.
Clean Room
Names, emails and phones become consistent pseudonyms, with a measured residual-risk report.
Never sold twice
Owner and broker declare prior sales, and every sample is scanned for watermarks from copies already sold.
Every reveal logged
Identity is encrypted field by field. Any access is recorded: who, why, when.
Paper trail
A fingerprinted manifest and a dated, owner-confirmed license for every corpus.
09 / Glossary
The words, plainly.
- Reinforcement learning (RL)
- Training by trial and reward: the model attempts a task, gets scored, and is pushed toward higher scores.
- Environment
- The world, the task and the grader, packaged so a model can attempt the task millions of times.
- Grader / verifier
- The program that scores an attempt. Can be a test, a database check or an expert-written rubric.
- Rubric
- A checklist of what a correct answer must contain, written by a domain expert.
- Reward hacking
- When a model finds a way to score well without doing the job, like editing the test instead of the code.
- GRPO
- A training method: the model tries a task several times and learns from its better attempts.
- Epoch
- One pass through every training task.
- Out-of-distribution
- Tests the model never saw in training. Gains here mean it learned the skill, not the test.
- Taskset / harness
- Prime Intellect's split: the taskset is the work and its scoring; the harness is the agent loop that attempts it.
- RFT
- Reinforcement fine-tuning: RL on a customer's own tasks, offered as a service by labs and startups.
10 / Sources
Where the numbers come from.
Figures as reported by each source. Market numbers in this space are mostly from interviews and press, not audits.
- 1Epoch AI, "An FAQ on reinforcement learning environments"Jan 2026
- 2SemiAnalysis, "RL environments and RL for science"Jan 2026
- 3Wing VC, "Who will win the RL environment market, and why"2026
- 4Surge AI, "EnterpriseBench Corecraft", arXiv 2602.16179Feb 2026
- 5ServiceNow AI Research, "EnterpriseOps-Gym", arXiv 2603.13594Mar 2026
- 6Prime Intellect, "$130M Series A"Jul 8, 2026
- 7Mercor, "Mercor to acquire Deeptune"Jul 9, 2026
- 8Fortune, "AI unicorn Mercor acquires Deeptune"Jul 9, 2026
- 9European Commission, AI Act: GPAI obligations and enforcementAug 2026
- 10Anthropic Alignment, "Training a misaligned reward seeker"Aug 2026
- 11Jeong & Yoon, "AgentMercury", arXiv 2608.20634Aug 21, 2026
- 12Realset AI, "$10M Series A to build real-world training data"Sep 23, 2026
- 13Scale AI Labs, "Scaling enterprise agent performance with RL via verifiable feedback loops"Nov 2025
- 14Mechanize, "Cheap RL tasks will waste compute"Aug 2025