Learn / RL environments / Updated October 2026

AI stopped reading. Now it practices.

Labs now pay to build small worlds where models do real work and get graded. Those worlds need real businesses inside them.

Reinforcement learning is how frontier models learn to finish jobs, not just write about them. rl.data.cool is named for its buyers: we supply the real-world history their environments are missing.

$1B+

One lab's discussed annual spend on RL environments1

$2B

Mercor's annual run-rate in June 2026, up from $1B a year earlier8

$200-2k

Typical price of a single training task1

4-5x

Premium for the same environment sold exclusively1

01 / What an RL environment is

A task, a place to attempt it, and a grader.

Repeat a million times. The internet taught models to write; nothing online taught them to work.

01 / The world

A working copy of a job: an inbox, a CRM, a ticket queue, a set of books. Shipped as a container the agent can act in.

e.g. a QuickBooks replica with three years of a firm's invoices

02 / The task

Something a real employee would be asked to do, over many steps, with the tools that job uses.

e.g. fix every invoice billed to the wrong client in March

03 / The grader

A program that decides, without a human, whether the job was actually done. This is the hard part.

e.g. every corrected invoice matches the ledger; nothing else changed

Act

The model clicks, types, queries, sends.

Score

The grader checks the end state.

Update

Training pushes toward higher scores.

A task burns about $2,400 of compute over its life. A bad input is the most expensive line in the budget.14,1

02 / Where the money is going

Training moved from the whiteboard to the back office.

Labs built environments in order of how easy they are to grade. Enterprise work is where they are buying now, and it's where small and mid-sized businesses keep their records.

  1. 01 / FirstMathThe answer is right or wrong. Grading is free.
  2. 02 / ThenCodeTests pass or fail.
  3. 03 / ThenSoftware engineeringAgents work inside real repositories.
  4. 04 / NowEnterprise workSpreadsheets, CRMs, ticket desks, approvals. Where labs are buying now.

Contracts

Six to seven figures a quarter at the top; $300-500k a quarter for smaller labs.1

Replicas

About $20k for a website replica, about $300k for a high-fidelity clone of a complex product.1

Market shape

Roughly 20 early-stage builders today, expected to narrow to 3-5 leaders by 2030.3

03 / 2026 so far

The year environments became the bottleneck.

  1. Jan 2026

    Labs buy from dozens of vendors

    SemiAnalysis reports Anthropic working with more than a dozen environment companies, often as their first customer, and pushing a common sandbox format so vendors are interchangeable.2

  2. Feb 2026

    A realistic company makes models better everywhere

    Surge AI's Corecraft: frontier models solve under 35% of its tasks. One pass of training lifts a model from 25% to 37%, and the gains carry over to benchmarks it never saw.4

  3. Mar 2026

    Enterprise work is still mostly unsolved

    ServiceNow's EnterpriseOps-Gym: 1,150 tasks across email, calendar, HR, IT and customer service, checked against the final database state. The best model completes 34%.5

  4. Jul 8 2026

    $130M for open RL infrastructure

    Prime Intellect raises a $130M Series A for compute, RL training, environments and sandboxes, and says companies can now "train directly on their own product".6

  5. Jul 9 2026

    "The constraint has shifted to the environments"

    Mercor buys Deeptune, which recreated hundreds of enterprise apps, from spreadsheets to Salesforce, for frontier labs. Mercor's CEO: the bottleneck is now "the places where models practice the work".7,8

  6. Aug 2 2026

    Provenance becomes enforceable in the EU

    The AI Office can now fine providers of general-purpose AI models up to 3% of global turnover. Their public training-data summaries have to confirm the licensing behind private datasets.9

  7. Aug 2026

    Reward hacking at production scale

    In a deliberate study, Anthropic trains a model on 80 real production environments with known exploits. By the end it cheats on 40% of episodes, and finds hacks nobody anticipated.10

  8. Aug 21 2026

    Synthetic business worlds, at scale

    AgentMercury generates 4,783 executable business environments across 14 industries. Training on them helps, which raises the bar for what real data has to add.11

  9. Sep 23 2026

    Real workflows, real outcomes

    Realset raises a $10M Series A to build RL environments from real workflows, scored with "verifiable rewards from real business outcomes".12

04 / Why business data

Real business data is what makes an environment worth training in.

Grounded, governed, gradable, connected, and owned. Synthetic worlds struggle to fake all of it at once.

01 / The world

Grounded scenarios

Real product catalogs, customer threads and financial pipelines, not templates. The agent learns the job as it's actually done.

02 / The rules

Rules and shared state

Check the record, follow the approval policy, update every system consistently. Real data carries the constraints with it.

03 / The answer key

Verifiable outcomes

History has endings: the invoice that was disputed, how the ticket closed, which lead signed. Deterministic answers, no human grader.

04 / The hard tasks

Cross-application work

Real companies run on email, a CRM, a ticket desk and a database at once. Agents get trained and scored on switching between them.

05 / The edge

Private tools and jargon

Models fail on in-house tools and vocabulary out of the box. Environments built from real business data close that gap.

06 / The rights

A licensed owner

A real company's history comes with someone who can say yes. Synthetic data has no owner to ask; scraped data has one who never agreed.
Synthetic companyReal company history
Cost to look realSomeone invents years of historyAlready exists
Mess and driftToo clean; models learn shortcutsContradictions, half-done processes, real quirks
Ground truthWhatever the author decidedWhat actually happened
ScaleThousands of worlds, cheaplyOne company at a time
Who owns itThe vendorA confirmed owner, licensed

A $300k Slack clone with no channels teaches nothing.

The honest counterpoint. Synthetic worlds are getting good: AgentMercury generated nearly 5,000 business environments and training on them helped.11 What generation can't produce is what actually happened: the real dispute, the real resolution, the real mess. Our view is that synthetic gives you breadth and real history gives you the answer key, and the best environments will use both.

05 / Proof

Trained on a company's real work, small models beat the giants at it.

21.9% → 46.9%

Correct database queries

Small open model trained on an insurer's real tables vs the best off-the-shelf model13

49% → 21%

Made-up answers

On a law firm's real contracts, when the answer wasn't in the document13

25% → 37%

Tasks solved after one epoch

On Corecraft tasks the model never practiced, plus +4.5 to +7.4 points on outside benchmarks4

34%

Best model on enterprise ops

1,150 stateful tasks across eight business domains5

Scale AI built environments from real client work: database questions over a global insurer's financial tables, and 4,000 questions on 30 real contracts checked by practicing lawyers. A 4B-parameter open model trained inside them beat GPT-5 at that work.13

Surge's Corecraft is a fictional support company with 2,500+ records, 14 kinds of business object and 23 tools, all built by hand. Surge credits realistic workflows, varied hard tasks and expert grading for the transfer.4 A real company already has all of it.

06 / What can go wrong

The grader is only as good as the truth behind it.

01 / Bad data costs points

-5 pts

Scale mixed 15% unverifiable examples into legal training. Accuracy fell by five points. Clean, checked data is the lever.13

02 / Models cheat the grader

90% → 32%

A model learned to slip instructions to its AI grader and scored 90%. With a hardened grader, its real score was 32%.13

03 / Cheating spreads

40%

Trained on 80 exploitable production environments, a frontier-class model reward-hacked 40% of episodes by the end.10

04 / Real data needs care

CURRENT_DATE

Old records plus today's date broke an insurer's queries. Scale had to freeze the clock inside the environment.13

Real outcomes make graders harder to fool, not impossible. Builders pay most for history that is clean, checked and tagged with what happened.

07 / The supply chain

We're the lumber yard. Builders and labs both buy lumber.

01 / The forest

Business owners

10-250 person operators with years of email, tickets, CRM and books they already own.

02 / The loggers

Brokers

M&A advisors and consultants who bring owners in, only with the owner's agreement.

03 / The lumber yardUs

rl.data.cool

Verifies, de-identifies, scans for relisted data, certifies the owner's sign-off. Then licenses it.

04 / The builders

Environment builders

Mercor, Surge, Scale, Turing, Mechanize, Prime Intellect and others turn lumber into worlds, tasks and graders.

05 / The homeowners

Frontier labs

Pay for the finished environment. Some build in-house and buy lumber direct.

08 / Rights

Every dataset needs an owner who said yes.

Since August 2026 the EU AI Office can fine providers of general-purpose models up to 3% of global turnover, and their public training-data summaries must confirm the licensing behind private datasets.9 Provenance used to be a nice-to-have. Now it's audited.

Owner signs off

Every listing is confirmed by the owner on a company-domain email. Brokers can't list without them.

Blind by default

Buyers never learn who the owner is unless the owner chooses to disclose to counsel.

Clean Room

Names, emails and phones become consistent pseudonyms, with a measured residual-risk report.

Never sold twice

Owner and broker declare prior sales, and every sample is scanned for watermarks from copies already sold.

Every reveal logged

Identity is encrypted field by field. Any access is recorded: who, why, when.

Paper trail

A fingerprinted manifest and a dated, owner-confirmed license for every corpus.

09 / Glossary

The words, plainly.

Reinforcement learning (RL)
Training by trial and reward: the model attempts a task, gets scored, and is pushed toward higher scores.
Environment
The world, the task and the grader, packaged so a model can attempt the task millions of times.
Grader / verifier
The program that scores an attempt. Can be a test, a database check or an expert-written rubric.
Rubric
A checklist of what a correct answer must contain, written by a domain expert.
Reward hacking
When a model finds a way to score well without doing the job, like editing the test instead of the code.
GRPO
A training method: the model tries a task several times and learns from its better attempts.
Epoch
One pass through every training task.
Out-of-distribution
Tests the model never saw in training. Gains here mean it learned the skill, not the test.
Taskset / harness
Prime Intellect's split: the taskset is the work and its scoring; the harness is the agent loop that attempts it.
RFT
Reinforcement fine-tuning: RL on a customer's own tasks, offered as a service by labs and startups.

10 / Sources

Where the numbers come from.

Figures as reported by each source. Market numbers in this space are mostly from interviews and press, not audits.

  1. 1Epoch AI, "An FAQ on reinforcement learning environments"Jan 2026
  2. 2SemiAnalysis, "RL environments and RL for science"Jan 2026
  3. 3Wing VC, "Who will win the RL environment market, and why"2026
  4. 4Surge AI, "EnterpriseBench Corecraft", arXiv 2602.16179Feb 2026
  5. 5ServiceNow AI Research, "EnterpriseOps-Gym", arXiv 2603.13594Mar 2026
  6. 6Prime Intellect, "$130M Series A"Jul 8, 2026
  7. 7Mercor, "Mercor to acquire Deeptune"Jul 9, 2026
  8. 8Fortune, "AI unicorn Mercor acquires Deeptune"Jul 9, 2026
  9. 9European Commission, AI Act: GPAI obligations and enforcementAug 2026
  10. 10Anthropic Alignment, "Training a misaligned reward seeker"Aug 2026
  11. 11Jeong & Yoon, "AgentMercury", arXiv 2608.20634Aug 21, 2026
  12. 12Realset AI, "$10M Series A to build real-world training data"Sep 23, 2026
  13. 13Scale AI Labs, "Scaling enterprise agent performance with RL via verifiable feedback loops"Nov 2025
  14. 14Mechanize, "Cheap RL tasks will waste compute"Aug 2025