AI Training Data Startup Micro1 Hit $500 Million Gross Run Rate in Eight Months

Demand for human-verified AI training data is minting fast-growing startups. Micro1's revenue explosion shows the market is big enough for several players, and the ethical debates are heating up too.

AI2Day Newsdesk3 min read
Photoreal news-editorial 16:9 image of a server room corridor with rows of glowing rack-mounted computers, cool blue and white lighting, shallow depth of field
Share

Key points

  • Micro1, a four-year-old data-labeling startup, grew its gross annual run rate from $100 million to $500 million in eight months as of mid-2025.
  • The company retains roughly 60 to 70 percent of that figure, putting its net annual run rate between $150 million and $200 million.
  • Competitor Mercor hit $2 billion in gross annualized revenue this summer; Handshake reached $1 billion earlier this year.
  • Micro1 raised its Series A at a $500 million valuation in September 2024 and may have recently closed another round at a significantly higher valuation.
  • Founder Ali Ansari publicly stated the company does not sell training data to Chinese AI developers.

AI companies need massive amounts of carefully checked, human-verified information to train their models. Gathering and labeling that information is a booming business, and a startup called Micro1 is one of the fastest-growing players in it.

First reported by TechCrunch, Micro1 grew its gross annual run rate, the amount of money it would earn in a year at its current pace, from $100 million to $500 million in just eight months. That is a fivefold jump in under a year.

How does Micro1 actually make money?

The company connects AI labs and corporations with contract experts: doctors, lawyers, scientists, and engineers who check and rate AI outputs, helping models learn what good answers look like. This process is called data labeling.

Micro1 keeps 60 to 70 percent of the revenue it collects, passing the rest to contractors. Its net annual run rate sits somewhere between $150 million and $200 million.

The startup is also building a second, higher-margin business around synthetic data, information generated automatically rather than by humans. Some of this data can be sold to multiple clients. When the same dataset ships to several buyers, the cost of creating it is shared, so margins on those "off-the-shelf" datasets can reach 80 to 90 percent, according to a person familiar with the company's finances.

One example: automated written descriptions of video content, created without a human watching every clip.

Does it matter who buys this data?

Yes, and it is already controversial. Critics have argued that selling the same datasets to Chinese AI developers helps those developers build models that rival the best American ones.

Micro1's founder, Ali Ansari, addressed this directly last month, saying the company does not sell its data to Chinese model makers. "Some human data companies work with foreign adversaries, and the results show today in Kimi K3," he posted on X, referring to a Chinese AI model. "We believe it's shameful to claim American AI dominance desires while selling millions worth of data to countries that we are in adversarial competition with."

What happens next?

Micro1 is not the only company racing to supply training data. Competitor Mercor hit $2 billion in gross annualized revenue this summer; Handshake reached $1 billion earlier this year. The market appears large enough for all of them.

Some researchers now suggest that future AI spending on data could eventually rival what the industry spends on compute, the chips and computing power that run AI systems. If that prediction holds, companies like Micro1 are positioned well.

Micro1 raised its Series A at a $500 million valuation in September 2024. The company may have recently closed another funding round at a significantly higher valuation, though details have not been confirmed. Micro1 did not respond to a request for comment.

© 2026 AI2Day