
23 Sep 2026
Realset AI and Flatkey raise $10M Series A for real-world AI training data
Realset AI and Flatkey announced a $10 million Series A to expand capture of expert human demonstrations in real workplaces, build reinforcement-learning environments from real workflows, and publish open benchmarks for embodied agents, with a household manipulation bench due in the fourth quarter of 2026.
FINANCE desk — labs already ate the easy internet text; the next training edge is paying skilled people to do real jobs on camera so robots and agents learn what the web never recorded.
Why the release says the round exists. The section title is “Why It Matters: The Internet Is Exhausted, the Physical World Is Not.” Frontier labs and robotics companies, it says, have largely consumed the text available on the internet. The next gains come from data about people doing real tasks in real places: how a worker folds laundry, loads a dishwasher, packs an order, or handles a customer return. That data does not exist on the web, and simulation does not reproduce it. Simulation, here, is a computer copy of a room. These sentences are the release’s argument. They are not a measurement this desk made of how much text is left online.
What the company says it does instead. Every dataset starts with a real person doing a real task in a real place. Realset uses expert demonstrators rather than crowd annotators, real environments rather than simulation, and delivers de-identified data to US-hosted cloud buckets, with quality measured at every step of a single pipeline. A crowd annotator is a person paid to label someone else’s clip. An expert demonstrator is the skilled person doing the task on camera. De-identified means personal details are stripped before the file is delivered. A cloud bucket is a storage folder on a server. US-hosted means that folder sits in the United States. A pipeline is the path from the camera to the delivered file. These lines are the release’s. This desk did not open a bucket.
Three product lines, as the release names them. Realset Body is embodied data for physical policy models: cameras on skilled workers doing manipulation tasks in homes, kitchens, warehouses, and light assembly lines, plus episodes where a person remotely drives a robot’s two hands. The files ship with a dense label on each action, for vision-language-action training. A vision-language-action model, shortened to VLA, is a robot model that sees, reads an instruction, and moves. Embodied, here, means the data is about a body in a room, not a chat transcript. Manipulation means hands moving objects. The release also says those remote-drive episodes include synchronized stereo video, a motion sensor, and action logs. Stereo video is two cameras, so the clip has depth. Those sensor lines are the release’s. This desk did not watch a tape.
Realset Field is training data for language models and agents, built as reinforcement-learning environments from real workflows. Reinforcement learning, shortened to RL, means the model practices inside an environment and gets a score for what it did. The environments mirror e-commerce operations, customer support, logistics dispatch, and manufacturing standard operating procedures. A standard operating procedure is the written way a shop says a job is done. Domain experts generate trajectories, preference pairs, and rubrics inside the environment, with rewards a buyer can check against real business outcomes, and with bilingual English and Chinese expert pools. A trajectory is the path of steps the person or the agent took. A preference pair is two answers, with a judgment of which is better. A rubric is a written score sheet. These lines are the release’s. The bilingual pool is Field’s sentence. It is not the six-language list below.
Realset Judge is expert evaluation of agents already in production. People who do the job the agent is replacing design the evaluation, diagnose failures, and keep watching. They then produce targeted fix-data for the top failure modes and re-score as models and prompts change. Fix-data means new examples aimed at the mistakes, not a fresh scrape of the web. The release says the same services are offered to data-integration and AI-solution companies that deploy agents for their own clients. These lines are the release’s. This desk did not sit in on a review.
All three run on the Realset Workspace. Domain experts in six languages — English, Simplified Chinese, Japanese, Korean, Spanish, and Arabic — answer real tasks, attach evidence and screen recordings, and pass an independent quality review before a record is approved. Every approved record exports as JSONL with provenance a customer can verify. JSONL is a text file with one structured record per line. Provenance is the trail of where that record came from and who checked it. Independent means the reviewer is not the person who made the record. These lines are the release’s.
The open tests, as the release states them. Realset publishes benchmarks so the field can measure whether a policy works outside the lab. A policy, here, is the model’s rule for what to do next. The Realset Household Manipulation Bench evaluates open-source VLA policies, including π0, OpenVLA, GR00T, and Octo, on folding, loading, sorting, and wiping in real kitchens and laundry rooms. The score is a success rate over three trials per task. Results are expected in the fourth quarter of 2026, October through December. A Light Assembly Bench, scored by the line workers who trained on the tasks, and a Commerce Ops Agent Bench for computer-use agents on real e-commerce seller operations, are planned. Planned means the release names them and does not give a date. Open-source, here, means the policies’ code is public. These lines are the release’s. This desk did not run the bench.
The founder, as a quote. Hunter Guo, founder of Realset AI, said: “The easy data is gone. What is left is the physical world, and you cannot scrape it.” He said you have to put a camera on a skilled person doing real work, structure what they did, and check it with people who know the job. He said that is a capture and quality problem, not a labeling problem, and it is the problem they are building a company around. That is his sentence on the release. A quote is not a second measurement. This desk did not interview him.
What the release says Flatkey is. Flatkey is an AI infrastructure platform that gives developers access to more than 100 official AI models and more than 1,000 AI tools through one key and one balance. One key means one credential instead of a separate login for each model. One balance means one prepaid account those calls draw from. More than 100 and more than 1,000 are the release’s counts. The page prints https://flatkey.ai. This desk did not count the catalog and did not open that site.
What the release says Realset is. A real-world data lab and a training-data provider for large language models and embodied AI. It captures expert human demonstrations, builds RL environments from real workflows, and evaluates AI agents with domain experts, for frontier labs, robotics companies, and data-integration and AI-solution providers. It is headquartered in San Jose, California. The page prints https://realset.ai. A large language model, shortened to LLM, is a model trained to continue text. An embodied agent is AI that acts in a physical room or a realistic workflow, not only in a chat box. These lines are the about box. They are not a customer list this desk checked.
Plain English for the rest of the card. $10 million is the Series A. 1:30 p.m. Eastern is the wire’s stamp. VLA is a model that sees, reads an instruction, and moves. An embodied agent acts in a room or a workflow, not only in chat. RL is practice inside an environment that scores the attempt. JSONL is one record per line. Provenance is the trail on that record. De-identified means personal details are stripped. Six languages are the Workspace list. English and Chinese are Field’s expert-pool sentence. Do not merge the two lists. The fourth quarter of 2026 is the window for the household bench results. Three trials is the scoring rule. π0, OpenVLA, GR00T, and Octo are the open-source policies the bench names. The release does not name who wrote the checks.
PRIMARY here: PR Newswire, source Realset AI, stamped Sep 23, 2026, 13:30 ET — Tier A PRIMARY, the companies’ own wire. The company homepage at realset.ai, read the same day, restates the Workspace languages, the JSONL export, and the household bench. It is not a second announcement. The $10 million Series A, the San Jose dateline, the three uses of proceeds, the internet-is-exhausted argument, the expert-demonstrator and de-identified US-hosted lines, Body, Field, and Judge, the six languages, the JSONL provenance line, the household bench and its four named policies, the three-trial score, the fourth-quarter window, the two planned benches, the Hunter Guo quote, the Flatkey one-key description, and the San Jose headquarters line are the wire’s. The homepage’s street address, its four-kind product list, and a gateway-trajectory product the wire does not name stay on the homepage. Do not merge that gateway product into Body, Field, or Judge. NOT claimed: the lead investor, a valuation, how the $10 million splits between Realset and Flatkey, that this desk counted Flatkey’s models or opened flatkey.ai, that the household bench has published scores, a stock tip, or investment advice. Distinct from the already-filed tekever-580m-series-d, bird-com-450m-agentic, ema-77m, and basecamp-research-140m.
RELATED
On 23 Sep 2026 Realset AI and Flatkey announced they raised $10 million in Series A funding. The record is the PR Newswire release. The stamp is Sep 23, 2026, 13:30 ET, which is 1:30 p.m. Eastern and 5:30 p.m. UTC. The dateline is San Jose, California. A Series A is an early growth round, the check that comes after a company has a product and wants to expand it. $10 million is the size of this round. The release says the money will expand Realset’s capture network of real workplaces and studio environments, grow its pool of expert demonstrators and domain experts, and support open benchmarks that measure whether AI policies work outside the lab. A capture network, here, is the set of real rooms and the people filmed in them. A domain expert is someone who already does the job. An open benchmark is a public test other labs can run, not a private score the company keeps. These lines are the release’s. This desk did not see the term sheet. The release does not name the investors.