Simulated Agent Systems Generate Coherent Enterprise Data at Scale
agents cohere training
| Source: HF Papers | Original article
Researchers propose a scalable simulation method that generates coherent enterprise tabular data, bypassing legal and business limits on training tool‑calling agents.
Researchers from SAP and partner institutions have unveiled a new method for producing enterprise‑grade data without ever touching real customer records. The approach, dubbed **Synthesis Through Simulation (STS)**, was posted on arXiv and described as a **schema‑free** data‑synthesis paradigm in which a large‑language‑model (LLM) agent interacts with policy‑enforcing APIs inside a simulated enterprise environment. By issuing tool‑calling operations that respect built‑in constraints, the agent can populate tables, trigger workflows and generate multi‑system records that mirror the statistical properties of genuine business data.
The paper highlights the **Generalist Populator (GP)**, a domain‑agnostic agent that achieved an average marginal fidelity of 0.88 while satisfying 100 % of the imposed constraints. The generated datasets are reproducible, openly available through a hosted service (https://console.era.eon.io) and as container images for offline use, allowing researchers and developers to train and evaluate tool‑calling agents at scale.
Why it matters: Tool‑calling agents are becoming central to enterprise AI, yet their development is hamstrung by legal and privacy restrictions on real data and database schemas. STS offers a practical workaround, delivering coherent, policy‑compliant synthetic data that can be used for supervised learning, benchmarking and stress‑testing without exposing sensitive information. The schema‑free nature also sidesteps the costly effort of hand‑crafting data models for each new system.
What to watch next: The community will be looking for early adopters integrating STS‑generated data into their training pipelines, as well as any performance gaps that emerge when agents move from simulated to production environments. Follow‑up studies may expand the method to more complex, cross‑system workflows, and the availability of containerised environments could spur open‑source tooling around synthetic enterprise data. As we reported on 2026‑10‑09, this work builds on the growing interest in scaling LLM‑agent learning through synthetic supervision, and its real‑world impact will become clearer as enterprises begin to experiment with the hosted service.
Sources
Back to AIPULSEN