Scalable Agent Simulations Generate Coherent Enterprise Data
agents cohere training
| Source: ArXiv | Original article
A new study introduces a simulation-driven method to synthesize coherent enterprise tabular data, enabling scalable training and evaluation of tool‑calling agents despite legal and data constraints.
A new arXiv pre‑print (arXiv:2610.10549v1) introduces **Synthesis Through Simulation (STS)**, a schema‑free approach for generating enterprise‑scale tabular data. The method lets a large language model (LLM) agent produce data by invoking policy‑enforcing APIs inside a simulated business environment, rather than by pulling from real customer records. The resulting datasets are reproducible across multiple systems, come with explicit statistical assumptions, and are offered both as a hosted service (https://console.era.eon.io) and as container images for offline use.
The announcement tackles a long‑standing bottleneck for tool‑calling agents that are increasingly central to enterprise AI. Training and evaluating such agents has been hampered by legal and privacy constraints that limit access to genuine enterprise databases and schemas. By synthesising coherent, policy‑compliant data in a sandboxed setting, STS provides a scalable source of supervision that sidesteps those restrictions while preserving the structural complexity of real‑world tables.
The development builds on the broader trend of **agentic data synthesis**, which we have covered previously as an emerging paradigm that uses autonomous agents, multi‑agent MDPs and recursive loops to create, curate and verify complex datasets. STS extends that concept to the enterprise domain, offering a concrete pipeline that could accelerate instruction‑tuning, alignment and multi‑system reasoning for LLM agents.
What to watch next is the uptake of STS by AI teams building enterprise assistants and the emergence of benchmark suites that benchmark tool‑calling performance on the synthetic data. Further research will likely explore how the schema‑free paradigm integrates with multi‑agent world models and whether the approach can be extended to other data modalities. The community will also be keen to see open‑source releases of the containerised environments, which could become a standard testbed for safe, large‑scale agent training.
Sources
Back to AIPULSEN