AutoSynthData: Generating Training Data for Enterprise Agents

AI小蝌蚪AI 前沿📡 Hugging Face Blog2026-10-02100 阅读💛 35 收藏
Back to Articles

AutoSynthData: Generating Training Data for Enterprise Agents

Enterprise Article
Published October 2, 2026
Esakkivel Esakkiraja's avatar
Esakkivel Esakkiraja
esakkivel
ServiceNow-AI's avatar ServiceNow-AI
Shruthan Radhakrishna's avatar
Shruthan Radhakrishna
shruthan-r
ServiceNow-AI's avatar ServiceNow-AI
Denis Akhiyarov's avatar
Denis Akhiyarov
dtanow
ServiceNow-AI's avatar ServiceNow-AI
Sagar Davasam's avatar
Sagar Davasam
davasam
ServiceNow-AI's avatar ServiceNow-AI

AutoSynthData_thumbnail_1200x648 (2)

Enterprises need agents that work well in their own environments. The work they ask these agents to do is shaped by the systems they use, the rules they follow, and the state of their data. A model may be broadly capable and still struggle with a particular environment: a workflow it handles poorly, a combination of tools it misuses, or a constraint it fails to respect. Those are the weaknesses an enterprise needs to improve.

The difficulty is turning those weaknesses into training data. An individual failure tells us something, but training a model requires many new tasks that exercise the same capability in different situations. Those tasks must also be possible to complete in the environment, resemble work someone would actually request, and have a reliable way to check whether the agent succeeded.

At ServiceNow CoreAI, we built AutoSynthData to turn those capability gaps into training data. It uses a target model’s failures and a stronger teacher’s successes to decide what the model should learn next, then generates and validates new tasks that exercise those capabilities. As the model improves, the curriculum shifts toward what it still finds difficult. We illustrate the pipeline with EnterpriseOps Gym (Malay et al., 2026), using the released dataset. We begin by describing the environment an agent operates in and what makes a task useful for training.

What makes a useful agentic task?

An agentic environment defines the world in which an agent operates: the state it can observe and modify, the tools and APIs it can invoke, and the state transitions produced by its actions.

A task is instantiated within this environment. We use the following abstraction:

task = (system specification, user prompt, verifier)

System specification

The system specification defines the constraints under which the agent operates, including system instructions, environment policies, and, when applicable, task-specific initialization such as a seeded database state or a set of knowledge articles.

The specification must be compatible with the environment’s tools, state, and supported actions. Its instructions should be clear and avoid arbitrary constraints introduced solely to manufacture difficulty.

Agent-facing task

The user prompt specifies what the user wants the agent to accomplish, together with any user-level constraints. A generated task should satisfy three properties.

Feasibility. There should exist at least one trajectory in the current environment that satisfies the user prompt while respecting the system specification. This rules out tasks that depend on unavailable tools, inaccessible knowledge, impossible state transitions, or actions prohibited by policy.

Realism. The user prompt should resemble something a user would plausibly ask in the target environment. The space of executable behaviors is usually much larger than the space of realistic workflows.

Difficulty. For training, the task should expose a weakness of the current agent. Tasks that are already solved reliably provide little new training signal. The useful region is therefore tasks that are feasible and realistic, but not yet consistently solved.

Verifier

The verifier determines whether the resulting trajectory successfully completes the task. It should satisfy three properties.

Consistency. It should agree with the user prompt, the system specification, and the task-specific environment state.

Soundness. It should reject trajectories that fail to satisfy the task or violate relevant constraints.

Completeness. It should accept valid solutions rather than encode one particular reference trajectory.

These properties matter directly during training. A lax verifier can reward incorrect behavior, while an overly restrictive verifier can penalize valid solutions.

Overview

Given an environment and a target model, AutoSynthData generates training tasks consisting of a system specification, user prompt, and verifier. The generated tasks are grounded in the environment and selected to provide useful training signal for the current model.

figure-01

AutoSynthData first evaluates the target model in the environment using diagnostic tasks and identifies patterns in the tasks it struggles to complete. A stronger teacher helps characterize which of those tasks are solvable and what successful behavior looks like. AutoSynthData turns the resulting capability gaps into new executable tasks, checks each task in the environment, and uses accepted samples for post-training. Evaluating the updated model reveals which gaps remain and can guide the next round of generation.

figure-02

From model failures to a curriculum

AutoSynthData uses evaluation runs in the target environment to identify what the model needs to learn next. In our EnterpriseOps Gym experiment, we run both the target model and a stronger teacher on the evaluation tasks. We examine those runs to identify:

  • the capability being tested;
  • the tools and workflow structure involved;
  • where the target model fails and how the teacher succeeds;
  • the properties that a correct final state must satisfy;
  • the dimensions that can vary while preserving the capability being tested.

We distill these findings into sanitized capability specification cards. The evaluation tasks guide what the model should learn, but the generator does not receive their original prompts, entities, trajectories, or verifier details. It receives the cards and uses them to create new tasks with different prompts, states, and solution paths.

figure-03

Generating and scaling tasks

Identifying a capability gap tells us what to teach, but training requires many varied tasks that exercise it. AutoSynthData uses the specification card to generate those tasks.

Suppose the target model struggles with tasks that require the following workflow:

figure-04

The generator creates new tasks that exercise this workflow, varying the entities, initial environment state, workflow composition, tool combinations, wording, and difficulty. The stronger teacher then demonstrates a successful trajectory for each task. For supervised fine-tuning (SFT), these demonstrations teach the target model how to apply the capability in new situations.

AutoSynthData builds the dataset in two phases: first generating and validating core samples, then expanding them into novel variants.

Target

The target phase creates the core set of training samples from the capability specifications. Workers generate independent tasks in parallel, picking up a new target when they finish. Each candidate goes through validation, execution, solver evaluation, and repair before acceptance. The result is a batch of vetted examples built around what the target model needs to learn.

Multiply

The multiply phase expands the dataset by creating novel variants of accepted target samples. Each variant has its own user request, environment state, entity configuration, reference trajectory, and verifier, and must pass the same validation and execution checks. A multiplied sample cannot seed another multiplied sample. This anchors expansion to the vetted target set and limits drift across generations.

Implementation details

To support both phases, AutoSynthData separates generation control from environment-specific execution. A shared controller coordinates generation, quality control, coverage, and dataset construction, while an adapter handles environment execution, task and state management, reference replay, deterministic verification, solver execution, and task profiling.

Together, parallel target generation and multiplication provide a path to training-scale datasets. Their usefulness depends on the checks applied to every candidate: the task must be executable, the solution must work, and the verifier must distinguish success from failure.

figure-05

High-quality synthetic data needs more than generation

Generating a plausible request is not enough to produce useful training data. A task may be impossible in the target environment, its reference solution may fail when executed, or its verifier may reward the wrong final state. AutoSynthData checks these properties before accepting a task for training.

AutoSynthData reviews quality at two levels: individual candidates must pass verification, and batches must provide useful coverage and diversity.

Sample-level verification and repair

Each candidate must clear a quality-control loop before entering the training dataset. We begin with solver evaluation to measure difficulty. In the configuration used here, we favor tasks the target model solves on no more than one of three trials and the stronger solver solves on at least two of three trials. Candidates also undergo positive and negative verification and a bounded repair process.

figure-06

Positive verification

The positive gate asks: Does the intended solution solve the generated task?

The pipeline executes the reference trajectory in the target environment and checks the resulting state against the candidate’s verifier. This reveals mismatches among the prompt, initial state, solution, and success criteria.

Negative verification

The negative gate asks: Do relevant incorrect outcomes fail?

For example, it can mutate parts of the expected outcome and confirm that those states no longer pass verification. This catches weak verifiers that award success without requiring the intended behavior.

Critique and repair

Failed candidates go to a critic before being discarded. The critic examines the sample and its failure, looking for inconsistent state, impossible workflows, incorrect task construction, bad reference trajectories, weak verifier logic, or a mismatch with the intended capability. The critic’s findings guide repairs, with a fixed limit on retries:

candidate
↓
failure
↓
critique / diagnosis
↓
targeted repair
↓
run the gates again
↓
accept or retry

A repaired task must pass the relevant checks again. The diagnosis guides repairs to the existing candidate rather than requiring generation to start over.

Passing these checks makes a sample eligible for training, but individually valid samples can still form a repetitive or unbalanced dataset. AutoSynthData therefore also reviews generation at the batch level.

Batch-level review

A batch may overrepresent a few easy task families, miss a capability, or reflect too much generation effort spent on a low-yield pattern.

A meta-review examines accepted samples, rejected samples, and generation behavior across each batch. It asks:

  • Which task families are overrepresented, and which capability dimensions are missing?
  • Are the same kinds of examples appearing repeatedly?
  • Do particular targets keep failing generation?
  • Are systematic problems appearing in critiques?
  • What guidance should change for the next batch?

The controller tracks coverage in the accepted dataset, reduces generation in overrepresented regions, and directs more work toward gaps. When a region repeatedly produces poor candidates, critiques and meta-review guide changes to the generation strategy. These adjustments balance useful learning signal, task quality, coverage, diversity, and low redundancy within the available generation budget and dataset size requirements.

Together, these feedback loops improve both individual tasks and the dataset they form: sample-level checks guide candidate repair, while batch-level review guides future generation.

Moving the training frontier

The useful training distribution changes as the model improves. AutoSynthData treats synthetic data generation as a search for tasks near the target model's capability boundary: difficult enough to expose weaknesses, but solvable enough for the teacher to provide reliable demonstrations.

After post-training, we evaluate the updated model in the same environment. Tasks it now solves reliably are less useful for the next training round; persistent failures point to capabilities that still need attention. Those results can guide the next generation round.

figure-07

Our experiments focus on SFT, but the same mechanism could support reinforcement learning (RL): generate tasks that challenge the current policy and provide reliable learning signal, train, then move the generation target with the updated policy. We plan to test this moving, difficulty-calibrated frontier beyond SFT.

EnterpriseOps Gym experiments

We use EnterpriseOps Gym to test whether this approach improves a model on tasks in a stateful enterprise environment. We generate training tasks in the Gym’s Hybrid and ITSM environments, fine-tune the target model on accepted samples, and evaluate the resulting checkpoints.

Hybrid

We tested the pipeline on the Hybrid domain of EnterpriseOps Gym using Gemma-4-26B-A4B-it as the target model and as the teacher.

AutoSynthData generated 2,000 synthetic training samples in about 18 hours. We fine-tuned Gemma on this dataset and evaluated the resulting checkpoints on the benchmark. The best checkpoint was epoch 5.

Hybrid results

The synthetic SFT checkpoint improves mean Pass@1 by 7.2 percentage points, a 35% relative improvement, and raises verifier success from 63.01% to 68.55%. It closes 59% of the original Pass@1 gap between Gemma and the reference model.

figure-08

The training tasks were newly generated from capability specifications; the generator did not receive the original evaluation tasks. The result demonstrates improvement in EnterpriseOps Gym Hybrid, the environment used for this experiment.

ITSM

We also applied AutoSynthData to the ITSM domain of EnterpriseOps Gym, using Gemma-4-26B-A4B-it as the target model and as the teacher. AutoSynthData generated 1,994 synthetic training samples in 66 hours. Generation took longer than in the subsequent Hybrid run described above, primarily because the ITSM run used a larger teacher model and preceded pipeline optimizations that improved throughput.

On ITSM, synthetic SFT raises mean Pass@1 from 18.77% to 27.18%, showing that the approach also improves performance in a second domain.

figure-09

Closing the loop

The tasks most useful for training depend on both the environment and the model working in it. AutoSynthData uses the model’s failures to choose what to generate, validates new tasks against the environment, and makes those tasks available for post-training. Our EnterpriseOps Gym results show the value of that approach in a controlled setting. As the model changes, the same process can focus on the gaps that remain.

文章评论(12)

星观澜8 小时前

实话,说的不明不白

回复
杨丽华10 小时前

楼主辛苦了,内容很有参考价值。

回复
龙文博10 小时前

讲解得很细致,新手也能看懂。

回复
林观澜7 小时前

支持作者,持续关注中。

回复
龙文博8 小时前

点赞,必须点赞

回复
杨丽华6 小时前

很有价值的分享,感谢整理。

回复
龙文博6 小时前

不错不错,已加入书签。

回复
云逐月4 小时前

实测过类似工具,作者说的基本属实。

回复
墨沐雨4 小时前

很有价值的分享,感谢整理。

回复
空向阳2 小时前

有没有更详细的教程,期待后续。

回复
黄小刚4 小时前

楼主辛苦了,内容很有参考价值。

回复
一叶知秋2 小时前

支持作者,持续关注中。

回复