AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

ServiceNow CoreAI introduced AutoSynthData, a pipeline described in a Hugging Face report for generating training tasks tailored to enterprise agents and their operating environments. The method uses target-model failures and a stronger teacher’s successful runs to identify capability gaps, then creates and checks new tasks; the report illustrates the approach with EnterpriseOps Gym.

ServiceNow CoreAI has described AutoSynthData, a pipeline designed to generate training data for enterprise agents by using their failures to identify skills they need to improve. In a report published on Hugging Face, the team says the system combines evaluations of a target model with successful runs from a stronger teacher, then generates and validates new tasks in the target environment. The report uses EnterpriseOps Gym to illustrate the approach; it does not establish that the method has been deployed in production or quantify its effect on real-world enterprise performance.

AutoSynthData is intended to address a mismatch between general model capability and the specific demands of an enterprise environment. An agent may struggle with a particular workflow, misuse a combination of tools, or fail to follow a relevant constraint even if it performs well on other tasks. The report frames the challenge as converting such weaknesses into enough varied examples to support training.

The pipeline starts by evaluating a target model on diagnostic tasks in the environment. The team compares its attempts with runs from a stronger teacher model to characterize what capability is being tested, where the target fails, and what a successful outcome requires. Those findings are distilled into sanitized capability specification cards. According to the report, the generator receives these cards—not the original evaluation prompts, entities, trajectories, or verifier details—and uses them to create new tasks with different prompts, states, and solution paths.

Each generated task includes a system specification, a user prompt, and a verifier. The specification sets the rules and environment conditions; the prompt states the requested work; and the verifier judges whether the outcome meets the requirements. AutoSynthData checks generated tasks in the environment and uses accepted samples for post-training. The updated model can then be evaluated again so remaining weaknesses can shape a further round of task generation.

At a glance
announcementWhen: Described in a Hugging Face report; the…
The developmentServiceNow CoreAI published a report describing AutoSynthData, a system for turning enterprise-agent capability gaps into generated, validated training tasks.

Training Agents for Local Workflows

The proposal addresses a practical limitation in enterprise AI: successful behavior depends not only on a model’s general abilities, but also on the tools, policies, data, and workflows in the particular environment where it operates. Training against tasks grounded in that environment could give developers a more targeted way to improve a model than relying only on broad, generic examples.

The approach also makes task quality part of the training problem. A generated exercise should be possible to complete with available tools, resemble plausible user work, and test something the current model does not already handle reliably. Its verifier matters just as much: a permissive check may reward incorrect results, while a check tied too closely to one reference path may reject other valid solutions. The report therefore presents feasibility, realism, difficulty, and reliable verification as requirements, not optional refinements.

If the method works as intended, teams could use evaluation failures to guide the next training curriculum as an agent’s performance changes. That would be relevant to organizations deploying agents across systems with distinct rules and data. However, the source describes a method and an experiment, not evidence that AutoSynthData improves deployed agents, lowers costs, or outperforms other data-generation approaches.

Amazon

AI training data generation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Evaluation to New Tasks

AutoSynthData treats an agent task as a combination of system specification, user prompt, and verifier, all tied to an environment that defines observable state, available tools, and the results of actions. This framing matters because a prompt that sounds realistic may still be unusable if it requires an unavailable tool or an impossible change to the environment.

In the reported EnterpriseOps Gym example, diagnostic evaluations indicate which capabilities need attention. The generator then uses a distilled description of those capabilities to produce different executable tasks, rather than simply copying evaluation examples. The intended result is a curriculum that changes as the target model improves: tasks already solved consistently offer less training value, while feasible and realistic tasks that expose persistent weaknesses can provide more useful signal.

The source cites EnterpriseOps Gym (Malay et al., 2026) and says it uses the released dataset to illustrate the pipeline. It provides a description of the method, but the material supplied here does not include a full account of measured results, comparisons with alternative methods, or a production deployment.

“At ServiceNow CoreAI, we built AutoSynthData to turn those capability gaps into training data.”

— ServiceNow CoreAI, in its Hugging Face report

Amazon

enterprise AI model training datasets

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evidence and Deployment Still Open

The report describes the pipeline and its use with EnterpriseOps Gym, but the source material provided does not establish how much model performance improved, how many generated tasks passed validation, or whether gains held across different enterprise environments. It also does not provide enough detail here to assess costs, generation speed, or the reliability of the teacher model’s judgments.

It remains unclear whether the capability cards and verifiers can be produced consistently for a wide range of tools and policies, and how teams would detect mistakes in generated tasks that pass automated checks. The report’s stated workflow includes validation, but validation alone does not confirm that every accepted task reflects useful or safe work in a live organization. No production deployment or independent evaluation is established in the supplied source.

Amazon

automated machine learning data labeling

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Further Testing Will Matter

The report describes an iterative next step: post-train the target model on accepted tasks, evaluate the updated model in the environment, and use the remaining gaps to guide another generation cycle. That process would test whether the curriculum shifts in response to learning rather than repeatedly producing tasks the model already handles.

For readers assessing the system, the next evidence to watch for is a fuller account of benchmark results, task-validation rates, and comparisons with other ways of producing training data. Testing across additional environments would also help show whether the approach generalizes beyond the EnterpriseOps Gym example. The source material does not specify a timeline for those results.

Amazon

AI model validation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is AutoSynthData?

AutoSynthData is a ServiceNow CoreAI pipeline described for generating and validating training tasks for enterprise agents. It aims to base those tasks on weaknesses found when evaluating a target model in a specific environment.

How does it choose what an agent should learn?

The system compares a target model’s evaluation runs with successful runs from a stronger teacher. The findings are summarized in capability specification cards that guide generation of new tasks, according to the report.

How does AutoSynthData check generated tasks?

Tasks are grounded in the environment and include a verifier that checks whether the resulting behavior meets the prompt and relevant constraints. The report says generated tasks are checked in the environment before accepted samples are used for post-training.

Has AutoSynthData been shown to improve production agents?

The supplied report describes the method and illustrates it with EnterpriseOps Gym. The source material does not establish production deployment or provide quantified evidence of improved real-world performance.

Source: rss

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Multi-Vector (Late Interaction) Embedding Models With Sentence Transformers

Sentence Transformers v6.0 adds MultiVectorEncoder, enabling ColBERT-style late-interaction retrieval for text and visual documents, with larger indexes and complex scoring.

A Beginner’s Guide To IEP Negotiation Rehearsal Simulators

A proposed simulator would help parents review IEP documents, prepare requests and rehearse school meetings. Its effect on outcomes remains untested.

What To Know About 14 AI Workflow Automation Tools In 2027

A comparison of 14 AI workflow automation guides, from no-code introductions to developer and industry-specific books, and what readers should verify.

Maximize Efficiency With AI Tools & Automation In 2026

Explore how AI tools and automation are transforming productivity and security in 2026, with confirmed developments and key insights for users.