ServiceNow CoreAI Debuts AutoSynthData Training Pipeline
ServiceNow CoreAI has introduced AutoSynthData, a pipeline that generates tailored synthetic training data to help enterprise AI agents overcome specific operational weaknesses.

ServiceNow CoreAI has developed AutoSynthData, a system designed to address the unique challenges of training enterprise AI agents. While general-purpose models are highly capable, they often struggle with the specific tools, rules, and workflows of individual corporate environments. AutoSynthData solves this by identifying a target model's failures, using a stronger teacher model to demonstrate successful outcomes, and generating new, validated training tasks that target those specific capability gaps.
The pipeline operates in two distinct phases. First, the target phase creates core training samples from capability specifications. Second, the multiply phase expands these into novel variants with distinct user requests and environment states. To ensure high data quality, AutoSynthData subjects each candidate task to rigorous validation. This includes a solver evaluation to measure difficulty, a positive gate to verify that the intended solution works, and a negative gate to confirm that incorrect outcomes are rejected. A critic agent diagnoses failed tasks to guide repairs, while a batch-level meta-review prevents repetitive data generation.
ServiceNow CoreAI tested the pipeline on the Hybrid domain of the EnterpriseOps Gym benchmark. Using Gemma-4-26B-A4B-it as the target model and Qwen3.8-27B as the teacher, AutoSynthData generated 2,000 synthetic training samples in approximately 18 hours. Fine-tuning Gemma on this dataset yielded a best checkpoint at epoch 5, which improved the mean Pass@1 score by 7.2 percentage points—a 35% relative improvement. This training also raised verifier success from 63.01% to 68.55% and closed 59% of the original Pass@1 performance gap between Gemma and the reference model.
The researchers also evaluated the pipeline on the ITSM domain of EnterpriseOps Gym. For this test, Gemma-4-26B-A4B-it remained the target model, while DeepSeek-V4.1-Flash served as the teacher. AutoSynthData generated 1,994 synthetic training samples over 66 hours, a longer duration attributed to the larger teacher model and a lack of subsequent pipeline optimizations. Despite the longer run, the resulting synthetic supervised fine-tuning successfully raised the mean Pass@1 score from 18.77% to 27.18%, proving the framework's effectiveness across multiple enterprise domains.
This is our own summary of reporting by Hugging Face Blog



