Self-improving data synthesis scales embedding models across domains and tasks

(SYNTRA)

† Equal contribution.   * Corresponding author.

University of Science and Technology of China, Hefei, China

Manuscript draft · 6 October 2026.

Coding and research agents retrieve information from distinct corpora. Different objectives change relevance even within one domain. Few-shot-guided self-improvement adapts SYNTRA to construct task-specific training records.
Task-specific supervision across domains and information needs. Few-shot guidance adapts a shared synthesis workflow. Examples are schematic and abbreviated.Select any figure to open it at full resolution.

SYNTRA adapts a shared retrieval-data synthesis workflow through few-shot-guided self-improvement, turning an unlabelled corpus and a few query–positive-document examples into task-specific training data.

Abstract

Adapting embedding models to new domains and retrieval tasks requires training data that reflect task-specific information needs and relevance criteria. Constructing such data often depends on human design. Here we introduce SYNTRA, a self-improving agentic framework for retrieval data synthesis across domains and tasks. Given an unlabelled corpus and a few examples of relevant query–document pairs, SYNTRA infers task requirements and uses feedback on trial outputs to refine shared instructions for large-scale synthesis. Automatic refinement improves downstream retrieval across six tasks, outperforming human calibration on four. Evaluations across 24 retrieval datasets spanning 15 domains demonstrate the effectiveness of the generated training data. These data also benefit strong embedding models with and without existing task-specific supervision. On MS MARCO, retrieval performance continues to improve over a 32-fold data expansion. SYNTRA offers a reusable approach to adapting retrieval data synthesis with minimal human intervention.

Method

Task preparation and test-mode adaptation

Corpus preprocessing selects diverse source documents and builds a full-corpus retrieval index. Few-shot examples guide a task definition specifying the query type, document type and relevance criteria. Small-sample generate–evaluate–refine loops then establish task-adapted guidance for identifier and attribute construction, query synthesis and relevance-label synthesis.

Each loop ends when verification passes or its iteration limit is reached. All loops finish before production, and the final shared instructions remain fixed throughout production. In the experiments, human involvement included reviewing task definitions and occasionally adjusting the document-selection threshold.

An unlabelled corpus and few-shot query-positive pairs establish a retrieval task. Three small-sample generate-evaluate-refine loops produce shared instructions for identifiers and attributes, queries, and relevance labels.
Task preparation and adaptation. Few-shot examples establish task requirements; optional user review is recommended. Test-mode feedback revises reusable synthesis guidance before production begins.

Production-scale synthesis

For each selected document, SYNTRA extracts document-grounded identifiers and constructs query attributes. Plans combine an identifier with attribute options, and query synthesis expresses each planned information need. Candidate documents are retrieved from the full corpus and assessed against task-specific relevance criteria. The resulting query, positive and negative records support embedding-model training after synthesis.

Three production modules reuse fixed task-adapted instructions: semantic identifier extraction and attribute construction, query planning and writing, and candidate retrieval followed by relevance annotation.
Production-scale synthesis. The shared instructions remain fixed while document-specific plans, queries, candidates and labels change. Content shown in the workflow is illustrative.

Self-improvement

Few-shot examples convey the intended information need and relevance criteria. In test mode, SYNTRA generates small trial samples, evaluates their alignment with the task and revises the shared instructions using diagnostic feedback. The revisions guide later synthesis across documents.

Automatic self-improvement increased downstream retrieval performance on all six tested tasks. On NFCorpus and FiQA, data generated with the initial instructions yielded scores below the MS MARCO-trained baseline; data generated with refined instructions exceeded it.

nDCG@10 increased from 24.828 to 36.007 on NFCorpus and from 28.873 to 35.529 on FiQA after self-improvement.

Automatic refinement also yielded higher scores than human calibration on four of six tasks. Text2SQL was close (62.240 versus 62.449), while human calibration remained stronger on TheoremQA (theorems), at 26.681 versus 22.199.

FiQA case study

For query synthesis, feedback redirected generation towards a user's financial information need. For relevance assessment, a known-positive example exposed an overly narrow rule: an answer explaining why a mortgage-payment plan may fail still helps resolve the user's question. Refinement incorporated that distinction into the shared annotation instruction.

FiQA case study: trial-query feedback revises the query instruction to express a user's financial need. Few-shot feedback revises relevance criteria so an answer explaining why a plan fails is labelled positive. Retrieval nDCG@10 rises from 28.873 to 35.529.
Few-shot-guided self-improvement on FiQA. Query examples are verbatim outputs from separate production runs with different refinement-round instructions; attribute plans were not held fixed. The relevance panel summarizes a recorded calibration episode. The retrieval gain reflects joint refinement of query and relevance-annotation instructions.

Results

One workflow across 24 datasets and 15 domains

Training on SYNTRA-generated data improved retrieval in all 23 task-specific comparisons against a common MS MARCO-trained baseline. The embedding backbone and adaptation protocol were held fixed, with data generated and a Qwen3-0.6B backbone adapted separately for each task. MS MARCO is included among the 24 datasets as the common training corpus, without a task-adaptation gain.

These comparisons evaluate the complete framework. The separate six-task ablation directly assesses the contribution of self-improvement.

Results across 24 datasets and 15 domains, a six-task comparison of initial instructions, automatic self-improvement and human calibration, task-specific query-source comparisons, and agreement with seven human annotators.
Cross-task applicability, self-improvement and supervision quality. Panels show coverage and relative retrieval gains, the six-task self-improvement ablation, query-source comparisons, and human-reference label agreement. Numeric cross-task labels are percentage gains; radial extensions use a nonlinear scale. nDCG@10 scores use a 0–100 scale.

Synthetic queries provide effective training supervision

In a matched MS MARCO experiment with 10,000 training examples per condition, SYNTRA-generated queries exceeded original Bing search-log queries on both evaluation suites. Document source, relevance annotation, negative mining and the training protocol were held fixed.

Mean nDCG@10 under matched training conditions
Evaluation suiteOriginal queriesSYNTRA queries
BEIR42.1842.87
NanoBEIR54.5355.51

Generated labels agree with human relevance judgements

Qwen3-30B-A3B labels were compared with reference labels derived from seven independent annotators. The evaluation covered 150 query–document pairs, with 50 each from MS MARCO, Text2SQL and TheoremQA (theorems).

Binary relevance agreement was 91.3% (Cohen's κ = 0.719), and three-class agreement was 68.7% (κ = 0.510). Pooled exact-category precision among LLM-labelled positives and hard negatives was 82.8%.

Most three-class disagreements concerned hard versus easy negatives. Pooled precision was 72/87 exact matches among pairs the LLM labelled positive or hard negative. Four three-class voting ties were assigned hard-negative reference labels.

Gains extend to strong pretrained embedding models

BGE-M3 and Qwen3-Embedding-8B both improved on all three evaluated tasks without existing task-specific labelled training sets. Both also improved on all three tasks where SYNTRA-generated data augmented existing labelled sets, compared with fine-tuning on those sets alone.

More synthetic data continue to improve retrieval

On MS MARCO, nDCG@10 rose from 30.88 to 33.20 as training data increased from 10,000 to 320,000 records, a 32-fold expansion. Performance improved at each of six evaluated scales with a fixed Qwen3-8B training recipe. Each scale was evaluated in one training run.

BGE-M3 and Qwen3-Embedding-8B adaptation with and without existing labelled data, NanoBEIR component ablations, task-specific training-source comparisons, MS MARCO data scaling, and multiple queries per document.
Model adaptation, synthesis components and scaling. The two adaptation settings use their corresponding pretrained or labelled-data-only baselines. Component ablations are averaged over 13 NanoBEIR tasks. Data scaling is evaluated on MS MARCO; query-count comparisons are evaluated on CodeTrans-DL and MedicalQA.

Task-adapted instructions and synthesis components matter

Mean NanoBEIR nDCG@10 was 54.64 for the complete workflow. Removing the task-adapted query instruction reduced it to 52.56; removing identifier–attribute planning reduced it to 53.96; replacing relevance annotation with hard-negative mining reduced it to 52.80. Generating one, two and four queries per document also improved performance on CodeTrans-DL and MedicalQA.

Discussion

Few-shot-guided self-improvement lets the synthesis system take on part of the task-specific design work needed to adapt retrieval models. The experiments support this approach within the evaluated retrieval settings, while retaining human review at task setup. Small-scale calibration establishes instructions that can be reused in large-scale production with a locally deployed model.

The evidence also identifies useful boundaries: negative difficulty is harder to label consistently than binary relevance; data scaling was tested on MS MARCO; and generating more queries changes both supervision quantity and query variation.

Future directions

Self-improvement could extend to training-data synthesis for other LLM capabilities, with feedback combining correctness checks and task-specific assessments. Model evaluation, task outcomes and user interactions could also reveal emerging training needs. These are research directions beyond the capabilities evaluated here.