Abstract
Adapting embedding models to new domains and retrieval tasks requires training data that reflect task-specific information needs and relevance criteria. Constructing such data often depends on human design. Here we introduce SYNTRA, a self-improving agentic framework for retrieval data synthesis across domains and tasks. Given an unlabelled corpus and a few examples of relevant query–document pairs, SYNTRA infers task requirements and uses feedback on trial outputs to refine shared instructions for large-scale synthesis. Automatic refinement improves downstream retrieval across six tasks, outperforming human calibration on four. Evaluations across 24 retrieval datasets spanning 15 domains demonstrate the effectiveness of the generated training data. These data also benefit strong embedding models with and without existing task-specific supervision. On MS MARCO, retrieval performance continues to improve over a 32-fold data expansion. SYNTRA offers a reusable approach to adapting retrieval data synthesis with minimal human intervention.
Method
Task preparation and test-mode adaptation
Corpus preprocessing selects diverse source documents and builds a full-corpus retrieval index. Few-shot examples guide a task definition specifying the query type, document type and relevance criteria. Small-sample generate–evaluate–refine loops then establish task-adapted guidance for identifier and attribute construction, query synthesis and relevance-label synthesis.
Each loop ends when verification passes or its iteration limit is reached. All loops finish before production, and the final shared instructions remain fixed throughout production. In the experiments, human involvement included reviewing task definitions and occasionally adjusting the document-selection threshold.
Production-scale synthesis
For each selected document, SYNTRA extracts document-grounded identifiers and constructs query attributes. Plans combine an identifier with attribute options, and query synthesis expresses each planned information need. Candidate documents are retrieved from the full corpus and assessed against task-specific relevance criteria. The resulting query, positive and negative records support embedding-model training after synthesis.
Self-improvement
Few-shot examples convey the intended information need and relevance criteria. In test mode, SYNTRA generates small trial samples, evaluates their alignment with the task and revises the shared instructions using diagnostic feedback. The revisions guide later synthesis across documents.
Automatic self-improvement increased downstream retrieval performance on all six tested tasks. On NFCorpus and FiQA, data generated with the initial instructions yielded scores below the MS MARCO-trained baseline; data generated with refined instructions exceeded it.
nDCG@10 increased from 24.828 to 36.007 on NFCorpus and from 28.873 to 35.529 on FiQA after self-improvement.
Automatic refinement also yielded higher scores than human calibration on four of six tasks. Text2SQL was close (62.240 versus 62.449), while human calibration remained stronger on TheoremQA (theorems), at 26.681 versus 22.199.
FiQA case study
For query synthesis, feedback redirected generation towards a user's financial information need. For relevance assessment, a known-positive example exposed an overly narrow rule: an answer explaining why a mortgage-payment plan may fail still helps resolve the user's question. Refinement incorporated that distinction into the shared annotation instruction.
Results
One workflow across 24 datasets and 15 domains
Training on SYNTRA-generated data improved retrieval in all 23 task-specific comparisons against a common MS MARCO-trained baseline. The embedding backbone and adaptation protocol were held fixed, with data generated and a Qwen3-0.6B backbone adapted separately for each task. MS MARCO is included among the 24 datasets as the common training corpus, without a task-adaptation gain.
These comparisons evaluate the complete framework. The separate six-task ablation directly assesses the contribution of self-improvement.
Synthetic queries provide effective training supervision
In a matched MS MARCO experiment with 10,000 training examples per condition, SYNTRA-generated queries exceeded original Bing search-log queries on both evaluation suites. Document source, relevance annotation, negative mining and the training protocol were held fixed.
| Evaluation suite | Original queries | SYNTRA queries |
|---|---|---|
| BEIR | 42.18 | 42.87 |
| NanoBEIR | 54.53 | 55.51 |
Generated labels agree with human relevance judgements
Qwen3-30B-A3B labels were compared with reference labels derived from seven independent annotators. The evaluation covered 150 query–document pairs, with 50 each from MS MARCO, Text2SQL and TheoremQA (theorems).
Binary relevance agreement was 91.3% (Cohen's κ = 0.719), and three-class agreement was 68.7% (κ = 0.510). Pooled exact-category precision among LLM-labelled positives and hard negatives was 82.8%.
Most three-class disagreements concerned hard versus easy negatives. Pooled precision was 72/87 exact matches among pairs the LLM labelled positive or hard negative. Four three-class voting ties were assigned hard-negative reference labels.
Gains extend to strong pretrained embedding models
BGE-M3 and Qwen3-Embedding-8B both improved on all three evaluated tasks without existing task-specific labelled training sets. Both also improved on all three tasks where SYNTRA-generated data augmented existing labelled sets, compared with fine-tuning on those sets alone.
More synthetic data continue to improve retrieval
On MS MARCO, nDCG@10 rose from 30.88 to 33.20 as training data increased from 10,000 to 320,000 records, a 32-fold expansion. Performance improved at each of six evaluated scales with a fixed Qwen3-8B training recipe. Each scale was evaluated in one training run.
Task-adapted instructions and synthesis components matter
Mean NanoBEIR nDCG@10 was 54.64 for the complete workflow. Removing the task-adapted query instruction reduced it to 52.56; removing identifier–attribute planning reduced it to 53.96; replacing relevance annotation with hard-negative mining reduced it to 52.80. Generating one, two and four queries per document also improved performance on CodeTrans-DL and MedicalQA.
Discussion
Few-shot-guided self-improvement lets the synthesis system take on part of the task-specific design work needed to adapt retrieval models. The experiments support this approach within the evaluated retrieval settings, while retaining human review at task setup. Small-scale calibration establishes instructions that can be reused in large-scale production with a locally deployed model.
The evidence also identifies useful boundaries: negative difficulty is harder to label consistently than binary relevance; data scaling was tested on MS MARCO; and generating more queries changes both supervision quantity and query variation.
Future directions
Self-improvement could extend to training-data synthesis for other LLM capabilities, with feedback combining correctness checks and task-specific assessments. Model evaluation, task outcomes and user interactions could also reveal emerging training needs. These are research directions beyond the capabilities evaluated here.
