Technical Report · 2026

Let Your Data Build and Improve Itself

DataEvolver is an autonomous dataset construction framework where goal-driven loop agents choose and compose text-to-image, image editing, text-to-video, and image-to-3D interfaces, then use VLM feedback and WebSearch research priors to turn user requests or papers into reproducible data pipelines.

Qisong Zhang, Wenzhuo Wu, Zhuangzhuang Jia, Yunhao Yang, Shuo Zhang School of Artificial Intelligence, Beijing University of Posts and Telecommunications (BUPT)
📄 arXiv 2605.01789 🏆 KDD Workshop Oral 🎖️ BAAI Poster
DataEvolver autonomous synthetic data construction pipeline poster
Autonomous synthetic data construction pipeline Text expansion → T2I → segmentation → 3D reconstruction → scene rendering → VLM review loop → training-ready data
Agent-guided
4
Generation Interfaces
2
Construction Tracks
24
Repair Actions
Web
Research Prior

Why Dataset Construction Needs Agent Strategy

Dataset construction is no longer a single fixed path. Some tasks need image generation, some need editing or video synthesis, some need 3D reconstruction and Blender insertion, and paper reproduction needs source-backed methodology rather than ad hoc prompting.

Single-Path Pipelines Are Too Rigid

A user request may require parallel T2I candidates, an image-edit branch, a text-to-video branch, or a serial image-to-3D route. The agent must choose the right composition.

Synthetic Outputs Need Semantic Repair

Rigid numeric thresholds cannot diagnose "this object floats" or "the lighting is inconsistent." Goal-driven agents use VLM feedback to decide targeted repairs.

Paper Reproduction Needs Sources

A paper or topic should produce a traceable dataset construction handoff: cited sources, official artifacts, evaluation references, and reusable requirements.

Two Connected Dataset Construction Tracks

DataEvolver combines an agent-orchestrated multimodal generation track with a WebSearch-guided research-prior track. The agent can run interfaces in parallel, in series, or as a hybrid strategy to build faster, higher-quality datasets.

Text-to-Image

Generate source objects, scene references, or candidate visual concepts from prompts.

Image Edit

Produce controlled variants, target images, and editing pairs from accepted sources.

Text-to-Video

Extend static generation into temporal object or scene workflows when sequence data is needed.

Image-to-3D

Reconstruct textured objects for Blender insertion, scene rendering, and VLM repair.

DataEvolver six-stage synthetic data generation pipeline
Representative serial 3D route The image-to-3D path remains a concrete serial strategy: prompt/image generation, segmentation, 3D reconstruction, Blender rendering, and VLM review.
1
Request Strategy
Agent maps the user goal to interface choices
2
Composable Interfaces
T2I, image edit, T2V, and image-to-3D
3
Branch Orchestration
Parallel, serial, or hybrid execution
4
3D Route
Object image to textured mesh to Blender
5
Scene Context
Optional HYWorld / WorldMirror support
6
Review & Export
VLM repair loop, metadata, and dataset handoff

WebSearch-Guided Reproducible Research Pipelines

The research-prior workflow turns a paper, topic, or generation request into a reproducible dataset-construction handoff. The active agent performs WebSearch and paper reading, then dynamically composes serial and parallel t2i, edit, i2v, and 3D routes while DataEvolver keeps source notes, trace events, gallery sheets, and research_prior.json attached for replay.

Research-Guided Pipeline

Single-paper and universal modes merge into one source-backed workflow for reproducible dataset construction and evaluation-ready handoff.

WebSearch-guided reproducible research pipeline from paper or topic input to evidence handoff
WebSearch-guided reproducible research pipelines The agent retrieves papers, official artifacts, code, and datasets, then records evidence, GT references, manifests, lineage, and research_prior.json.
Single-Image Route Case

Robot Safety Signal: T2I → Edit → I2V

A compact three-stage example from one accepted image. The route starts with a text-to-image base render, promotes the selected frame through targeted image editing, then uses the edited result as the anchor for short motion synthesis.

1Base Image
1Edit Pass
1I2V Clip
MixStrategy
Robot base image generated by text-to-image
T2I Base
Initial robot render

Clean lab robot scene with a blue transparent cube held at the center.

Robot edited image with safety signal and green status lights
Edit Pass
Safety signal added

Targeted edit adds a green status light and a front safety mark while preserving pose and scene layout.

I2V Output
Motion-ready result

The edited image becomes the visual anchor for a short image-to-video generation pass.

Serial / Parallel Strategy Orchestration

Use parallel search where alternatives are cheap, then switch to serial promotion once a stable image anchor is selected.

Parallel Explore base candidates

Fan out prompts, seeds, or style variants to quickly compare possible robot scenes.

Select Promote one anchor

Choose the frame with the clearest pose, object placement, and scene consistency.

Serial Apply targeted edit

Lock geometry and camera first, then add the safety indicator and status lights.

Serial Generate motion

Use the edited image as the input anchor for I2V so identity and visual edits carry forward.

Parallel T2I search Anchor selection Serial edit Serial I2V Replayable handoff

Serial / Parallel Route Planner

The agent turns the research prior into a route decision: run cheap branches together, cascade dependent steps only when needed, and keep the selected plan replayable through manifests and gallery records.

Parallel
Fan-out candidates

Run t2i, edit, or i2v candidates side by side when the task needs diversity, fast comparison, or multiple visual hypotheses.

Serial
Cascade dependencies

Chain image -> edit -> i2v or image -> 3D -> Blender when later stages need validated upstream artifacts.

Plan Choice
Score route utility

Select the lowest-risk plan from cost, source coverage, reference availability, expected quality, and replay requirements.

WebSearch prior-> Source plan-> Route scoring-> Serial / parallel execution-> Gallery handoff
Paper
Single-Paper Mode

Reproduce a target paper by separating official artifacts from generated proxies, then recording dataset method, metrics, GT references, and missing-assets decisions.

Topic
Universal Mode

Search 3-5 related papers, extract shared dataset requirements, and synthesize a reusable plan for generation, evaluation, and traceable outputs.

Research Handoff

  • Agent-run WebSearch and paper reading, with DataEvolver preserving source notes and trace events.
  • Dynamic route planning chooses when to run t2i, edit, i2v, and 3D in sequence or in parallel.
  • Official open datasets and evaluation criteria are treated as GT references, not replaced by VLM-only acceptance.
  • Every accepted route exports prompts, manifests, review sheets, and replayable handoff records.
WebSearch-guided single-paper universal dynamic plan research_prior.json

HYWorld / WorldMirror as a 3D Route Enhancer

HYWorld / WorldMirror is an optional world_model route for strengthening the 3D reconstruction track. It builds scene context before object insertion, then requires contract-backed geometry, pure-scene multi-view renders, manifests, and lineage records.

HYWorld automated 3D scene reconstruction pipeline
Optional scene context route Prompt or image → HY-Pano → HYWorld / WorldMirror geometry → Blender scene contract → pure-scene validation → object insertion.

Contract-Backed Geometry

HY-Pano creates a 360° context; HYWorld and WorldMirror reconstruct depth, camera priors, support surfaces, and scene geometry instead of accepting a fixed panorama shell.

Traceable Review

The route preserves pure-scene renders, placement contracts, support checks, and lineage hashes before any object-scene render can be promoted.

Route Records

  • Preserved run artifacts and SHA-256 manifests for HY-Pano and full worldgen.
  • Pure-scene multi-view renders before object insertion.
  • Mesh-raycast support placement with scene-camera matching.
  • Final object-scene reports include RGB, masks, depth, normals, metadata gates, and lineage instead of VLM-only acceptance.
world_model profile HYWorld WorldMirror Blender contract
720p 3D Route Case Pack

HYWorld Qwen2.5-12B In-place Object Rotation

A compact preview extracted from the zipped report. It packages two HYWorld scenes, fixed-camera source RGB backgrounds, mesh-zbuffer object placement, eight yaw angles per object, and validation records for replayable 3D dataset construction.

2Scenes
2Objects
8Yaw Views
720pOutput
34Enhanced RGB
PassValidation
HYWorld scene 001 object 001 in-place rotation contact sheet
Scene 001 / object 001 Original bundled contact sheet: static camera, fixed world position, eight yaw samples.
HYWorld scene 002 object 009 in-place rotation contact sheet
Scene 002 / object 009 Original bundled contact sheet: second scene, same in-place rotation review checks.

Validation Summary

  • Scene static, camera static, and source RGB background preserved.
  • Object remains fixed in world space while rotating in place.
  • All yaw angles are visible after harmonization and super-resolution.
  • Manifest, lineage, masks, metadata, and validation report stay bundled with the archive.

Goal-Driven Loop Agents: Perceive, Diagnose, Act, Repeat

In the 3D reconstruction route, a goal-driven loop agent perceives rendered outputs via VLM review, diagnoses semantic issues, selects targeted rendering adjustments from a structured action space, and repeats until quality goals are met.

DataEvolver closed-loop scene refinement pipeline
Closed-loop scene refinement Render, review, act, and stop when the quality gate returns keep.
Blender Render
Cycles 512spp at 1024×1024
VLM Review
Qwen3.5-35B free-form critique
Agent Decision
Reads review, selects from 24 actions
Quality Gate
Verdict: keep / revise / reject
Loop continues until reviewer says keep

Anti-Oscillation Control

Sign-flip tracking, dead-zone detection, and step-scale scheduling prevent infinite loops and parameter thrashing.

Scene-Aware Rendering

Objects can be inserted into configured Blender scenes or optional HYWorld / WorldMirror scene context with support-surface checks.

24 Atomic Actions

Structured action space across lighting, object transform, scene environment, and material property groups.

DataEvolver-Rotate: View-Controlled Rotation Editing

A benchmark dataset for rotation-conditioned image editing. Each sample pairs a canonical front-view image with a target view specified in natural language.

50
Unique Objects
8
Viewpoints / Object
350
Training Pairs
3×A800
Infrastructure
Python — Load Dataset
import json
from pathlib import Path
from PIL import Image

root = Path("dataset_scene_v7_full50_rotation8_...")
rows = []
with (root / "pairs/train_pairs.jsonl").open("r") as f:
    for line in f:
        rows.append(json.loads(line))

row = rows[0]
source = Image.open(root / row["source_image"]).convert("RGB")
target = Image.open(root / row["target_image"]).convert("RGB")
instruction = row["instruction"]

24 Structured Atomic Actions

In the 3D reconstruction route, the agent selects from a discrete, structured action space to address VLM-identified rendering issues. Each action targets a bounded Blender parameter.

Key Light Intensity

×1.2 / ×0.8 multiplicative, bounded [0.5, 2.0]

Key Light Rotation

±15° yaw step, bounded [-90°, 90°]

Env Light Intensity

×1.2 / ×0.8 multiplicative, bounded [0.5, 2.0]

Env Rotation (Z)

±30° step, bounded [-180°, 180°]

Object Elevation

±0.02 step, bounded [-0.1, 0.1]

Material Roughness

±0.08 step, bounded [-0.3, 0.6]

+ 18 more actions — see scene_action_space.json

Key Differentiators

What sets DataEvolver apart is not one model call, but the way agents choose routes, preserve run records, repair failures, and export reusable dataset construction traces.

Agent-Orchestrated Interfaces

The agent can compose T2I, image edit, T2V, and image-to-3D routes in parallel, in series, or as hybrid strategies based on dataset goals.

VLM-as-Feedback, Not Score

Free-form natural language feedback provides semantic diagnosis that numeric scores cannot. The reviewer identifies why a render fails.

Research-Grounded Pipelines

WebSearch-guided research records source notes, dataset methods, official artifacts, and evaluation references before downstream generation.

Evaluation-to-Optimization Path

Accepted traces, failure cases, and generated evaluation sets can guide future LoRA tuning for stronger reconstruction and editing models.

Built With

The models, frameworks, and traceability records powering DataEvolver's multimodal generation, 3D reconstruction, and research-prior workflows.

Python 3.10+ Blender 4.2+ Cycles Path Tracing PyTorch 2.8 Qwen-Image-2512 Image Edit Workflow Text-to-Video Workflow SAM3 Hunyuan3D-2.1 HYWorld / WorldMirror Qwen3.5-35B-A3B Qwen-Image-Edit-2511 DiffSynth-Studio WebSearch Research research_prior.json LoRA (PEFT) 3×A800 80GB

Quick Start

Start with the lightweight dry-run path. It validates the setup shape and prints the local plan without downloading models, writing tokens, or launching GPU jobs.

Shell — Setup
git clone https://github.com/PRIS-CV/DataEvolver.git
cd DataEvolver

# Safe onboarding dry-run
bash src/dataevolver/cli/bootstrap_dataevolver_default.sh \
  --profile quick \
  --dry-run

# Review the reproducible research handoff format
find outputs -name research_prior.json -print

# Read the public documentation
cat README.md

Citation

If you use DataEvolver or DataEvolver-Rotate in your research, please cite our work.

BibTeX
@misc{zhang2026dataevolverletdatabuild,
  title        = {DataEvolver: Let Your Data Build and Improve
                  Itself via Goal-Driven Loop Agents},
  author       = {Qisong Zhang and Wenzhuo Wu and Zhuangzhuang Jia
                  and Yunhao Yang and Huayu Zhang and Xianghao Zang
                  and Zhixiang He and Zhongjiang He and Kongming Liang
                  and Zhanyu Ma},
  year         = {2026},
  eprint       = {2605.01789},
  archivePrefix= {arXiv},
  primaryClass = {cs.AI},
  url          = {https://arxiv.org/abs/2605.01789}
}

Honors & Recognition

Public recognition for DataEvolver across academic workshop review and agent-for-science competition presentation.

🏆 Oral Presentation

AIDataSci 2026, KDD 2026 Workshop

DataEvolver was accepted as an Oral Presentation for the AIDataSci 2026 workshop.

🎖️ Poster Presentation

2026 BAAI Conference Agent for Science Competition

DataEvolver received a Certificate of Poster Presentation and was selected for on-site poster presentation.

DataEvolver BAAI Agent for Science poster presentation certificate
BAAI Agent for Science Competition Certificate of Poster Presentation, 2026.

Ready to Build Research-Grounded Data Pipelines?

Clone the repository, run the safe dry-run, and choose either the multimodal generation route or the WebSearch-guided research route.

Get Started on GitHub