Overview of DataFlow-Instruct-10K Dataset
mainDataFlow-Instruct-10K is a unified multi-domain instruction dataset generated using the DataFlow framework. It is constructed through automated data preparation pipelines covering mathematical reasoning, code, and general text instructions.
Each pipeline follows a “generate-evaluate-filter-refine” workflow to synthesize high-quality "instruction-response" pairs. The dataset contains approximately 10K samples designed to enable base models to achieve performance levels comparable to full-scale instruction-tuned models with significantly fewer training samples.