Understand the RLHF dataset format
mainThe RLHF datasets are stored in Parquet files. Data is organized using a chat-based format within the prompt field to support multi-turn conversations. To facilitate answer extraction, instruction-following text is often included directly in the prompt (e.g., asking the model to output the final answer after a specific delimiter like ####).
Key fields in the dataset schema include:
data_source: The origin of the data (e.g.,openai/gsm8k).prompt: A list of message objects containingroleandcontent(standard chat format).ability: The capability being tested (e.g.,math).reward_model: A configuration object defining how the model is evaluated. It includes:style: The reward calculation method (e.g.,rule).ground_truth: A list of expected correct answers.
{
"data_source": "openai/gsm8k",
"prompt": [{"role": "user", "content": "Natalia sold clips to 48 of her friends in April, and then she sold half as many clips in May. How many clips did Natalia sell altogether in April and May? Let's think step by step and output the final answer after \"####\""}],
"ability": "math",
"reward_model": {
"style": "rule",
"ground_truth": ["72"]
}
}