Prepare SFT data for Speaker Completion
mainSpeaker Completion data allows the model to adapt to a specific speaker's style by providing a portion of the speaker's audio as context.
Implementation Details:
- The input should include a portion of the speaker's audio (typically 5-10 seconds).
- It is recommended to use varying segment lengths rather than a fixed length to improve robustness.
Input Format: Includes the text prompt followed by a segment of the speaker's audio tokens.
<|im_start|>
<|text_start|>text<|period|><|text_end|>
<|audio_start|>
[Partial Audio Tokens]Target Format: The remaining audio tokens required to complete the sequence.
[Remaining Audio Tokens]
<|audio_end|>
<|im_end|><|im_start|>
<|text_start|>this<|space|>is<|space|>a<|space|>test<|period|><|text_end|>
<|audio_start|>
this<|t_0.15|><|27|><|1789|><|379|><|1236|><|1465|><|1326|><|1584|><|889|><|183|><|1283|><|794|><|space|>
is<|t_0.09|><|1281|><|903|><|1521|><|319|><|230|><|1533|><|906|><|space|>