Hyperparameter tuning and best practices
mainWhen using Schedule-Free learning, consider the following guidance:
- Learning Rate Warmup: Highly recommended. Use the
warmup_stepsparameter. - Momentum ($\beta$): Training is sensitive to $\beta$. The default is $0.9$, but for very long training runs, you may need to increase this to $0.95$ or $0.98$.
- SGD Learning Rates: A good starting point is $10x$ to $50x$ larger than classical rates.
- AdamW Learning Rates: A good starting point is $1x$ to $10x$ larger than schedule-based approaches.
- Regularization: This method requires tuning; it may not outperform scheduled approaches without proper tuning of regularization and learning rate parameters.