Files
wehub-resource-sync 593b94c120
pytest / Unit Tests (push) Has been cancelled
pytest / Integration (integration_tests_a) (push) Has been cancelled
pytest / Integration (integration_tests_b) (push) Has been cancelled
pytest / Integration (integration_tests_c) (push) Has been cancelled
pytest / Integration (integration_tests_d) (push) Has been cancelled
pytest / Integration (integration_tests_e) (push) Has been cancelled
pytest / Integration (integration_tests_f) (push) Has been cancelled
pytest / Integration (integration_tests_g) (push) Has been cancelled
pytest / Integration (integration_tests_h) (push) Has been cancelled
pytest / Integration (integration_tests_i) (push) Has been cancelled
pytest / Integration (integration_tests_j) (push) Has been cancelled
pytest / Distributed (distributed_a) (push) Has been cancelled
pytest / Distributed (distributed_b) (push) Has been cancelled
pytest / Distributed (distributed_c) (push) Has been cancelled
pytest / Distributed (distributed_d) (push) Has been cancelled
pytest / Distributed (distributed_e) (push) Has been cancelled
pytest / Distributed (distributed_f) (push) Has been cancelled
pytest / Minimal Install (push) Has been cancelled
pytest / Event File (push) Has been cancelled
pytest (slow) / py-slow (push) Has been cancelled
Publish JSON Schema / publish-schema (push) Has been cancelled
chore: import upstream snapshot with attribution
2026-07-13 12:49:20 +08:00

76 lines
2.5 KiB
Markdown

# Optimizer Comparison: Schedule-Free, Muon, Adafactor, and More
[![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/ludwig-ai/ludwig/blob/main/examples/optimizers/optimizer_comparison.ipynb)
## Why optimizer choice matters
The optimizer is more than a training detail — it controls how fast gradients
are translated into weight updates, whether training is stable in early epochs,
how much memory the optimizer state consumes, and whether you need to tune a
separate learning-rate schedule at all.
Ludwig 0.11 added five production-ready optimizers beyond the classic Adam/SGD
family: **RAdam**, **Adafactor**, **Schedule-Free AdamW**, **Muon**, and **SOAP**.
This example shows how to configure each one and compares them on a real dataset.
## What this example shows
- How to set `trainer.optimizer.type` in a Ludwig YAML config
- The one rule for Schedule-Free AdamW: no `learning_rate_scheduler`
- Side-by-side training curves (validation loss + accuracy) for all optimizers
- A summary table of final metrics and wall-clock training time
## Prerequisites
```bash
pip install ludwig
```
No GPU required. The notebook runs on CPU in a few minutes.
## Quick start
### Run the notebook (recommended)
Open [`optimizer_comparison.ipynb`](optimizer_comparison.ipynb) in Jupyter or
click the Colab badge above.
### Run the script
```bash
python optimizer_comparison.py
```
This downloads the UCI Wine Quality dataset, trains all five configs, and
prints a comparison table.
### Use a standalone YAML config
Each optimizer has its own config file you can use directly with the Ludwig CLI:
```bash
ludwig train --config config_schedule_free_adamw.yaml --dataset winequality-red.csv
```
| File | Optimizer |
| --------------------------------- | ------------------- |
| `config_adamw.yaml` | AdamW (baseline) |
| `config_radam.yaml` | RAdam |
| `config_adafactor.yaml` | Adafactor |
| `config_schedule_free_adamw.yaml` | Schedule-Free AdamW |
| `config_muon.yaml` | Muon |
## Key insight: Schedule-Free AdamW needs no LR scheduler
```yaml
trainer:
optimizer:
type: schedule_free_adamw
lr: 0.001
# Do NOT add learning_rate_scheduler here
```
Adding a `learning_rate_scheduler` on top of `schedule_free_adamw` fights the
built-in schedule and hurts convergence. See the notebook for a detailed
explanation.