btw fine-tuning open-source models at 2x the speed is easy now. you just need:
- your favorite base model (Llama, Qwen, Mistral)
- a single YAML config file
- Halo from
@whitecircle
up to 2.8x faster throughput, lower peak VRAM, and 0 checkpoint conversion scripts