We just released two new encoder models (MLM) in 2026 ๐๐๐
They're super fast, easy to train, and strongly multilingual.
Try them today: we created 5 demos on @huggingface
We're open-sourcing a new benchmark on @huggingface: IFStruct
The goal is to measure output validity and schema following across diverse prompts.
We also detail how we generated it and the design choices behind it in our blog post.
Fun surprise: DeepSeek used my open-perfectblend dataset to train their new DSpark drafter
Time to promote it again! It's an open-source reproduction of "The Perfect Blend" paper.
If you ever need >1M diverse prompts in math, chat, and code, it does the job.