New set of LFM2.5-Encoders just dropped.
These are super fast, even at long context and on CPU.
Our team adapted them from their LFM2.5 backbones:
1. Replaced the causal attention mask with a bidirectional one
2. Made the LFM short convolutions non-causal
3. Trained with a masked language modeling objective
The result is a set of two tiny, super fast encoders.
Encoding a document of 12 to 15 pages (about 8k tokens) takes less than 30s on a CPU.
Release blog:
Models on Hugging Face: