Register and share your invite link to earn from video plays and referrals.

tomaarsen
@tomaarsen
Sentence Transformers, SetFit & NLTK maintainer Machine Learning Engineer at 🤗 Hugging Face
473 Following    4.8K Followers
Very cool: Tencent has updated the license of their SOTA multimodal embedding models to Apache-2.0! Tencent released WeMM-Embedding this week with a technical report; it's the top trending paper on Papers with Code The models outperform Gemini Embedding 2, Qwen3-VL Embedding, and @VoyageAI embedding models
Show more
BIG ANNOUNCEMENT FROM HUGGING FACE TODAY: We're unveiling Microduck 🐥🤖 It's a tiny $399 open-source robot you can teach new tricks with reinforcement learning. It can walk, pick things up, get back up when it falls, and even roller-skate. Welcome to the era of open-source affordable robots to democratize physical AI and world models! 🤗🤗🤗
Show more
0
623
12K
983
Forward to community
Summer of open-source ain't over yet 🔥🔥🔥
📈 New blog post: Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers. As a practical example, I finetuned a ColBERT-style model for medical retrieval. 14.5 hours on one RTX 3090, and it beats every general-purpose retriever I could find. Thread 🧵
Show more
Regarding Sentence Transformers v6.0 update: A dense model compresses a whole text into one vector, then compares two vectors. A multi-vector model keeps one vector per token, scores every query token against every document token, takes the best match for each, and sums those.
Show more
@jeremyphoward @answerdotai He's right! 200k monthly downloads and not for nothing, here's the link for those who want to try it:
🚨I've just released Sentence Transformers v6.0! MultiVectorEncoder joins the family: ColBERT-style late interaction models are now a first-class model type, for training, inference & interpretation, alongside dense, sparse & reranker models. Big thread 🧵
Show more