AI & ML interests

The AI community building the future.

Recent Activity

nielsrΒ  updated a dataset 22 minutes ago
huggingface/documentation-images
nielsrΒ  updated a Space about 1 hour ago
huggingface/paperswithcode
View all activity

Articles

Add MVE train images

#654 opened about 1 hour ago by
tomaarsen
nielsrΒ 
updated a bucket about 7 hours ago
tarekziadeΒ 
updated a bucket about 10 hours ago
tomaarsenΒ 
posted an update 7 days ago
view post
Post
3560
🚨 I've just published Sentence Transformers v6.0, introducing MultiVectorEncoder: ColBERT-style late interaction models are now a fourth model type, for training, inference, and interpretation, alongside the dense, sparse, and reranker models! Details:

Where a regular embedding model compresses a whole text into one vector, a multi-vector model keeps one vector per token and scores query against document with the MaxSim operator. That preserves token-level matching information that a single vector has to average away. It is also the state of the art for visual document retrieval, where a text query is matched against page images directly, charts and tables included, with no OCR step in between.

Any PyLate, Stanford ColBERT, or ColPali checkpoint loads straight into the same familiar API: model.encode_query(), model.encode_document(), and model.similarity() just work, whether the documents are texts or page images.

Does it help? LightOn trained LateOn (multi-vector) and DenseOn (dense) on the same data with the same 149M ModernBERT backbone, and the multi-vector model wins on 9 of the 13 NanoBEIR datasets: 0.6868 vs 0.6764 mean NDCG@10. The price is a bigger index, and the new HierarchicalTokenPooling module halves it at roughly no retrieval cost.

Antoine Chaffin, RaphaΓ«l Sourty, and I wrote a blog post walking through multi-vector models in practice: loading the various checkpoint formats, encoding and scoring, plugging them into a search stack, running them on page images, and keeping the index affordable. Check it out if you want to get started, or just point your Agent to the URL: https://huggingface.co/blog/multi-vector-encoder

pip install sentence-transformers==6.0.0

Release notes: https://github.com/huggingface/sentence-transformers/releases/tag/v6.0.0
julien-cΒ 
posted an update 28 days ago
view post
Post
4408
who's working on an NVFP4 version of Kimi-K3?
  • 4 replies
Β·
badaouiΒ 
posted an update about 1 month ago
view post
Post
2381
432 GB of ultra-fast HBM4 and up to 23.3 TB/s of memory bandwidth on a single GPU 🀯.

Two weeks ago, we got early access to AMD's new Instinct MI455X, and our first goal was simple: make sure πŸ€— Transformers works on day one.

Over the past few weeks, we worked closely with the AMD team to validate the platform, enable Flash Attention, add torchcodec support for multimodal models, and resolve issues uncovered during testing.

The result:
βœ… 99.5% success rate across our 24 core Transformers model architectures - already on par with our daily CI on previous AMD and NVIDIA platforms.

The hardware is just as exciting. With 432 GB of HBM per GPU, our early capacity experiments showed more than 3Γ— the concurrent long-context requests compared to MI300, thanks to the much larger KV cache capacity.

A huge thanks to the AMD team for the early access and the great collaboration!

Read the full blog πŸ‘‡
https://huggingface.co/blog/badaoui/transformers-on-amd-mi455
  • 1 reply
Β·