Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
💼
Hiring
254.0
TFLOPS
Dan Petrovic
PRO
dejanseo
44
6
100
Follow
adamfallon's profile picture
Mi6paulino's profile picture
Aifanatic123's profile picture
78 followers
·
16 following
https://dejan.ai/
dejanseo
dejanmarketing
seoguy
dejanseo.bsky.social
AI & ML interests
DEJAN - The AI SEO agency. [DEJAN](https://dejan.ai/)
Recent Activity
posted
an
update
about 11 hours ago
10 Search Subqueries in 200 Microseconds: 1-Bit Consistency Fanout We built Fanout Diffusion, an ultra-low-latency model that expands a search query into 10 diverse subquery vectors in a single forward pass. Inspired by the continuous retrieval framework in R4T (arXiv:2603.06397), our goal was to make query expansion fast enough for production search without running language models. How it works: • Direct Embedding Space: Takes a query embedded via google/embeddinggemma-300m (768-dim) and predicts 10 distinct subquery vectors simultaneously. • 1-Step Consistency Denoiser: Generates all 10 slots analytically in a single pass with zero ODE integration loops. • 1-Bit Hardware Tensor Cores: Quantized ternary weights (-1, 0, +1) running on Ampere/Ada sub-byte PTX instructions with INT4 outer projections. Key numbers on an RTX 4090: • Latency: 0.200 ms (ONNX) / 0.329 ms (native C++ CUDA graph) • Throughput: 417,035 queries/second in batch mode • Model Size: 1.64 MB (PyTorch QAT) / 32.1 MB (ONNX graph) • Quality: 0.683 prompt alignment across 540k search queries Interactive Space: https://huggingface.co/spaces/dejanseo/fanout-diffusion Model weights, ONNX graph, and C++ engine: https://huggingface.co/dejanseo/fanout-diffusion
updated
a Space
about 11 hours ago
dejanseo/fanout-diffusion
updated
a model
about 11 hours ago
dejanseo/fanout-diffusion
View all activity
Organizations
dejanseo
's models
61
Sort: Recently updated
dejanseo/fanout-diffusion
Updated
about 11 hours ago
dejanseo/chrome_models
Updated
Aug 13
•
21.7k
•
12
dejanseo/LinkjeBERT
Token Classification
•
0.3B
•
Updated
Jul 5
•
16
•
1
dejanseo/LinkjeBERT4
0.3B
•
Updated
Jul 5
•
12
dejanseo/LinkjeBERT3
0.3B
•
Updated
Jul 5
•
9
dejanseo/LinkjeBERT2
0.3B
•
Updated
Jul 3
•
10
dejanseo/gemotions
Updated
May 17
•
5
dejanseo/latent-entity
0.3B
•
Updated
Mar 24
•
5
•
1
dejanseo/reverse-prompter
Text Generation
•
0.3B
•
Updated
Mar 18
•
32
•
3
dejanseo/ecommerce-query-volume-classifier
Text Classification
•
0.2B
•
Updated
Mar 12
•
26
•
1
dejanseo/google-links
Token Classification
•
0.3B
•
Updated
Feb 16
•
119
•
9
dejanseo/ecommerce-taxonomy-classifier
Text Classification
•
Updated
Nov 20, 2025
•
4
dejanseo/Intent-XS
Text Classification
•
11.7M
•
Updated
Nov 20, 2025
•
32
•
3
dejanseo/LinkBERT
Token Classification
•
Updated
Nov 20, 2025
•
389
•
9
dejanseo/confidence-distribution-threshold-detector
Updated
Oct 10, 2025
•
1
dejanseo/kt
0.3B
•
Updated
Sep 17, 2025
•
3
dejanseo/link-prediction
Token Classification
•
Updated
Aug 14, 2025
•
23
•
3
dejanseo/query-fanout
Text Generation
•
1B
•
Updated
Aug 9, 2025
•
27
•
1
dejanseo/query-reformulation
60.5M
•
Updated
Jul 7, 2025
•
7
dejanseo/query-reformulator-large
0.2B
•
Updated
Jul 7, 2025
•
3
dejanseo/gemma-embed-large
Updated
Jul 6, 2025
dejanseo/gemma-embed
Feature Extraction
•
Updated
Jun 28, 2025
•
2
dejanseo/gemma-embed-stage-3
Updated
Jun 27, 2025
•
6
dejanseo/universal-query-classifier-base
0.2B
•
Updated
Jun 27, 2025
•
7
dejanseo/gemma-embed-stage-2
Updated
Jun 27, 2025
•
4
dejanseo/gemma-embed-stage-1
Updated
Jun 27, 2025
•
6
dejanseo/universal-query-classifier-large
0.4B
•
Updated
Jun 17, 2025
•
6
dejanseo/universal-query-classifier-xsmall
70.6M
•
Updated
Jun 16, 2025
•
7
dejanseo/universal-query-classifier-small
0.1B
•
Updated
Jun 16, 2025
•
6
dejanseo/ucq
0.4B
•
Updated
Jun 14, 2025
•
10
Previous
1
2
3
Next