Qwen3.6 also excels at classification. This dataset was synthesized using the 27B variant of the model series, with approximately 28K image samples and their corresponding guardrail classification pairs.
Prithiv Sakthi PRO
AI & ML interests
computer vision, nlp, multimodality
- HuggingFace Fellow ML ๐ค
Recent Activity
updated a Space about 1 hour ago
prithivMLmods/Wan2.2-Fast updated a Space about 4 hours ago
prithivMLmods/FireRed-Image-Edit-1.0-Fast new activity about 4 hours ago
prithivMLmods/Qwen-Image-Edit-2511-LoRAs-Fast:update requirementsOrganizations
replied to their post 1 day ago
Post
2263
ImageShield-MMCF โ Multimodal Content Filter is a multimodal content-safety classifier built on top of Qwen3.5 and is now available on Hugging Face!
This is the preview initial version (v1.0) of the model, designed to classify visual content as Safe or Unsafe, with a particular focus on detecting Non-Consensual Intimate Imagery (NCII) and other potentially sensitive visual content.
The demo is implemented in the prithivMLmods/opencaption-4b-vl-sft Space, which serves as an active content-safety layer for computer vision tasks. It helps block NCII content generation and paves the way for more meaningful and responsible creativity.
โน ImageShield-MMCF-0.8B: prithivMLmods/ImageShield-MMCF-0.8B
โน ImageShield-MMCF-2B: prithivMLmods/ImageShield-MMCF-2B
This is the preview initial version (v1.0) of the model, designed to classify visual content as Safe or Unsafe, with a particular focus on detecting Non-Consensual Intimate Imagery (NCII) and other potentially sensitive visual content.
The demo is implemented in the prithivMLmods/opencaption-4b-vl-sft Space, which serves as an active content-safety layer for computer vision tasks. It helps block NCII content generation and paves the way for more meaningful and responsible creativity.
โน ImageShield-MMCF-0.8B: prithivMLmods/ImageShield-MMCF-0.8B
โน ImageShield-MMCF-2B: prithivMLmods/ImageShield-MMCF-2B
posted an update 2 days ago
Post
2263
ImageShield-MMCF โ Multimodal Content Filter is a multimodal content-safety classifier built on top of Qwen3.5 and is now available on Hugging Face!
This is the preview initial version (v1.0) of the model, designed to classify visual content as Safe or Unsafe, with a particular focus on detecting Non-Consensual Intimate Imagery (NCII) and other potentially sensitive visual content.
The demo is implemented in the prithivMLmods/opencaption-4b-vl-sft Space, which serves as an active content-safety layer for computer vision tasks. It helps block NCII content generation and paves the way for more meaningful and responsible creativity.
โน ImageShield-MMCF-0.8B: prithivMLmods/ImageShield-MMCF-0.8B
โน ImageShield-MMCF-2B: prithivMLmods/ImageShield-MMCF-2B
This is the preview initial version (v1.0) of the model, designed to classify visual content as Safe or Unsafe, with a particular focus on detecting Non-Consensual Intimate Imagery (NCII) and other potentially sensitive visual content.
The demo is implemented in the prithivMLmods/opencaption-4b-vl-sft Space, which serves as an active content-safety layer for computer vision tasks. It helps block NCII content generation and paves the way for more meaningful and responsible creativity.
โน ImageShield-MMCF-0.8B: prithivMLmods/ImageShield-MMCF-0.8B
โน ImageShield-MMCF-2B: prithivMLmods/ImageShield-MMCF-2B
reacted to GoktugD's post with โค๏ธ 2 days ago
Post
2292
๐น๐ท We trained a 1B OCR model specifically for Turkish enterprise documents.
**Werea-DocOCR-1B v2**
The result surprised us:
LightOnOCR-2 base โ **64.2% CER**
Werea-DocOCR v1 โ **~8.1% CER**
Werea-DocOCR v2 โ **0.15% CER** ๐
Evaluated on a held-out 72-page test set across 12 Turkish document types and 3 different capture conditions.
๐ 12 Turkish enterprise document types
๐งช 12,960 synthetic training pages
๐ฑ Digital + scanned + phone photos
๐ Tables โ structured Markdown
โ๏ธ Full-parameter fine-tuning
๐ฅ๏ธ Trained on a single RTX 3090
It handles:
โข e-Invoices
โข rental contracts
โข bank receipts
โข payroll documents
โข insurance policies
โข vehicle documents
โข official correspondence
โข trade registry documents
โข SGK-style tables
โข and more.
**Model ๐ค**
Werea-co/Werea-DocOCR-1B
**Dataset ๐**
Werea-co/werea-tr-doc-ocr-enterprise-v2
**Werea ๐น๐ท**
Werea-co
We're building open AI models from Tรผrkiye.
This is just the beginning.
#HuggingFace #OCR #DocumentAI #TurkishAI #OpenSourceAI #ComputerVision
**Werea-DocOCR-1B v2**
The result surprised us:
LightOnOCR-2 base โ **64.2% CER**
Werea-DocOCR v1 โ **~8.1% CER**
Werea-DocOCR v2 โ **0.15% CER** ๐
Evaluated on a held-out 72-page test set across 12 Turkish document types and 3 different capture conditions.
๐ 12 Turkish enterprise document types
๐งช 12,960 synthetic training pages
๐ฑ Digital + scanned + phone photos
๐ Tables โ structured Markdown
โ๏ธ Full-parameter fine-tuning
๐ฅ๏ธ Trained on a single RTX 3090
It handles:
โข e-Invoices
โข rental contracts
โข bank receipts
โข payroll documents
โข insurance policies
โข vehicle documents
โข official correspondence
โข trade registry documents
โข SGK-style tables
โข and more.
**Model ๐ค**
Werea-co/Werea-DocOCR-1B
**Dataset ๐**
Werea-co/werea-tr-doc-ocr-enterprise-v2
**Werea ๐น๐ท**
We're building open AI models from Tรผrkiye.
This is just the beginning.
#HuggingFace #OCR #DocumentAI #TurkishAI #OpenSourceAI #ComputerVision
Post
5220
The Qwen3.8 27B demo for object grounding is now available on Hugging Face Spaces.
It features three tasks: Object Detection (Bounding Boxes), Point Localization (Keypoints), and Spatial Guidance (Path Mapping).
Try it now: prithivMLmods/Qwen3.8-27B-Object-Detection
It features three tasks: Object Detection (Bounding Boxes), Point Localization (Keypoints), and Spatial Guidance (Path Mapping).
Try it now: prithivMLmods/Qwen3.8-27B-Object-Detection
posted an update 11 days ago
Post
5220
The Qwen3.8 27B demo for object grounding is now available on Hugging Face Spaces.
It features three tasks: Object Detection (Bounding Boxes), Point Localization (Keypoints), and Spatial Guidance (Path Mapping).
Try it now: prithivMLmods/Qwen3.8-27B-Object-Detection
It features three tasks: Object Detection (Bounding Boxes), Point Localization (Keypoints), and Spatial Guidance (Path Mapping).
Try it now: prithivMLmods/Qwen3.8-27B-Object-Detection
Post
5522
Made a demo for Text/Image-to-3D Video and Image-to-3D Video asset generation using TRELLIS.2. It is paired with Z-Image-Turbo to accelerate the input image preprocessing pipeline, streamlining the Image-to-3D workflow. The generated GLB (GL Transmission Format) files are converted into MP4 (MPEG-4) videos, making them easy to preview and share. Try it now on Hugging Face Spaces.๐ค
โ Image-to-3D-Video-Asset-Generator: prithivMLmods/Image-to-3D-Video-Asset-Generator
โ collection: https://huggingface.co/collections/prithivMLmods/multimodal-implementations
โ github: https://github.com/PRITHIVSAKTHIUR/Image-to-3D-Video-Asset-Generator
โคท To learn more, visit the app page or the respective model pages.
โ Image-to-3D-Video-Asset-Generator: prithivMLmods/Image-to-3D-Video-Asset-Generator
โ collection: https://huggingface.co/collections/prithivMLmods/multimodal-implementations
โ github: https://github.com/PRITHIVSAKTHIUR/Image-to-3D-Video-Asset-Generator
โคท To learn more, visit the app page or the respective model pages.
posted an update about 1 month ago
Post
5522
Made a demo for Text/Image-to-3D Video and Image-to-3D Video asset generation using TRELLIS.2. It is paired with Z-Image-Turbo to accelerate the input image preprocessing pipeline, streamlining the Image-to-3D workflow. The generated GLB (GL Transmission Format) files are converted into MP4 (MPEG-4) videos, making them easy to preview and share. Try it now on Hugging Face Spaces.๐ค
โ Image-to-3D-Video-Asset-Generator: prithivMLmods/Image-to-3D-Video-Asset-Generator
โ collection: https://huggingface.co/collections/prithivMLmods/multimodal-implementations
โ github: https://github.com/PRITHIVSAKTHIUR/Image-to-3D-Video-Asset-Generator
โคท To learn more, visit the app page or the respective model pages.
โ Image-to-3D-Video-Asset-Generator: prithivMLmods/Image-to-3D-Video-Asset-Generator
โ collection: https://huggingface.co/collections/prithivMLmods/multimodal-implementations
โ github: https://github.com/PRITHIVSAKTHIUR/Image-to-3D-Video-Asset-Generator
โคท To learn more, visit the app page or the respective model pages.