AI & ML interests

Computer Vision

Recent Activity

prithivMLmods 
posted an update 1 day ago
view post
Post
1779
Qwen-Image-2.1 Plug and Play LoRA App is now live on Hugging Face Spaces.

🔗 Space: prithivMLmods/Qwen-Image-2.1-LoRAs-PnP

It supports standard inference, 4-step Turbo inference, custom LoRA lazy repacks, and LoRA Plug and Play (PnP), all in one setting!

🔗 Qwen-Image-2.1 Image-to-Image LoRAs: https://huggingface.co/collections/prithivMLmods/qwen-image-21-image-to-image-loras

🔗 GitHub: https://github.com/PRITHIVSAKTHIUR/Qwen-Image-2.1-LoRAs-PnP

To learn more, visit the app page or the respective model pages.
prithivMLmods 
posted an update 9 days ago
view post
Post
789
VisionGuardrail EVO-2, a multimodal image-classification content-safety model based on Qwen/Qwen3.8-27B, is now available on the Hub!

Stricter image classification than before, with a dense 27-billion-parameter multimodal model, more precise reasoning, and improved captions for classifying visual media.

➠ Models: prithivMLmods/VisionGuardrail-Evo2-27B, prithivMLmods/VisionGuardrail-Evo2-27B-GGUF

➠ Collection: https://huggingface.co/collections/prithivMLmods/visionguardrail-evo2

➠ Previous Models: https://huggingface.co/collections/prithivMLmods/visionguardrail-collection

⤷ To learn more, visit the app page or the respective model pages.
prithivMLmods 
posted an update 13 days ago
prithivMLmods 
posted an update 20 days ago
view post
Post
3818
VisionGuardrail, a multimodal content-safety classifier based on Qwen3.5, is now available on Hugging Face in 4B and 9B variants. It is a direct upgrade to ImageShield-MMCF, providing improved parental controls through conservative visual content-safety filtering.

More About:
➠ hf.co/blog — https://huggingface.co/blog/prithivMLmods/vision-guardrail-mini-blog

➠ Models:
✦ VisionGuardrail-4B: prithivMLmods/VisionGuardrail-4B
✦ VisionGuardrail-9B: prithivMLmods/VisionGuardrail-9B

➠ Dataset:
✦ ImageShield-Guardrail-Pro: prithivMLmods/ImageShield-Guardrail-Pro

⤷ To learn more, visit the app page or the respective model pages.
prithivMLmods 
posted an update 29 days ago
view post
Post
3090
ImageShield-MMCF — Multimodal Content Filter is a multimodal content-safety classifier built on top of Qwen3.5 and is now available on Hugging Face!

This is the preview initial version (v1.0) of the model, designed to classify visual content as Safe or Unsafe, with a particular focus on detecting Not Safe for Work (NSFW) and other potentially sensitive visual content.

The demo is implemented in the prithivMLmods/opencaption-4b-vl-sft Space, which serves as an active content-safety layer for computer vision tasks. It helps block Not Safe for Work (NSFW) content generation and paves the way for more meaningful and responsible creativity.

⊹ ImageShield-MMCF-0.8B: prithivMLmods/ImageShield-MMCF-0.8B
⊹ ImageShield-MMCF-2B: prithivMLmods/ImageShield-MMCF-2B
  • 2 replies
·
prithivMLmods 
posted an update about 1 month ago
view post
Post
5279
The Qwen3.8 27B demo for object grounding is now available on Hugging Face Spaces.

It features three tasks: Object Detection (Bounding Boxes), Point Localization (Keypoints), and Spatial Guidance (Path Mapping).

Try it now: prithivMLmods/Qwen3.8-27B-Object-Detection
Nymbo 
posted an update about 2 months ago
view post
Post
2296
Anthropic gave me six months of Claude Max 20x through the Claude for Open Source program, granted based on my Hugging Face work. Thank you
Anthropic
for supporting open source.

So far I've been pointing it at Markdown Minimap, an Obsidian plugin that adds a scrollable IDE-style minimap to your notes. This week I've been clearing a backlog of user-reported issues on it, with Claude often handling them end to end.

https://github.com/Nymbo/Markdown-Minimap — issues and PRs welcome.
prithivMLmods 
posted an update about 2 months ago
view post
Post
5561
Made a demo for Text/Image-to-3D Video and Image-to-3D Video asset generation using TRELLIS.2. It is paired with Z-Image-Turbo to accelerate the input image preprocessing pipeline, streamlining the Image-to-3D workflow. The generated GLB (GL Transmission Format) files are converted into MP4 (MPEG-4) videos, making them easy to preview and share. Try it now on Hugging Face Spaces.🤗

➠ Image-to-3D-Video-Asset-Generator: prithivMLmods/Image-to-3D-Video-Asset-Generator
➠ collection: https://huggingface.co/collections/prithivMLmods/multimodal-implementations
➠ github: https://github.com/PRITHIVSAKTHIUR/Image-to-3D-Video-Asset-Generator

⤷ To learn more, visit the app page or the respective model pages.
Nymbo 
posted an update 2 months ago
view post
Post
6011
Introducing Inflect-v2, two exceptionally small, open-weight English TTS models at just 3.9M and 9.3M parameters. Both generate speech multiple times faster than real-time on CPU. Despite their size, Inflect-v2 delivers quality that is competitive with much larger lightweight TTS systems, including KittenTTS, Piper, and Supertonic-3.

CPU, CUDA, PyTorch, and ONNX are supported. Apache 2.0.

See it for yourselves:
owensong/Inflect-Micro-v2
owensong/Inflect-Nano-v2

Try the Demos:
Nymbo/Inflect-TTS (unlimited CPU usage)
owensong/Inflect-v2 (ultra-fast ZeroGPU usage)
  • 6 replies
·