Atif Saleem
atifsal
AI & ML interests
AI safety and security research and its super-alignment with human race by using sup-intelligence that follows ethics, compliance and mimics human emotions in real time with empathy. I recently have been interested in Quantum computing and Molecular computing for AI where efficient low energy computing is leveraged to develop AI agents and Robots for our everyday use.
Recent Activity
liked a model about 14 hours ago
empero-ai/Qwythos-9B-Claude-Mythos-5-1M-GGUF liked a model about 14 hours ago
LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V6-GGUF updated a collection about 14 hours ago
STT-TTS-Audio-ModelsOrganizations
VTON-Papers
-
FW-VTON: Flattening-and-Warping for Person-to-Person Virtual Try-on
Paper • 2507.16010 • Published -
Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off
Paper • 2508.04825 • Published • 60 -
ART-VITON: Measurement-Guided Latent Diffusion for Artifact-Free Virtual Try-On
Paper • 2509.25749 • Published • 1 -
MC-VTON: Minimal Control Virtual Try-On Diffusion Transformer
Paper • 2501.03630 • Published
World-Models
Embeddings
Text-to-Image
ComfyUI-Models-Workflows
Fashion-Models
Image-to-Image_Models
Video-Text_to_Text_Models
Image-to-Video_Models
Vision-Models
VTON_Models
AI-Datasets
Research-Papers
-
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Paper • 2404.03413 • Published • 28 -
RepVideo: Rethinking Cross-Layer Representation for Video Generation
Paper • 2501.08994 • Published • 15 -
Hierarchical Cross-modal Prompt Learning for Vision-Language Models
Paper • 2507.14976 • Published • 2
VTON-Datatsets
Multimodals
Reasoning_Models
Rerankers
STT-TTS-Audio-Models
Text-to-Video_Models
Graph-Learning_Models
Audio-Text-to-Text_Models
-
onnx-community/Voxtral-Mini-3B-2507-ONNX
Audio-Text-to-Text • Updated • 401 • 29 -
mmwillet2/Kokoro_GGUF
Text-to-Speech • 87.7M • Updated • 2.42k • 28 -
Qwen/Qwen3-ASR-0.6B-hf
Automatic Speech Recognition • 0.8B • Updated • 144k • 42 -
Qwen/Qwen3-ASR-0.6B
Automatic Speech Recognition • 0.9B • Updated • 2.39M • • 322
Text-Gen_Models
Any-to-Any_Models
Embedding-Models
AI-Models
-
microsoft/Orca-2-13b
Text Generation • Updated • 770 • • 668 -
SG161222/Realistic_Vision_V6.0_B1_noVAE
Text-to-Image • Updated • 34.9k • 315 - Running on ZeroAgentsFeatured83
UDOP
🏃83Generate answers or summaries from document images with prompts
-
Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca
Paper • 2304.08177 • Published • 2
Prompt-Engineering
Agentic-Models
VTON-Datatsets
VTON-Papers
-
FW-VTON: Flattening-and-Warping for Person-to-Person Virtual Try-on
Paper • 2507.16010 • Published -
Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off
Paper • 2508.04825 • Published • 60 -
ART-VITON: Measurement-Guided Latent Diffusion for Artifact-Free Virtual Try-On
Paper • 2509.25749 • Published • 1 -
MC-VTON: Minimal Control Virtual Try-On Diffusion Transformer
Paper • 2501.03630 • Published
Multimodals
World-Models
Reasoning_Models
Embeddings
Rerankers
Text-to-Image
STT-TTS-Audio-Models
ComfyUI-Models-Workflows
Text-to-Video_Models
Fashion-Models
Graph-Learning_Models
Image-to-Image_Models
Audio-Text-to-Text_Models
-
onnx-community/Voxtral-Mini-3B-2507-ONNX
Audio-Text-to-Text • Updated • 401 • 29 -
mmwillet2/Kokoro_GGUF
Text-to-Speech • 87.7M • Updated • 2.42k • 28 -
Qwen/Qwen3-ASR-0.6B-hf
Automatic Speech Recognition • 0.8B • Updated • 144k • 42 -
Qwen/Qwen3-ASR-0.6B
Automatic Speech Recognition • 0.9B • Updated • 2.39M • • 322
Video-Text_to_Text_Models
Text-Gen_Models
Image-to-Video_Models
Any-to-Any_Models
Vision-Models
Embedding-Models
VTON_Models
AI-Models
-
microsoft/Orca-2-13b
Text Generation • Updated • 770 • • 668 -
SG161222/Realistic_Vision_V6.0_B1_noVAE
Text-to-Image • Updated • 34.9k • 315 - Running on ZeroAgentsFeatured83
UDOP
🏃83Generate answers or summaries from document images with prompts
-
Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca
Paper • 2304.08177 • Published • 2
AI-Datasets
Prompt-Engineering
Research-Papers
-
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Paper • 2404.03413 • Published • 28 -
RepVideo: Rethinking Cross-Layer Representation for Video Generation
Paper • 2501.08994 • Published • 15 -
Hierarchical Cross-modal Prompt Learning for Vision-Language Models
Paper • 2507.14976 • Published • 2