Krea 2 Identity Edit
Identity-preserving instruction image editing on Krea 2
Apps started here then claimed by the ð authors
Identity-preserving instruction image editing on Krea 2
Multi-view character sheet from one image (FLUX.2 LoRA)
Efficient native-resolution image generation and editing
Hum a melody, get a finished song
Extend images into larger canvases with Krea 2 outpaint
Codec-native video & image understanding with Mage-VL 4B
Bilingual TTS with voice design, cloning, and direction
Live interactive world rollout from an image
Calibrated typed decisions + System 2 reasoning
Layout-controlled comics with a draggable bbox canvas
Text-to-image model for UI, infographics and poster design
Distilled LTX-2.3 identity video from a reference photo
Parse digital & photographed documents into Markdown
Unified audio-text intelligence
Turn a 3D mesh into a parametric CadQuery program
Edit one frame, ripple the change through the whole video
Streaming 3D reconstruction from video with local context
Multilingual CPU-only ASR with a 1.58-bit BitNet decoder
Find when a described sound happens in a recording
Realtime VLM for image and video understanding
One photo into a frozen-instant 360° camera orbit
Bilingual EN/KO speech LM - transcribe, ask, and speak
Zero-shot voice cloning TTS with Audio8 0.1B
Verbatim + intended transcripts with word-level timing
Infographic generation & editing with SenseNova-U1 V3
Studio-quality generative speech enhancement
Explore all seven Video-ORA task families
Steer a 3D camera path through any still image
8M-param English TTS distilled from Kokoro-82M
Zero-shot TTS with explicit word-level prosody control
Aerial object detection - YOLO Models
ClinFusion medical multimodal LLM for 2D images
Generative 4x video upscaling with LTX-2.3 IC-LoRA
Remove people & vehicles from video, keep the background
104M-param text-to-image DiT, trained from scratch
Embodied spatial reasoning, pointing and arm traces
Typed decision model: state in, choices + probabilities
The PlagueKind V1.5 workflow for MiniMax-H3 video + audio
Parse document images into structured Markdown
Segment and track object instances in video with QueenVIS
Depth-controlled image generation with Krea-2 Turbo
Multilingual translation across 46 languages with MiLMMT
Box an object and let Qwen-Image-2.1 erase it.
Zero-shot voice cloning TTS with 120M params
Walk through a world Zing-0.5 hallucinates live around you
Clean and normalize raw ASR transcripts with S1-mini
MobileWan text-to-video generation (Qualcomm AI Research)
Style-guided image generation with Krea 2 Turbo
Militant roots reggae LoRAs for YuE2-3B
Camera-controlled video world model with 3D-aware memory
Krea 2 Turbo HD text-to-image with enhanced VAE
1.06M-param pure-loop transformer with six effort levels
Spatial physics video generation with MiniMax-H3 LoRA
Play table tennis against a PPO agent
Animate bounding boxes into video, one prompt per box
Detection, segmentation & pose with compact edge ViTs
FLUX.2 Klein 9B fine-tune for image generation and editing
Hierarchical parallel document parsing with a 1B VLM
Zero-shot probabilistic time-series forecasting
Audio reasoning with evolving rubric rewards
Unified controllable video-to-audio Foley generation
Multilingual PII detection & LLM safety moderation
Relight exterior video clips with a light-direction ball
Voice-preserving speech dubbing across zh/en/es/ja
Native size-sensitive matroid audit
Block-diffusion zero-shot text-to-speech by Resemble AI
Tiny any-to-any multimodal GPT (text + image) prototype demo
Dense fine-grained image captioning with Qwen3-VL-4B
Zero-shot voice tuning for Kokoro-82M from a reference clip
Watch a 574K-param non-Transformer LM think
Explore, interact and converse in first-person video
Offline ASR with Confucius4-R2T2
Qwen-Image-2.1 editing with a 0.8B text encoder
4-step text-to-audio-video generation, distilled from H3
Edit images or generate from Canny edges with NK2E on Krea 2
DINOv3 feature upsampling with ViT-Up
LightOn multimodal document reranker with pointwise scoring
Cinematic product commercial style video LoRA for LTX-2.3
Zero-shot CT findings with the Jolia 3D CT foundation model
Play chess against a 32.5M-param search-free policy
Chat with a 5B model that only studied grades K-5
TinyCast zero-shot probabilistic time-series forecasting
Proactive real-time commentary on audio-video streams
Danish ASR: RoPE Whisper encoder + Qwen3 decoder
Listwise document reranker with jina-reranker-v3.5
Occult philosophy chat LLM fine-tuned on Gemma 4 12B
Compact multilingual ASR model (324M params)
Tiny 31.7M SLM that solves arithmetic expressions
Faithful x4 image super-resolution via FLUX.1-dev dual-LoRA
Generate scene-linear HDR images with Krea 2 + LogC4 LoRA
Four voices in one 10M-param CPU model, with blending
Photorealistic skin-texture LoRA for Krea 2 Turbo
Predict 3D protein structures with OpenDDE
An RL plate that balances 21 objects in MuJoCo
Minute-scale human animation with Wan2.2-Animate + LoRA
Sequence-only protein-ligand binding scorer (LULA-1.1)
Multilingual text-to-speech with zero-shot voice cloning
SAGE retrieves matching place images from a gallery
Recognize text from WordArt / artistic scene text images
Arabic speech to fully-diacritized text (tashkeel)
Bilingual streaming ASR with turn-taking prediction
Image and video understanding with MOSS-VL multimodal model
Synthesize a 3D chest CT from a radiology report
Cell Painting images to single-cell transcriptomes
One-pass GUI (element, action) decision scorer
Single-pass typed decisions with calibrated probabilities
From-scratch Glow-TTS + HiFi-GAN English voice
Roll a neural ocean emulator forward against GFDL-OM4
Multi-shot video with LayerRecall long-horizon memory
Subject-driven text-to-video from reference images (Wan2.2)
Recognize molecular structure images as E-SMILES
AlienLM privacy layer for black-box LLMs
Compact 60M T5 translator for 15 languages
Unified AR model for image understanding & generation
Zero-shot multilingual NER with GLiNER on LFM2.5-350M
Multi-speaker meeting transcription with diarization
Remove adverse weather effects (Histoformer, ECCV 2024)
Multilingual translation across 46 languages with MiLMMT
Compositional instruction-guided video editing
Few-step action-conditioned Minecraft world model
Rebuild any 3D mesh as a clean artist-style mesh
Predict where a photo was taken â full chipoint-2, both arms
Fill-mask demo for a French medical ModernBERT
Four-step 720p text-to-video preview with SANA-Video 2.0
GUI grounding & tool calling with Supertron3-0.8B
Paint inside your line art with Krea 2
Video-based quantitative physical reasoning VLM
Multimodal reasoning VLM with Grug-style thinking
Multi-shot narrated video with cross-shot memory
Define JSON tools and watch Lumma-0.6B generate tool calls
Typed questions in, calibrated probabilities out
GUI grounding & agent with a block-diffusion VLM
150-language translation with instruction following
Mask faces and license plates in street-level images
Configurable multi-stage crop disease and pest diagnosis
Detect the spoken language of a recording, 99 languages
Prompt-free all-in-one restoration (rain/haze/blur/noise)
Calibrated probabilities for decisions, no generated text
Karakalpak speech-to-text with fine-tuned Whisper Medium
Speak to a drone using an image, video, text, or voice.
Medical reasoning VLM based on Qwen3.6-27B
Agentic video understanding with active perception loop
Multimodal bio foundation model for molecules & proteins
Forecast 3D point trajectories from video and language
VLA model for scientific laboratory robotics
Distractor-free 2D enhancer for radiance field renderings
3D-centric world-spatial-action robot policy demo
Multi-image instruction-guided image editing
Encode a context into a LoRA and chat with Qwen3-8B
Draft hardware plans, review sources, and record test notes
Persian TTS demo using Ava-82M (Kokoro-based)
Detect watermarks, signatures, and artifacts in images
Proposal-only typed personal actions with a 35M-param model.
Multi-shot cinematic text-to-video generation
Generate market-grounded business ideas from images with MBA
Video verification & temporal grounding with VideoSearch-R1
Multilingual translation across 46 languages with MiLMMT
Compact OCR model for Japanese, Chinese & English text
Ask questions about traffic anomalies in video
Dictation cleanup for transcribed speech (0.8B GGUF)
VLM-guided tree-search keyframe extraction from videos
Neural rendering of 3D scenes with a transformer
8-camera multi-task driving perception, 12 heads at once
Multilingual manga speech-bubble OCR (JA/ZH/EN)
In-context classification, regression & imputation on tables
Next-frame stop-motion generation with Qwen-Image 2.1 + LoRA
Counterfactual robot futures from a world-action model
Estimate optical flow between two frames with FreeFlow
RGBA object cut-out and object removal
Voice-preserving speech dubbing across zh/en/es/ja
Raspy female rock-soul songs with YuE2 + GRVL LoRAs
Danish speech-to-text with Edda v0.1 Whisper model
Typed questions in, calibrated probabilities out
Speech-in, reasoning, tool calls, speech-out in one model
Turkish TTS via a 65M flow-matching DiT (48 kHz)
Chinese speech or video to bilingual timestamped subtitles
Persian speech recognition with BuzzASR (Whisper-large-v3)
Typed decisions (yes/no, choice, score) in one pass
Hand-drawn keyframe animation LoRAs for MiniMax-H3
A tiny transformer that runs Python programs
Hyperbolic vision-language zero-shot classification
4-step RL text-to-image with MeanFlowNFT on SD3.5-Medium
EN subtitle cues to Taiwan Traditional Chinese (0.6B model).
Forecast Valence & Arousal state change from posts
Graph-native LLM scientific reasoning with graph viz
Speaker-conditioned TTS with emotion & energy control
MrFlow training-free diffusion acceleration demo
Unified multimodal video generation and editing (5B)
Multi-agent interleaved text-image generation pipeline
P2R fine-grained visual reasoning with Qwen3-VL
Human-object interaction video from image+audio+text
Facial affect estimation with uncertainty via rectified flow
Classify time series in-context with TimEE foundation model
Novel camera viewpoint from a video via depth-warp IC-LoRA
Ultrasound image understanding VLM with MoE architecture
Training-free typographic attack defense for CLIP
Reconstruct video from event-camera voxel grids via LongE2V
4-class lung ultrasound video classifier with attention
Detect AI-generated audio-visual content (DAV-Det)
Spatial reasoning VLM for 3D relations and perspective
Synthesize contrast-enhanced breast MRI from pre-contrast
Deobfuscate and detoxify Korean text with KOTOX
Context-aware multilingual subtitle translation with Gemma 4
Martial-arts action video with a MiniMax-H3 LoRA
JoyAI-Echo on LTX-2.5, few-step video with audio
DINOv2-based semi-supervised semantic segmentation
Detect robot planning & execution failures with a VLM
Detect AI-generated images with DEAR (Dissect and Prune)
Iwin Transformer for ImageNet image classification
In-context tabular prediction from a CSV
Translate between French and 18 African languages
In-context tabular prediction from a CSV
Online video QA with event-centric memory + tool use
Camera-controlled interactive video world model
Turn one musical property up with a LoRA slider
Whispered ASMR audio-video LoRA for MiniMax-H3
Predict robot action chunks from an image + instruction
Generate electron micrographs from text prompts
Agentic multimodal reranking with visual tool use
Temporally grounded camera-motion analysis for video clips
State in, a probability distribution out. Nothing decoded.
Document page to markdown with rakedoc-nano (1.2B)
Typed questions on an image, answered with probabilities
Read Azerbaijani text lines and whole page scans
Qawwali and dark Sufi fusion LoRAs for YuE2-3B
Answer typed questions about a text in one forward pass
Zero-shot quantile forecasts for any time series
Ternary diffusion transformer on CPU via LiteRT
Streaming Persian (Farsi) speech recognition
Anime text-to-image with SilvermoonMix Anima 2.9B
Typed decisions in one forward pass â probability per answer
SOTA one-step 512Ã512 generation with FLUX.2 [klein] 4B