VL / Multimodal plans for Ling-3.0-tiny?
#16
by Sariel00 - opened
Hi team,
Thanks for open-sourcing Ling-3.0-tiny! The 1.3B active parameter MoE design is impressive for edge deployment.
I'm wondering if there are any plans to release a Vision-Language (VL) or multimodal variant of Ling-3.0-tiny in the future? Given its lightweight architecture, a VL version would be very attractive for on-device multimodal applications.
Also, has anyone in the community tried integrating Ling-3.0-tiny with external vision encoders (e.g., SigLIP, InternVL) or building a custom VL pipeline on top of it? I'd love to hear about your experiences and any challenges you've encountered.
Looking forward to the discussion!