VL / Multimodal plans for Ling-3.0-tiny?

#16
by Sariel00 - opened

Hi team,
Thanks for open-sourcing Ling-3.0-tiny! The 1.3B active parameter MoE design is impressive for edge deployment.
I'm wondering if there are any plans to release a Vision-Language (VL) or multimodal variant of Ling-3.0-tiny in the future? Given its lightweight architecture, a VL version would be very attractive for on-device multimodal applications.
Also, has anyone in the community tried integrating Ling-3.0-tiny with external vision encoders (e.g., SigLIP, InternVL) or building a custom VL pipeline on top of it? I'd love to hear about your experiences and any challenges you've encountered.
Looking forward to the discussion!

Sign up or log in to comment