V2 = V1 + multimodal

#5
by Duonglv - opened

Hello Google team,

With the embeddinggemma2, I see that the general text-only benchmark is similar to v1. Only improving in code and support multimodal.
It means, if I use it for general text only, there is no significant different between these two versions, right?

We have been waiting for your new models for a long time. Hope that gemma5: 2b, 4b to 26b, 31B will appear soon. And they should be for general purpose like gemma4, please not focus on coding.

The gemma4 is smart enough and it expresses things very naturally, without the rigidity of code. This is the biggest different from other models in similar sizes. I love this one.

Thank so much.

Hi @Duonglv ,

Thanks for addressing the issue. Your observation is correct. As outlined in the official Developer Guide in Benchmarks & Evaluation section, general text performance remains consistent with version 1 (retaining the same multilingual text accuracy, with MTEB scores around 61.36 vs. 61.15 as provided in model card section Overall Evaluation Results (768d)). Version 2's major improvements are focused on code retrieval (scoring ~14% higher, moving from 68.76 to 78.68 on MTEB Code as provided in model card section Overall Evaluation Results (768d)) and native multimodality.
Thanks for the feedback related to Gemma 4 and inputs for Gemma 5!

Hello @thnamratha and the Google team,

To the extent that my "random HuggingFace user's opinion" has an effect, I would also like to offer some feedback. Similarly to @Duonglv , I think many would appreciate it if potential future Gemma 5 models remained as 'generalist' models similar to the Gemma 4 ones. If open weight LLMs can be said to possess competitive moats, I'd say that for the Gemma models this moat would be that it fulfills a 'generalist + creative writing' niche. For instance, Gemma 4 31B appears to be a much beloved model within the local LLM creative writing / RP space, and I think this is for good reason. Gemma 4 31B is pretty much my go-to for non-coding local LLM tasks. While a model like Qwen3.8 27B is excellent at coding and demonstrates remarkable intelligence per parameter, many would likely agree that Gemma 4 31B is the better model in terms of "softer" capabilities – prose quality, emotional intelligence etc. So if a future Gemma 5 lineup were simply 'codingmaxxed' then I feel like that would leave an important niche unfulfilled.

So if I were the Gemma team at Google, I would focus on preserving those strengths. That being said, I would also appreciate it if a new Gemma lineup improved the models' agentic functionality (e.g. improving tool calling further, making the models work nicer within harnesses like Pi, etc. – these sorts of things would greatly enhance the Gemma models' capabilities) and what I'd call the model's 'situation awareness'.

An example related to the latter: One kind of "dumb" behavior that users have been reporting is that the Gemma 4 models tend to insist that things happening after their training cutoff must be fictional. Now admittedly that is just me parroting information from someone else, I've not actually tried Gemma 4 31B with web search. But that's the thing: even though this is a relatively minor issue, it just seems like it'd be annoying to deal with – at least in contrast to Qwen3.8 27B's behavior here (I gave Qwen3.8 27B within Pi a recent LLM research paper and it basically just went "wait, so that's what has happened after my training cutoff date, a'ight that's kinda cool", which I think is a more appropriate reaction).

(Keeping possible future Gemma 5 models about as 'uncensored' as the Gemma 4 ones would also be much appreciated – I feel like that is one of Gemma 4 (31B)'s "moats" as well. For some that might mean 'unhinged RP scenarios', but I also find that the lack of censorship makes Gemma 4 31B a great model when doing for example a kind of 'rubber ducking' on one's thinking about a sensitive topic.)

I hope this doesn't come across as me being all like "make the perfect LLM for my use case lol!!!!!"; rather I just think that what I'm saying here most likely aligns with a lot of other users' perspectives as well.

Anyway that's just my two cents; thanks for EmbeddingGemma 2!

I have tried all of them, includes: Gemma4 26B, 31B and qwen3.8 27B

Qwen3.8 says very β€œtechinal”. This is quite understandable, as it focuses on code.

Gemma4 26B says very naturally, however it often skip some inportance information in tool results.

Gemma4 31B says very naturally, good output format, smart,... It can focus on importance sections to answer,.. overcomes the drawbacks of Gemma-4 26B.

However, both gemma4 26B and 31B take quite much Vram to load model and kvcache (vLLM). Hope that gemma5 can enhance this one as well.

For what it's worth, I also think Gemma 4's generality and flexibility is what makes it exceptional. This allows it to fill some important niches overlooked by more specialised models; but it also seems to make Gemma 4 more capable in some specific ways, even within those other models' niches. For example, I use Qwen to assist with programming but fall back on Gemma to advise on more abstract aspects of the programming workflow, like math, algorithms, and neural network architecture. I imagine it would also be better for generating code comments, moderating project level decisions, and such. Needless to say, there are other domains where Gemma 4 is simply the best choice. If I want to have a natural conversation or review a research paper, the models designed to emphasise programming inevitably fall short.

Sign up or log in to comment