Qwen/Qwen3.8-27B
Image-Text-to-Text • 28B • Updated • 6.84M • • 16.5k
Time will come! human with 86 billion brain cells already took us this far.
You're 100% correct on the ecosystem point of view. But DGX Station 748GB coherent memory has a major draw back, 496 GB out of 748GB memory are LPDDR5X memory with memory bandwidth of 396GB/s.
If your model weight exceed the real VRAM (252GB in GB300), inference speed will be bounded by 396GB/s, wasting 7.1TB bandwidth GB300 has.
Best practice is using a model that fits in 252GB(such as GLM5.3-Flash NVFP4, DeepseekV4-Flash...) and utilize 496GB LPDDR5X as KV cache.