Inference

#2
by islameissa - opened

Hi,
I wonder if it works with vLLM, SGLang, FreeToken?
MTP? is it included? if so, is it quantized too?
Vision?r

Under my test, it works on vllm-backport v0.13.1, but mtp layer has namespace issue than can't be loaded by vllm properly.

By disabling mtp, this model can reach about prefill ~1600-2200tps, decode ~37-47tps with 4x170hx (TP=4) at pcie2.0 x4, without p2p link.


Add on:
After de-quanterizing mtp layers into bf16 and setup layer exclude filters correctly, this model can now get intialized with mtp support, currently decode speed reaches avg ~66tps and sometimes burst at ~80-90tps

AlexYzhov
How is your experience with using multiple NVIDIA CMP 170HX 8GB HBM2 GPU Gen2 x4 Unlocked 64GB? is the PCIe2 x4 causing a bottle neck or not much?

AlexYzhov
How is your experience with using multiple NVIDIA CMP 170HX 8GB HBM2 GPU Gen2 x4 Unlocked 64GB? is the PCIe2 x4 causing a bottle neck or not much?

Actually I've benchmarked 4 cards communication performance under pcie2.0 x4 (with no p2p enabled):

2026-09-30T06:28:02,P0.5-shm,allreduce|8KB|75.6|0.11|0.16
2026-09-30T06:28:02,P0.5-shm,allreduce|512KB|778.7|0.67|1.01
2026-09-30T06:28:02,P0.5-shm,allreduce|8192KB|10159.3|0.83|1.24

The 1st column is payload size, 2nd is time(us), 3rd is algobw(GB/s), 4th is busbw.

And with the help of this model, I let it done the math to calculate the average throughput per token with previous performance data (TP=4, prefill ~1600-2200tps, decode ~37-47tps), based on the GLM5 paper. It is about 0.55-1.1MB/s per token on each card.

So technically pcie2.0 x4 caused bottle neck, but very little especially in decode stage.
If current prefill performance has already statisfied your needs, it won't benefit too much on decode performance.

I let GLM to explore more, it showed me benefits on each step:
· step 1. enable p2p: prefill +25-65%, decode +15-20%
· step 2. pcie2.0 x4 -> pcie2.0 x16: prefill +30%(which may turn prefill stage from communicational bound into computational bound), decode +~0%
· step 3. unlocking pcie gen3: prefill +0% (already bounded by SMs), decode +3-5% (sightly increase from faster TLP level handshake, not bandwidth)

AlexYzhov
How is your experience with using multiple NVIDIA CMP 170HX 8GB HBM2 GPU Gen2 x4 Unlocked 64GB? is the PCIe2 x4 causing a bottle neck or not much?

Actually I've benchmarked 4 cards communication performance under pcie2.0 x4 (with no p2p enabled):

2026-09-30T06:28:02,P0.5-shm,allreduce|8KB|75.6|0.11|0.16
2026-09-30T06:28:02,P0.5-shm,allreduce|512KB|778.7|0.67|1.01
2026-09-30T06:28:02,P0.5-shm,allreduce|8192KB|10159.3|0.83|1.24

The 1st column is payload size, 2nd is time(us), 3rd is algobw(GB/s), 4th is busbw.

And with the help of this model, I let it done the math to calculate the average throughput per token with previous performance data (TP=4, prefill ~1600-2200tps, decode ~37-47tps), based on the (GLM5 paper](https://arxiv.org/abs/2602.15763). It is about 0.55-1.1MB/s per token on each card.

So technically pcie2.0 x4 caused bottle neck, but very little especially in decode stage.
If current prefill performance has already statisfied your needs, it won't benefit too much on decode performance.

I let GLM to explore more, it showed me benefits on each step:
· step 1. enable p2p: prefill +25-65%, decode +15-20%
· step 2. pcie2.0 x4 -> pcie2.0 x16: prefill +30%(which may turn prefill stage from communicational bound into computational bound), decode +~0%
· step 3. unlocking pcie gen3: prefill +0% (already bounded by SMs), decode +3-5% (sightly increase from faster TLP level handshake, not bandwidth)

Thanks you very much. These cards are sold for around 500USD unlocked to 64GB. That is the cheapest 64GB of any type of RAM I can ever get and with 4 of them, that’s a very good amount of VRAM to work with. There is a version that’s PCIe Gen2x16. If PCIe Gen 3 can be unlocked on it too, that would be great as at that stage, P2P communication should not be affected much anymore. Cheaper than Apple with maybe similar performance?!
That’s why I was asking. Seems to me worth a try.

islameissa

I'd like to say 500 usd for 64GB 170hx would be very much reasonable.

If the hardware works fine, the only problems for 170hx is it's #1 residual value and #2 scalable issue.

#1 Residual value is very easy to understand, unlike G-Force cards still have game players to maintain its price, 170hx would not worth a penny if better solution appears.

#2 Scalable issue is an engineering problem. The size of mainstream llms are not that friendly for 4x/5x/8x GPU arrays. having 256GB/320GB/512GB vram does not change things too much. The trending size of next generation "Flash" level LLM would be larger, also the context size of COT may increase (which makes decode performance very important). I'm a little pessimistic for the scaling potential of 170hx based solution. From my perspective, models larger than DeepSeek 4.1 Flash would be out of 170hx(s) ability.

Sign up or log in to comment