Even though the accuracy seems comparable, there is significant token inflation on a per task basis
#8
by jakubjaniak - opened
Included a writeup here with more details on the problem
https://collectgarbage.substack.com/p/quantized-endpoints-charge-less-per