The real insight here is that per-token price is half the math. GLM's verbosity advantage eats its price advantage on prose tasks; Qwen's efficiency flips on code. It's the product that hits your invoice, not either factor in isolation. Nobody benchmarks that part.