The speed table would be more useful with prompt processing separated from token generation. The quoted 25–40 tokens per second for a 7–8B model on an M3 Pro tells a reader about decoding, but a long pasted document can make the wait before the first token dominate a short answer.
A compact benchmark recipe could use two fixed prompt lengths and a fixed output length, reporting time to first token and subsequent decoding speed separately. Record the actual context setting and offload configuration too. That helps distinguish a model that fits comfortably from one that feels responsive for the workload the reader intends.
The speed table would be more useful with prompt processing separated from token generation. The quoted 25–40 tokens per second for a 7–8B model on an M3 Pro tells a reader about decoding, but a long pasted document can make the wait before the first token dominate a short answer.
A compact benchmark recipe could use two fixed prompt lengths and a fixed output length, reporting time to first token and subsequent decoding speed separately. Record the actual context setting and offload configuration too. That helps distinguish a model that fits comfortably from one that feels responsive for the workload the reader intends.