The speed table would be more useful with prompt processing separated from token generation. The quoted 25–40 tokens per second for a 7–8B model on an M3 Pro tells a reader about decoding, but a long pasted document can make the wait before the first token dominate a short answer.
A compact benchmark recipe could use two fixed prompt lengths and a fixed output length, reporting time to first token and subsequent decoding speed separately. Record the actual context setting and offload configuration too. That helps distinguish a model that fits comfortably from one that feels responsive for the workload the reader intends.