@Rahul Sai Indeevar V - adding this here since the reply box on your thread would not submit for me.
If you do write that update, the experiment that makes the prefill and decode split click is measuring both phases separately at two or three batch sizes. Prefill throughput scales with batch until you saturate compute; decode throughput barely moves, because you re-read the whole KV cache either way. Seeing those two curves next to each other explains why continuous batching exists, and why a single tokens-per-second figure hides most of what is actually happening.