The dense-51M comparison will be especially informative if it uses the same tokenizer, token stream, context length, and training budget. Active parameter count is a useful architectural description, but it does not establish equivalent wall-clock cost on a T4: dispatching many small experts can introduce overhead that a dense matrix avoids.
Alongside validation loss, I would report tokens per second, peak memory, and total training time for both models. A second comparison at equal wall-clock budget would answer a different practical question: which model produces the better checkpoint before Colab preemption? Keeping those two comparisons separate would make the eventual result easier to interpret.
The dense-51M comparison will be especially informative if it uses the same tokenizer, token stream, context length, and training budget. Active parameter count is a useful architectural description, but it does not establish equivalent wall-clock cost on a T4: dispatching many small experts can introduce overhead that a dense matrix avoids.
Alongside validation loss, I would report tokens per second, peak memory, and total training time for both models. A second comparison at equal wall-clock budget would answer a different practical question: which model produces the better checkpoint before Colab preemption? Keeping those two comparisons separate would make the eventual result easier to interpret.