Does a Wider Embedding Space Help More on Diverse Data?
TL;DR
Increasing the embedding dimension from 32 to 64 improved validation loss in both the low- and high-diversity training conditions.
Most of the improvement appeared between 32 and 48 dimensions
curious-pm.hashnode.dev7 min read