Superiority of Multi-Head Attention in In-Context Linear Regression
We present a theoretical analysis of the performance of transformer with softmax
attention in in-context learning with linear regression tasks. While the
existing literature predominantly focuses on the convergence of transformers
with single-/multi-...
solvingmatters.hashnode.dev1 min read