Chasing down why it overfit being the most interesting part is the right instinct, and intent classification is a domain where it happens almost immediately. A transformer encoder has far more capacity than a few thousand short utterances justify, so it memorises phrasing rather than learning intent.
The tell worth checking is whether the errors concentrate on the rare classes. With an uneven intent distribution a model can post a respectable overall score while being useless on everything outside the top few.
At that dataset size the boring fixes usually beat architectural ones: fewer layers, aggressive dropout, and starting from a pretrained multilingual encoder instead of random init.