Thanks, John!
I think this specific corpus includes quite a few cultural variances since the questions comes from real people from different backgrounds, so it's not all in a specific scientific or academic dialect, it's all over the place. More data is always more accurate in this scenario though, I'll think about additional data sources for the future.
Thanks!
JohnLogan
This is a fascinating exploration of embedding models! I appreciate how you highlighted the nuances between different models when it comes to Hebrew. It might be interesting to consider additional contextual factors, like cultural variances in language use, which could further impact model performance. Thanks for sparking such an insightful discussion! monkey mart