Search before speaking is the right default for a low-resource language, because the model's prior on local prices and facts is thin and confident errors are worse than slow answers. The Ge'ez token cost point is underrated too: if each answer costs 3x the tokens of English, retrieval that keeps prompts short pays for itself. Did you measure how often the model still ignored the retrieved snippet and answered from memory?
iin1005h1628