Narrowing to one product and one pipeline is the move I wish more people made, a general agent has too large a surface to ever really trust. "No claim without proof" is also what makes it evaluable: a narrow agent has a small enough action space that you can actually score whether each answer was grounded. Did the narrow scope cut your failure rate mostly by removing wrong tool calls, or by making retrieval land on the right context?