The 0.7 semantic / 0.3 keyword blend is a useful experiment, but the scales determine what those weights actually mean. Your keyword score counts matching query tokens, while cosine similarity lives on a different scale; a longer query can therefore increase the keyword contribution without any change in relevance.
I would print both raw components beside the combined score and try the same question with a few unnecessary words added. Normalizing the keyword component or comparing rank fusion would help separate a deliberate weighting choice from an accidental consequence of query length.
praveenkulharee
Thanks for the valuable feedback! The 0.7/0.3 blend is part of my learning implementation, so I’ll test query variations and compare raw, normalized, and rank-fusion approaches next.