The filename-set check is the right shape but it'll still miss the exact case you flagged at the end, a file that gets edited but keeps its name. Hashing file content instead of comparing names closes that gap without much more code, and it survives renames too since you'd key by hash rather than path. One thing worth watching as the corpus grows: "compare indexed vs on-disk" means loading the whole FAISS index and unpickling the metadata before you even know a rebuild is needed. Worth splitting out a small manifest file (filename to hash) so the staleness check itself stays cheap and doesn't require deserializing the full store first.
The filename-set check is the right shape but it'll still miss the exact case you flagged at the end, a file that gets edited but keeps its name. Hashing file content instead of comparing names closes that gap without much more code, and it survives renames too since you'd key by hash rather than path. One thing worth watching as the corpus grows: "compare indexed vs on-disk" means loading the whole FAISS index and unpickling the metadata before you even know a rebuild is needed. Worth splitting out a small manifest file (filename to hash) so the staleness check itself stays cheap and doesn't require deserializing the full store first.