CLIP for Cross-Modal Retrieval: From Embedding to Search
While working on a recent project, I faced a challenge that seemed simple at first: I needed to generate embeddings for both images and text and use them together for retrieval. But the deeper I went, the more I realized the core issue — the image an...
my-ml-journey.hashnode.dev8 min read