Loading lesson page...
Open-Vocabulary Vision — CLIP
Build + UsePythonNo prerequisitesTrain an image encoder and a text encoder together so that matching (image, caption) pairs land at the same point in a shared space. That is the whole trick.
Train an image encoder and a text encoder together so that matching (image, caption) pairs land at the same point in a shared space. That is the whole trick.