clip
78 个项目 · ⭐ 58.9kState-of-the-art CLIP/SigLIP embedding models finetuned for the fashion domain. +57% increase in evaluation metrics vs FashionCLIP 2.0.
World's fastest and most compact embedded vector database: exact by default, multimodal, local-first, and GPU-accelerated
The most impactful papers related to contrastive pretraining for multimodal models!
[SOICT 2024] LLM-Powered Video Search: A Comprehensive Multimedia Retrieval System
[CVPR 2024] The official implementation of paper "synthesize, diagnose, and optimize: towards fine-grained vision-language understanding"
[IEEE TMI 2024] MultiEYE: Dataset and Benchmark for OCT-Enhanced Retinal Disease Recognition from Fundus Images
Vision Transformers Needs Registers. And Gated MLPs. And +20M params. Tiny modality gap ensues!
CLIP-based aesthetics predictor inspired by the interface of 🤗 huggingface transformers.
NeurIPS 2024 Track on Datasets and Benchmarks (Spotlight)
[CVPR 2025] Enhanced OoD Detection through Cross-Modal Alignment of Multi-modal Representations
共 78 条 · 第 3 / 4 页