computer-vision
507 个项目 · ⭐ 2087.8kRecurrence Meets Transformers for Universal Multimodal Retrieval
🏠 OpenStreetMap, AI import tool for buildings in Poland
Summit Vitals: Multi-Camera and Multi-Signal Biosensing at High Altitudes
Engage in a semantic segmentation challenge for land cover description using multimodal remote sensing earth observation data, delving into real-world scenarios with a dataset comprising 70,000+ aerial imagery patches and 50,000 Sentinel-2 satellite acquisitions.
building various transformer model architectures and its modules from scratch.
Unsupervised specificity-guided optimization of Image Captioning models to encourage meaningful diversity in the generated captions. Code for the paper Generating Diverse and Meaningful Captions: Unsupervised Specificity Optimization for Image Captioning (Lindh et al., 2018).
ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval
[ICCV 2025] Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP
共 507 条 · 第 25 / 26 页