computer-vision
507 个项目 · ⭐ 2087.8kvision language models finetuning notebooks & use cases (Medgemma - paligemma - florence .....)
[ICLRW 2024] Efficient Remote Sensing with Harmonized Transfer Learning and Modality Alignment
SPVD: Efficient and Scalable Point Cloud Generation with Sparse Point-Voxel Diffusion Models
CrossCLR: Cross-modal Contrastive Learning For Multi-modal Video Representations, ICCV 2021
Code repository for the work "Multi-Domain Incremental Learning for Semantic Segmentation", accepted at WACV 2022
Official code for "Amodal Completion via Progressive Mixed Context Diffusion" [CVPR 2024 Highlight]
PhyX: Does Your Model Have the "Wits" for Physical Reasoning?
This repository includes all computer vision, audio, document AI, and multimodal projects.
STB-VMM: Swin Transformer Based Video Motion Magnification (official repository)
[NeurIPS 2024] Official code for DiffCut: Catalyzing Zero-Shot Semantic Segmentation with Diffusion Features and Recursive Normalized Cut
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
[CVPR 2024] The official implementation of paper "synthesize, diagnose, and optimize: towards fine-grained vision-language understanding"
Human parsing model for fashion and virtual try-on applications
共 507 条 · 第 21 / 26 页