multimodal
972 个项目 · ⭐ 708.2kOfficial Repository of paper VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
Code and Pretrained Models for ICLR 2023 Paper "Contrastive Audio-Visual Masked Autoencoder".
Prompts of GPT-4V & DALL-E3 to full utilize the multi-modal ability. GPT4V Prompts, DALL-E3 Prompts.
Official repo for "GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization"
Phi-3, -3.5, and -4 for Mac: Locally-run Vision and Language Models for Apple Silicon
Archived snapshot of Thinking-with-Visual-Primitives
Code for the CVPR 2024 paper highlight and demo "PIGEON: Predicting Image Geolocations".
Code/Data for the paper: "LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding"
Multi-Modal Transformer for Video Retrieval
llama.cpp (GGUF LLMs) and llava.cpp (GGUF VLMs) for ROS 2
(AAAI 2024) BLIVA: A Simple Multimodal LLM for Better Handling of Text-rich Visual Questions
From scratch implementation of a vision language model in pure PyTorch
Language Models Can See: Plugging Visual Controls in Text Generation
共 972 条 · 第 12 / 49 页