vision-language-model
97 个项目 · ⭐ 117.0kAndroid native AI inference library, bringing text, image, video, STT, TTS inference
Mark web pages for use with vision-language models
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
[CVPR 2024] The official implementation of paper "synthesize, diagnose, and optimize: towards fine-grained vision-language understanding"
Code repository for "Post-pre-training for Modality Alignment in Vision-Language Foundation Models" (CVPR2025)
[ICLR 2025] Official code repository for "TULIP: Token-length Upgraded CLIP"
a simple lightweight large language model pipeline framework.
An open-source implementaion for fine-tuning Pixtral by MistralAI.
[ICLR 2026] - Spectral Concept Selection and Cross-modal Representation Learning for Generalized Category Discovery
共 97 条 · 第 4 / 5 页