vlm
117 个项目 · ⭐ 386.2k[NeurIPS 2025] Deep Memory Backtracking for Long Video Understanding
vision language models finetuning notebooks & use cases (Medgemma - paligemma - florence .....)
RLLaVA is a user-friendly framework for multi-modal RL research and optimized for resource-constrained teams.
[CVPR 2025] Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
Batch Deployment for Document Parsing with AWS Batch & Qwen-2.5-VL
Search knowledge by what documents mean and how they look — not one or the other.
EVA: Efficient Reinforcement Learning for End-to-End Video Agent
Probing the limitations of multimodal language models for chemistry and materials research
Awesome-HCI (Ubiquitous, LLM, MLLM, Agent, RAG, Embodied-AI, RLHF)
[ICLR 2026] - Spectral Concept Selection and Cross-modal Representation Learning for Generalized Category Discovery
Android 16 fork. AI as a platform primitive. Twelve capabilities, one shared runtime, every app. OEM-pluggable. Apache 2.0.
共 117 条 · 第 5 / 6 页