vision-language-model
97 个项目 · ⭐ 117.0k[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
👀 Train a 65M-parameter VLM from scratch in just 2h!
MineContext is your proactive context-aware AI partner(Context-Engineering+ChatGPT Pulse)
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
Align Anything: Training All-modality Model with Feedback
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
Collection of AWESOME vision-language models for vision tasks
The Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs, vision, Whisper speech-to-text, Stable Diffusion, tool calling, and local-network servers. Runs on your CPU, GPU, or NPU. No account, no API key, zero data leaves your device.
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
The Cradle framework is a first attempt at General Computer Control (GCC). Cradle supports agents to ace any computer task by enabling strong reasoning abilities, self-improvment, and skill curation, in a standardized general environment with minimal requirements.
An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.
NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.
Get clean data from tricky documents, powered by vision-language models ⚡
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
共 97 条 · 第 1 / 5 页