llava
54 个项目 · ⭐ 67.4k[ICLR'24] Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Official Repository of paper VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
AI Device Template Featuring Whisper, TTS, Groq, Llama3, OpenAI and more
Code/Data for the paper: "LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding"
llama.cpp (GGUF LLMs) and llava.cpp (GGUF VLMs) for ROS 2
From scratch implementation of a vision language model in pure PyTorch
A Ruby gem for interacting with Ollama's API that allows you to run open source AI LLMs (Large Language Models) locally.
Embed arbitrary modalities (images, audio, documents, etc) into large language models.
中文医学多模态大模型 Large Chinese Language-and-Vision Assistant for BioMedicine
Self-host a ChatGPT-style web interface for Ollama 🦙
[NAACL 2024] MMC: Advancing Multimodal Chart Understanding with LLM Instruction Tuning
RLLaVA is a user-friendly framework for multi-modal RL research and optimized for resource-constrained teams.
This repository contains a web application designed to execute relatively compact, locally-operated Large Language Models (LLMs).
共 54 条 · 第 2 / 3 页