vision
81 个项目 · ⭐ 227.7kStream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across various modality combinations.
Program that lets you ask questions about your documents, audio, and video files.
Open-source spreadsheets platform for deep research and document processing
[ICLR'24] Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Multi-Modal Transformer for Video Retrieval
Curiso is an infinite canvas for your thoughts
Official code for Paper "Mantis: Multi-Image Instruction Tuning" [TMLR 2024 Best Paper]
Recrafting Video Ads with Generative AI
Easiest way of fine-tuning HuggingFace video classification models
Agentic RAG, Multi-Agent Systems, and Vision Reasoning are three pipelines to find the perfect LLM
Convert PowerPoint files into semantically rich text using vision language models
An open source chat bot architecture for voice/vision (and multimodal) assistants, local(CPU/GPU bound) and remote(I/O bound) to run.
共 81 条 · 第 2 / 5 页