multimodal
973 个项目 · ⭐ 708.4kThe Official Repo for "Quick Start Guide to Large Language Models"
Creates an index of images, queries a local LLM and adds tags to the image metadata
Virtual Sparse Convolution for Multimodal 3D Object Detection
Code for the paper "LLark: A Multimodal Instruction-Following Language Model for Music" by Josh Gardner, Simon Durand, Daniel Stoller, and Rachel Bittner.
PyTorch Implementation for Paper "Emotionally Enhanced Talking Face Generation" (ICCVW'23 and ACM-MMW'23)
🏭 Mega Scale Multimodal DataPipeline for SOTA Foundation Models
Joint speech-language model - respond directly to audio!
Research Trends in LLM-guided Multimodal Learning.
A customizable pipeline for multimodal data extraction from MIMIC-IV!
共 973 条 · 第 10 / 49 页