multimodal
972 个项目 · ⭐ 708.0kCode and models for the paper "2D3MF: Deepfake Detection using Multi Modal Middle Fusion"
Your intelligent ally for effortless information retrieval and seamless browsing across documents and the web.
AI NourishBot is an AI-powered nutrition assistant that leverages advanced vision models and natural language processing to detect ingredients from food images, filter ingredients based on dietary restrictions, estimate calories, provide detailed nutrient analysis, and generate recipe suggestions.
A bug-free and improved implementation of LLaVA-UHD, based on the code from the official repo
Multimodal Chatbot with Amazon Bedrock Knowledge Bases Integration
[ACL 2021] Learning Relation Alignment for Calibrated Cross-modal Retrieval
Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator
ProteinGPT: Multimodal LLM for Protein Property Prediction and Structure Understanding
ICLR 2026 | ChemEval: 4-level, 13-dimension, 62-task text/multimodal chemistry benchmark for evaluating LLMs and MLLMs.
Code for LifelongMemory: Leveraging LLMs for Answering Queries in Long-form Egocentric Videos
VOXRAD is a voice transcription application for radiologists leveraging locally deployed ASR and LLM models.
[MLHC 2025] ECG-Byte: A Tokenizer for End-to-End Generative Electrocardiogram Language Modeling
Panacea is a framework for building collaborative, intelligent multi agent AI systems. The framework provides a robust infrastructure for creating and managing multiple AI agents, and enables developers and organizations to build, deploy, and optimize AI agents that work well in dynamic, complex environments.
共 972 条 · 第 31 / 49 页