multimodal
972 个项目 · ⭐ 707.9kOfficial implementation of JAMIE, Joint variational Autoencoders for Multimodal Imputation and Embedding
Finalist at Brainhack TIL 2024: Team 12000SGDPLUSHIE
Disentangled Graph Variational Auto-Encoder for Multimodal Recommendation with Interpretability, IEEE TMM
Collects a multimodal dataset of Wikipedia articles and their images
Learning Adaptive Fusion Bank for Multi-modal Salient Object Detection
[ACL 2023] VSTAR is a multimodal dialogue dataset with scene and topic transition information
Miscellaneous codes and writings for MLOps
This open-source project delivers a complete pipeline for converting multi-page documents (PDFs/images) into structured JSON using Vision LLMs on Amazon SageMaker. The solution leverages the SWIFT Framework to fine-tune models specifically for document understanding tasks.
CLI & async Python library for free AI chat, image & video generation.
Characterize Anything: A Wondrous Chemical Reaction between vision models and AI Characters
共 972 条 · 第 41 / 49 页