multimodal
972 个项目 · ⭐ 708.0kImplementation of GATO style Generalist Multimodal model capable of image, text, RL and Robotics tasks
[MM 2025] A Multimodal Finance Benchmark for Expert-level Understanding and Reasoning
Visual Instruction Tuning for Qwen2 Base Model
An end-to-end image captioning project using a CNN encoder (ResNet-50) and LSTM decoder in PyTorch. Includes vocabulary building, preprocessing, training with BLEU evaluation, and inference. Generates natural language captions for images with saved metrics, model checkpoints, and visualization outputs.
Repository for the paper "MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation)" (NAACL 2022).
[NeurIPS-2024] The offical Implementation of "Instruction-Guided Visual Masking"
[ACL 2024] An Easy-to-use Hallucination Detection Framework for LLMs.
[MobiCom 2022] InFi is a library for building input filters for resource-efficient inference.
共 972 条 · 第 28 / 49 页