[CVPR 2026 Highlight] ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
[Reproduce] Code for the ACL2019 paper "Multimodal Transformer for Unaligned Multimodal Language Sequences".
Light-MER for efficient multimodal emotion recognition with sub-1B MLLMs.
For our ACL25 Paper: Can Language Models Replace Programmers? RepoCod Says ‘Not Yet’ - by Shanchao Liang and Yiran Hu and Nan Jiang and Lin Tan
[2025] ModalFormer: Multimodal Transformer for Low-Light Image Enhancement
:speaking_head: :keyboard: Speech-to-text on key for Linux
Offline-first voice dictation, transcription, cleanup, and text-to-speech for macOS, iOS, and Android.
A fluent, scalable, and easy-to-use LLM data processing framework.
PySpark custom data source for Hugging Face Datasets
Real-time visualization of sentiment analysis on text input
共 9100 条 · 第 415 / 455 页