← 返回专题广场
multimodal
972 个项目 · ⭐ 707.9k781
This repo contains the original implementation of VAuLT, the Vision-and-Augmented-Language Transformer. We provide instructions to download some multimodal social-media datasets, and scripts to experiment with. VAuLT is a stack of Transformers, a LM like BERT that preprocesses the text input of ViLT
Python
⭐ 18
⑂ 1
MIT
· 2025-09-24推送
2025-09-24
最近推送
782
2025-06-24
最近推送
783
Open Translator: Speech To Speech and Speech to text Translator with voice cloning and other cool features
Python
⭐ 18
⑂ 4
NOASSERTION
· 2026-03-26推送
2026-03-26
最近推送
784
2026-01-13
最近推送
785
2021-12-09
最近推送
786
18 天前
最近推送
787
AI eyes that roll through video footage — watch, understand, act
Python
⭐ 17
⑂ 1
MIT
· 2026-05-15推送
2026-05-15
最近推送
788
2026-07-06
最近推送
789
2026-05-16
最近推送
790
2026-03-06
最近推送
791
2025-09-19
最近推送
792
2024-10-18
最近推送
793
2023-12-11
最近推送
795
[WACV 2025] I Dream My Painting: Connecting MLLMs and Diffusion Models via Prompt Generation for Text-Guided Multi-Mask Inpainting
Jupyter Notebook
⭐ 17
⑂ 0
MIT
· 2025-12-30推送
2025-12-30
最近推送
796
2024-11-01
最近推送
797
2026-03-27
最近推送
798
15 天前
最近推送
800
2025-01-23
最近推送
共 972 条 · 第 40 / 49 页