Pytorch code for "BAH Dataset for Ambivalence/Hesitancy Recognition in Videos for Digital Behavioural Change"
Repository to train multimodal latent reasoning tokens for Qwen 2.5 VL.
Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data
MultiModal Classifier Hierarchy (MMoCHi)
A comprehensive framework to explore whether embodied multimodal models are plausibly resilient
Self-hosted multimodal AI workspace — chat, vision QA, text-to-image, image-to-image in one conversation
Unified-modal Salient Object Detection via Adaptive Prompt Learning
[Reproduce] Code for the EMNLP2018 paper "A Visual Attention Grounding Neural Model for Multimodal Machine Translation".
An out-of-tree vLLM plugin for Mobilint NPU runtime integration.
vLLM tool parser for Qwen2.5-Coder models using <tools> tag format.
共 7255 条 · 第 355 / 363 页