Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data
MultiModal Classifier Hierarchy (MMoCHi)
A comprehensive framework to explore whether embodied multimodal models are plausibly resilient
Self-hosted multimodal AI workspace — chat, vision QA, text-to-image, image-to-image in one conversation
Unified-modal Salient Object Detection via Adaptive Prompt Learning
[Reproduce] Code for the EMNLP2018 paper "A Visual Attention Grounding Neural Model for Multimodal Machine Translation".
An out-of-tree vLLM plugin for Mobilint NPU runtime integration.
vLLM tool parser for Qwen2.5-Coder models using <tools> tag format.
FlashHead: Efficient Drop-In Replacement for the Classification Head in Language Model Inference
Benchmark OpenAI-compatible AI endpoints and AI Accelerators in a reproducible structured way
共 9100 条 · 第 447 / 455 页