Incorporate Image, Text and Tabular Data with HuggingFace Transformers
[CVPR 2026] Flow Matching for Multimodal Distributions
OMERO.web plugin for the Vitessce multimodal data viewer.
Multimodal document QA: vision + retrieval over PDFs (LLaVA + LlamaIndex)
A high-throughput LLM serving engine with non-uniform KV cache compression, built on vLLM
GPU Service Manager for LLM workloads on Linux/NVIDIA systems.
The Continuous Verification and Optimization layer for self-hosted LLMs
SAM — Smart Agentic Model: CLI coding agent for open-source LLMs. pip install sam-agent
CI scripts designed to build a Pascal-compatible version of vLLM.
Unsupervised specificity-guided optimization of Image Captioning models to encourage meaningful diversity in the generated captions. Code for the paper Generating Diverse and Meaningful Captions: Unsupervised Specificity Optimization for Image Captioning (Lindh et al., 2018).
T5 Fine-tuning on SQuAD Dataset for Question Generation
共 7255 条 · 第 352 / 363 页