Engage in a semantic segmentation challenge for land cover description using multimodal remote sensing earth observation data, delving into real-world scenarios with a dataset comprising 70,000+ aerial imagery patches and 50,000 Sentinel-2 satellite acquisitions.
Pytorch implementation of Multimodal Neural Machine Translation(MNMT).
building various transformer model architectures and its modules from scratch.
Incorporate Image, Text and Tabular Data with HuggingFace Transformers
[CVPR 2026] Flow Matching for Multimodal Distributions
OMERO.web plugin for the Vitessce multimodal data viewer.
Multimodal document QA: vision + retrieval over PDFs (LLaVA + LlamaIndex)
LLM, Fine Tuning, Llama 2, Gemma, Mixtral, vLLM, LangChain, RAG, ChromaDB, FAISS
A high-throughput LLM serving engine with non-uniform KV cache compression, built on vLLM
共 40674 条 · 第 2005 / 2034 页