Self-hosted multimodal AI workspace — chat, vision QA, text-to-image, image-to-image in one conversation
Multimodal and multilingual topic model with pretrained embeddings
[NeurIPS2023] LoRA: A Logical Reasoning Augmented Dataset for Visual Question Answering
Unified-modal Salient Object Detection via Adaptive Prompt Learning
[Reproduce] Code for the EMNLP2018 paper "A Visual Attention Grounding Neural Model for Multimodal Machine Translation".
Exploring Visual Interpretability for Contrastive Language-Image Pretraining
Batch LLM Inference with Ray Data LLM: From Simple to Advanced
An out-of-tree vLLM plugin for Mobilint NPU runtime integration.
LLM inference engine built from scratch in C++. No PyTorch, no frameworks.
vLLM tool parser for Qwen2.5-Coder models using <tools> tag format.
FlashHead: Efficient Drop-In Replacement for the Classification Head in Language Model Inference
共 40673 条 · 第 2013 / 2034 页