vlm
117 个项目 · ⭐ 386.2k🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
[CVPR 2026🔥] 🧑🎨 OmniLottie, an open-sourced multi-modal instructed vector animation generator that produces Lottie JSONs.
Dingo: A Comprehensive AI Data, Model and Application Quality Evaluation Tool
[ECCV 2026] SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation
🤗 Optimum Intel: Accelerate inference with Intel optimization tools
Flame is an open-source multimodal AI system designed to translate UI design mockups into high-quality React code. It leverages vision-language modeling, automated data synthesis, and structured training workflows to bridge the gap between design and front-end development.
全网最全的、持续更新的、最火爆的 100+ 人格 蒸馏skills 合集| 多agent系统 |名人/导演/天涯大神/古籍/二次元/职场/情感全品类
A hub for various industry-specific schemas to be used with VLMs.
🏭 Mega Scale Multimodal DataPipeline for SOTA Foundation Models
Seamlessly integrate state-of-the-art transformer models into robotics stacks
Phi-3, -3.5, and -4 for Mac: Locally-run Vision and Language Models for Apple Silicon
llama.cpp (GGUF LLMs) and llava.cpp (GGUF VLMs) for ROS 2
[CVPR 2026] LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
Official code for Paper "Mantis: Multi-Image Instruction Tuning" [TMLR 2024 Best Paper]
Boot your PC straight into an LLM. Rust, UEFI-resident, no operating system underneath.
共 117 条 · 第 3 / 6 页