vision-transformer
63 个项目 · ⭐ 118.2kOpenMMLab Detection Toolbox and Benchmark
pix2tex: Using a ViT to convert images of equations into LaTeX code.
This repository contains demos I made with the Transformers library by HuggingFace.
[NeurIPS 2024 Best Paper Award][GPT beats diffusion🔥] [scaling laws in visual generation📈] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *ultra-simple, user-friendly yet state-of-the-art* codebase for autoregressive image generation!
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
OpenMMLab Pre-training Toolbox and Benchmark
Scenic: A Jax Library for Computer Vision Research and Beyond
Towhee is a framework that is dedicated to making neural data processing pipelines simple and fast.
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
[ECCV2024] Video Foundation Models & Data for Multimodal Understanding
[CVPR 2025] Official PyTorch Implementation of MambaVision: A Hybrid Mamba-Transformer Vision Backbone
An all-in-one toolkit for computer vision
Get clean data from tricky documents, powered by vision-language models ⚡
[CVPR 2025 Highlight] Official code and models for Encoder-only Mask Transformer (EoMT).
共 63 条 · 第 1 / 4 页