vit
22 个项目 · ⭐ 39.6kpix2tex: Using a ViT to convert images of equations into LaTeX code.
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
Towhee is a framework that is dedicated to making neural data processing pipelines simple and fast.
Turn any computer or edge device into a command center for your computer vision projects.
[CVPR 2025 Highlight] Official code and models for Encoder-only Mask Transformer (EoMT).
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
HugsVision is a easy to use huggingface wrapper for state-of-the-art computer vision
An unofficial implementation of ViTPose [Y. Xu et al., 2022]
Combining ViT and GPT-2 for image captioning. Trained on MS-COCO. The model was implemented mostly from scratch.
CLIP-based aesthetics predictor inspired by the interface of 🤗 huggingface transformers.
This is a cross-modal benchmark for industrial anomaly detection.
A Persian Image Captioning model based on Vision Encoder Decoder Models of the transformers🤗.
共 22 条 · 第 1 / 2 页