image-captioning
30 个项目 · ⭐ 28.5kLAVIS - A One-stop Library for Language-Vision Intelligence
InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)
Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
Caption-Anything is a versatile tool combining image segmentation, visual captioning, and ChatGPT, generating tailored captions with diverse controls for user preferences. https://huggingface.co/spaces/TencentARC/Caption-Anything https://huggingface.co/spaces/VIPLab/Caption-Anything
Simple Swift class to provide all the configurations you need to create custom camera view in your app
Tag manager and captioner for image datasets
Language Models Can See: Plugging Visual Controls in Text Generation
Official Code for 'RSTNet: Captioning with Adaptive Attention on Visual and Non-Visual Words' (CVPR 2021)
Improving Chest X-Ray Report Generation by Leveraging Warm-Starting
Pytorch implementation of image captioning using transformer-based model.
Combining ViT and GPT-2 for image captioning. Trained on MS-COCO. The model was implemented mostly from scratch.
This will contain the code for the 2nd edition of NLP with TensorFlow (Edition 2)
Transformer & CNN Image Captioning model in PyTorch.
An end-to-end image captioning project using a CNN encoder (ResNet-50) and LSTM decoder in PyTorch. Includes vocabulary building, preprocessing, training with BLEU evaluation, and inference. Generates natural language captions for images with saved metrics, model checkpoints, and visualization outputs.
共 30 条 · 第 1 / 2 页