computer-vision
507 个项目 · ⭐ 2087.8kBuilding an Image-to-Video (I2V) Model from Scratch
Empowering Data Driven insights through hands-on projects, SQL challenges and practical tools.
Code for Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense? [COLM 2024]
VL-JEPA inspired pipeline — compress images/text locally via Ollama, send compact payloads to any LLM API. Cut token costs by ~80%.
Implementation for the different ML tasks on Kaggle platform with GPUs.
OCR system for handwritten medical prescriptions using Donut transformer and zero-shot classification
Application for playing Quickdraw powered by a Vision Transformer
Grounding Language Models for Compositional and Spatial Reasoning
Identifying reasons for human actions in lifestyle vlogs.
[WACV 2025] I Dream My Painting: Connecting MLLMs and Diffusion Models via Prompt Generation for Text-Guided Multi-Mask Inpainting
RiverSnap - estimation of river hydraulic parameters using machine learning/AI models
Forest Fire Detection By Convolutional Neural Network
共 507 条 · 第 24 / 26 页