Official Repository of paper VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
Dashboard to collect, analyze, and respond to reported phishing emails.
Segment Anything Model for large-scale, vectorized road network extraction from aerial imagery. CVPRW 2024
Pytorch implementation of "Genie: Generative Interactive Environments", Bruce et al. (2024).
FASHN VTON v1.5: Efficient Maskless Virtual Try-On in Pixel Space
Llama-github is an open-source Python library that empowers LLM Chatbots, AI Agents, and Auto-dev Solutions to conduct Agentic RAG from actively selected GitHub public projects. It Augments through LLMs and Generates context for any coding question, in order to streamline the development of sophisticated AI-driven applications.
Working online speech recognition based on RNN Transducer. ( Trained model release available in release )
ICML 2026 · Plug-and-play long-term memory for LLM agents
Better Aligning Text-to-Image Models with Human Preference. ICCV 2023
Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis (ECCV 2024 Oral) - Official Implementation
命令执行不回显但DNS协议出网的命令回显场景解决方案(修改为使用ceye接收请求,添加自定义DNS服务器)
Code and Pretrained Models for ICLR 2023 Paper "Contrastive Audio-Visual Masked Autoencoder".
共 9118 条 · 第 252 / 456 页