SAM + CLIP + DIFFUSION for image to edit objects in images using plain text
Open-source persistent memory infrastructure for LLM applications.🐬
Model Context Protocol (MCP) server implementation for semantic vector search and memory management using TxtAI. This server provides a robust API for storing, retrieving, and managing text-based memories with semantic vector database search capabilities. You can use Claude and Cline AI as well.
Pytorch Implementation of CLIP-Lite | Accepted at AISTATS 2023
PegasusX: The Future of Multimodal Embeddings 🦄 🦄
Code for the paper [SIGIR'26]“Towards Mixed-Modal Retrieval for Universal Retrieval-Augmented Generation”
Summit Vitals: Multi-Camera and Multi-Signal Biosensing at High Altitudes
An Interactive Game-based Vision Planning benchmark
[ICRA26] OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation
A 5-way embedding model for text, audio, image, video, and 3D point clouds.
A lean, local AI research agent framework that runs on your machine
共 9101 条 · 第 441 / 456 页