multimodal
972 个项目 · ⭐ 707.9kAwesome-HCI (Ubiquitous, LLM, MLLM, Agent, RAG, Embodied-AI, RLHF)
Code for Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense? [COLM 2024]
Self-hosted AI knowledge base with hybrid semantic search (pgvector + FTS + RRF), MCP server, multi-provider LLM inference (Ollama, OpenAI, OpenRouter, llama.cpp), multimodal ingestion (vision, audio transcription, speaker diarization), and knowledge graph. Rust + PostgreSQL.
VL-JEPA inspired pipeline — compress images/text locally via Ollama, send compact payloads to any LLM API. Cut token costs by ~80%.
🚁 Can Vision-Language Models Think from the Sky? UAVReason for Aerial Reasoning and Generation
Implementation for the different ML tasks on Kaggle platform with GPUs.
[ICLR 2026] - Spectral Concept Selection and Cross-modal Representation Learning for Generalized Category Discovery
Multimodal summarization of user-generated videos from wearable cameras
Telephony Server is a powerful bridge that connects telephony providers (Twilio, Vonage, Plivo, etc.) with real-time communication platforms (LiveKit, Jay.so, Pipecat, etc.). It enables seamless call routing, robust metrics collection, and observability features for enhanced telephony operations.
Android 16 fork. AI as a platform primitive. Twelve capabilities, one shared runtime, every app. OEM-pluggable. Apache 2.0.
共 972 条 · 第 36 / 49 页