[ACL-IJCNLP 2021] Self-Supervised Multimodal Opinion Summarization
Awesome-HCI (Ubiquitous, LLM, MLLM, Agent, RAG, Embodied-AI, RLHF)
Automate the creation of soccer match highlights with the power of Generative AI and AWS. This solution leverages AWS Bedrock (Anthropic’s Claude 3 Sonnet model), AWS MediaConvert, Lambda, Step Functions and other AWS services to identify and compile exciting game moments without manual editing.
Search knowledge by what documents mean and how they look — not one or the other.
A Benchmark Dataset for Multimodal Scientific Fact Checking
EVA: Efficient Reinforcement Learning for End-to-End Video Agent
A stand-alone application with GUI for OpenAI's Whisper
Reproducible voice-AI benchmarking — TTS / STT latency and accuracy.
This is an OpenAI Whisper automatic speech recognition microservice
Probing the limitations of multimodal language models for chemistry and materials research
Kubernetes scanner that discovers LLMs running on vLLM and extracts their deployment and runtime facts.
共 7255 条 · 第 328 / 363 页