YesBut - Multimodal Satire Comprehension Dataset
小红书图文笔记 Agent Skill,支持看图理解、主动提问、智能选题、实时热梗搜索、标题正文、标签评论钩子和风险检查。
Code for "MIFNet: Learning Modality-Invariant Features for Generalizable Multimodal Image Matching", TIP2025
Implementation of "With a Little Help from my Temporal Context: Multimodal Egocentric Action Recognition, BMVC, 2021" in PyTorch
ABC: Achieving Better Control of Multimodal Embeddings using VLMs [TMLR2025]
React component and hook for speech recognition, with voice commands and pluggable STT engines
MCP server for transcript processing — formatting, contextual repair & smart summarization with deep-thinking LLMs
An Android app using whisper.cpp for voice-to-text transcription
A Flutter-based space exploration app with NASA APIs, ISS tracker, space missions, and AI-powered chatbot.
An implementation for "Conformer: Convolution-augmented Transformer for Speech Recognition" Paper
共 40708 条 · 第 1966 / 2036 页