Code for "MIFNet: Learning Modality-Invariant Features for Generalizable Multimodal Image Matching", TIP2025
Implementation of "With a Little Help from my Temporal Context: Multimodal Egocentric Action Recognition, BMVC, 2021" in PyTorch
ABC: Achieving Better Control of Multimodal Embeddings using VLMs [TMLR2025]
An implementation for "Conformer: Convolution-augmented Transformer for Speech Recognition" Paper
Speech-to-Speech conversational AI using Azure OpenAI Service and Azure Speech Services
Manalyzer: End-to-end Automated Meta-analysis with Multi-agent System
Proxima lets existing GPUs serve 4x more concurrent requests
A lightweight post-training framework for LLMs and VLMs. 51 algorithms, 38 verified models. Scales with DeepSpeed, vLLM, and Ray.
This repository provides an easy way to train your models on the datasets of DCASE task 1.
共 7255 条 · 第 336 / 363 页