The SpeechBrain project aims to build a novel speech toolkit fully based on PyTorch. With SpeechBrain users can easily create speech processing systems, ranging from speech recognition (both HMM/DNN and end-to-end), speaker recognition, speech enhancement, speech separation, multi-microphone speech processing, and many others.
This repository provides the code for a multimodal realistic simulation framework using CARLA and MATLAB.
This repository implements a learning-based beamforming approach leveraging multimodal feature fusion. It includes data preprocessing, augmentation, and a transformer-based network for efficient beam prediction.