dataset
155 个项目 · ⭐ 720.3kDeveloped a sophisticated machine learning model capable of generating diverse interview questions aligned with specific topics, ensuring depth of conversation. Integrated advanced Natural Language Processing (NLP) algorithms to analyse spoken responses, identifying grammatical errors & offering accurate corrections after the interview.
Generative AI-based CyberSecurity-focused Prompt Dataset for Benchmarking Large Language Models
Repository for the paper "MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation)" (NAACL 2022).
A tiny language model that talks like a house cat named Miso.
The first large-scale summarization corpus for the Indonesian language. AACL 2020.
A instruction data generation system for multimodal language models.
Wav2vec resources and models for Brazilian Portuguese
A merged version of multiple open-source German speech datasets.
Automatic Speech Recognition Dataset for Oromo Language
:page_facing_up: Verified Ethereum Smart Contract dataset
:hugs: A multilabel lymph node segmentation dataset from contrast CT
Web Scraping, Document Deduplication & GPT-2 Fine-tuning with a newly created scam dataset.
Multimodal dataset for ad text generation in Japanese [Mita+, ACL2024]
共 155 条 · 第 6 / 8 页