multilingual
49 个项目 · ⭐ 115.4kEthicore Engine™ is an AI safety, ethics, and compliance platform. This repo consists of the open-source components of Ethicore Engine™ - Guardian SDK; designed to protect your AI applications from prompt injection, jailbreaks, role hijacking, system-prompt extraction, and 100+ additional threat categories through a multi-layer analysis pipeline
Easy fine-tuning for Qwen3-TTS: Fast voice cloning and high-quality multilingual speech synthesis.
OpenAi-Sora (SoraFlows) is an open-source, cross-platform web application for AI-powered video creation and editing using the latest OpenAI Sora model. Effortlessly generate, edit, and share videos from text or images, with integrated support for voice-to-text, text-to-speech, voice cloning, and multi-language interfaces .
Fine-tuning toolkit for Chatterbox TTS & Chatterbox TURBO models. Supports 23 languages with smart vocabulary extension. Features offline preprocessing, automatic VAD trimming, and voice cloning capabilities. Train custom TTS models with your own dataset in LJSpeech and file-based format.
A configurable engine for analysing multi-lingual and multi-modal content.
The largest multilingual image-text classification dataset. It contains fashion products.
Taiwan Tongues ASR CE 是一個開源語音辨識(Automatic Speech Recognition, ASR)模型專案,專為台灣多元語言環境設計。 本模型支援 國語、台語、客語與英語,提供本地多語混合語音辨識,讓開發者與資訊服務業者可運用此開源模型進行 ASR 模型訓練、微調與發展在地化應用,以低成本、高效率進行 ASR 語音應用落地與智慧服務創新。 本專案為數位發展部數位產業署「114年數位產業跨域軟體基盤系統建置案」之實證成果之一,旨在推動台灣語音技術開源生態,協助資訊服務業者強化智慧應用能量,落實在地 AI 技術自主發展,由台灣大哥大執行與維護。
Lightweight multilingual language models for consumer hardware
Don Cheli — SDD Framework. The most comprehensive Specification-Driven Development framework for AI agents. 88+ commands, 51 skills, 15 reasoning models. TDD mandatory, OWASP audit, Autonomous Mode, Crash Recovery, PRD Generator. Works with Claude Code, Gemini/Antigravity, Cursor, Codex, Warp, Amp, OpenCode, Continue.dev. ES/EN/PT.
AI-powered semantic search engine for emojis in 50+ languages, developed in Python
Repository for the paper "MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation)" (NAACL 2022).
Master thesis with code investigating methods for incorporating long-context reasoning in low-resource languages, without the need to pre-train from scratch. We investigated if multilingual models could inherit these properties by making it an Efficient Transformer (s.a. the Longformer architecture).
共 49 条 · 第 2 / 3 页