ocr
157 个项目 · ⭐ 880.8kTransforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Tesseract Open Source OCR Engine (main repository)
OCR software, free and offline. 开源、免费的离线OCR软件。支持截屏/批量导入图片,PDF文档识别,排除水印/页眉页脚,扫描/生成二维码。内置多国语言库。
A community-supported supercharged document management system: scan, index and archive all your documents
Pure Javascript OCR for more than 100 Languages 📖🎉🖥
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
🌈一个跨平台的划词翻译和OCR软件 | A cross-platform software for text translation and recognition.
pix2tex: Using a ViT to convert images of equations into LaTeX code.
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
OCR & Document Extraction using vision models
A fast, helpful, and open-source document parser
OCR model that handles complex tables, forms, handwriting with full layout.
BISHENG is an open LLM devops platform for next generation Enterprise AI applications. Powerful and comprehensive features include: GenAI workflow, RAG, Agent, Unified model management, Evaluation, SFT, Dataset Management, Enterprise-level System Management, Observability and more.
共 157 条 · 第 1 / 8 页