Plumb a PDF for detailed information about each char, rectangle, line, et cetera — and easily extract text and tables.
-
Updated
Aug 6, 2026 - Python
Plumb a PDF for detailed information about each char, rectangle, line, et cetera — and easily extract text and tables.
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
Polyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus code intelligence for 371 languages. Fifteen bindings, with CLI, REST API, and MCP server.
Table Transformer (TATR) is a deep learning model for extracting tables from unstructured documents (PDFs and images). This is also the official repository for the PubTables-1M dataset and GriTS evaluation metric.
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
PDF to markdown using vision LLMs — tables, layouts, and structure preserved
img2table is a table identification and extraction Python Library for PDF and images, based on OpenCV image processing
Document Layout Analysis resources repos for development with PdfPig.
ParseBench - A Document Parsing Benchmark for AI Agents
Python library to extract tabular data from images and scanned PDFs
Visual document analysis studio powered by Docling — configure the extraction pipeline, inspect text, tables and bounding boxes in the browser, then chunk, embed and index into OpenSearch and Neo4j.
A Curated List of Awesome Table Structure Recognition (TSR) Research. Including models, papers, datasets and codes. Continuously updating.
Extract tables from PDF files (port of tabula-java)
Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
A carefully-designed OCR pipeline for universal boarded table recognition and reconstruction.
✂️ Extract Tables from Microsoft Word Documents with R
Best PDF Converter! PDF to any format, pdf2word/excel/xml/html/txt...
An MCP server that gives your AI agent agentic RAG over your PDFs, one file or a whole folder: hybrid semantic + keyword search, selective page reads, tables, images, OCR, chart data, and multi-column/CJK layouts. The agent decides when to search; pdf-mcp does the retrieval.
CCKS2019评测任务五-公众公司公告信息抽取,第3名
🔍 Table Extraction Tool: A powerful open-source solution combining OCR and computer vision for extracting structured tabular data from images. Ideal for LLM preprocessing, data analysis, and automation. 🚀
To associate your repository with the table-extraction topic, visit your repo's landing page and select "manage topics."