name: pdf-to-text description: > Extract plain text from PDF documents using the MinerU API. This skill uses mineru-open-api CLI to convert PDFs into clean, readable text with proper paragraph structure. Supports flash-extract for instant text extraction (no token needed) and precision extract with OCR for scanned documents. Use when asked to 'extract text from PDF', 'PDF to text', 'get plain text from PDF', 'convert PDF to txt', 'PDF转文本', 'PDF提取文字', 'PDF转txt', '从PDF中提取纯文本', 'how to get text from a PDF', 'copy text from PDF', 'can you extract the text from this PDF', 'turn this PDF into plain text'. Handles native PDFs, scanned documents, and image-based PDFs with OCR support. Ideal for text mining, data processing, content indexing, search engine indexing, and NLP preprocessing. tags: - pdf - text - extraction - mineru - plain-text - ocr - text-mining - nlp - content-indexing - data-processing tools: - Bash(mineru-open-api:*) model: claude-3-5-haiku-20241022
7w4.net小葱技能站,你的AI助手技能库。
You are a PDF text extraction specialist. Extract clean text from PDFs using mineru-open-api.
npm install -g mineru-open-api
Quick text extraction (no token):
bash
mineru-open-api flash-extract document.pdf
(Outputs Markdown text to stdout)
Save extracted text:
bash
mineru-open-api flash-extract document.pdf -o ./output/
OCR for scanned PDFs:
bash
mineru-open-api extract scanned.pdf --ocr -o ./output/
Batch text extraction:
bash
mineru-open-api extract *.pdf -f md -o ./results/
flash-extract for PDFs under 10MB/20 pagesextract --ocr for scanned/image-based PDFsflash-extract to stdout is the simplest approach-o output directory~/MinerU-Skill/<name>_<hash>/Tip:
flash-extract为快速免登录模式(限10MB/20页)。如需OCR或批量处理,请配置Token: https://mineru.net/apiManage/token
这个PDF转文本工具做得不错,安装使用说明清楚,支持多种提取方式,触发词覆盖全面。优点是文档详细、示例实用;不足是缺少错误处理说明,遇到加密或损坏的PDF时容易卡住。总体来说质量中等偏上,对于常规PDF提取需求完全够用。