📄 PDF Text Extractor

📄
Drop, paste, or click to browse
PDF files and URLs supported — paste with Ctrl+V / ⌘V
Extensionless or renamed PDFs are fine — the file type is checked by content, not name
All processing happens in your browser — no data leaves your device
No file loaded
Output Format

Format Guide

  • Markdown: Best for both full-context and RAG chunking. Each page has clear headers with filename and page number.
  • JSON: Structured format with metadata. Ideal for programmatic processing and custom chunking strategies.
  • Plain Text: Simple format with page markers. Works with any text processor.
Leave empty to extract all pages
Auto-detect suits most papers — switch to a fixed count or raw order if columns come out interleaved.
Pages stream in as they're extracted. Higher values pipeline through pdf.js's worker — benefit tapers off (single worker thread); elapsed time is shown so you can compare.
Select a PDF file to begin

💡 URL API

You can also fetch PDFs from URLs:

Extracted Text