Docento.app
Notebook, pen and laptop
All Posts

Preparing Long PDFs for AI Chat: Splitting, Trimming and Converting

By The Docento.app TeamPublished 3 min read
Try Docento's free PDF editor — No sign-up, 100% private — sign, annotate, and stamp PDFs in your browser.Open the editor

Modern AI assistants advertise context windows of hundreds of thousands of tokens, enough to swallow a whole annual report. In practice, answers about long documents still go wrong in predictable ways: details from the middle get missed, tables come back scrambled, and page references drift. A little preparation fixes most of that.

Why big context windows are not the whole answer

Fitting a document into the window is not the same as reading it carefully. Models pay more attention to some parts of a long input than others, and every irrelevant page (cover, legal boilerplate, appendices you do not care about) competes with the pages you do. A focused 20-page excerpt nearly always beats a 300-page dump.

There are also practical limits: file upload caps, per-message limits in chat tools, and cost if you pay per token through an API.

Step 1: cut what you do not need

Before uploading, remove pages that do not bear on your question:

  • Covers, tables of contents, blank pages and indexes.
  • Appendices and annexes, unless your question is about them.
  • Repeated boilerplate such as standard terms you already know.

See how to delete PDF pages and how to extract pages from a PDF. Browser tools like Docento.app do this locally, which matters if the document is confidential.

Step 2: split by topic, not by page count

If the whole document is relevant, split it into logical sections, such as chapters or contract schedules, and ask questions section by section. Splitting by bookmarks is fastest when the PDF has them; see how to split a PDF by bookmarks.

Step 3: check the text layer

Scanned PDFs need OCR. Some assistants OCR automatically, but quality varies. Run OCR yourself first and check a page by copying its text. See how to make a PDF searchable with OCR.

Step 4: convert tables you care about

If your question depends on numbers in tables, extract those tables to CSV or a spreadsheet and give the AI the table directly. Tables are the most common source of wrong answers. See how to convert a PDF to CSV.

Step 5: consider Markdown

Converting a PDF to Markdown keeps headings and lists while removing layout noise like headers, footers and page numbers repeated on every page. Many retrieval systems convert to Markdown internally for this reason. See how to convert a PDF to Markdown.

Step 6: ask for page references

Ask the assistant to cite the page or section for every claim. Then spot-check two or three. If a citation is wrong, the answer around it deserves suspicion too. Prompts that ask for quotes rather than paraphrases are easier to verify. See prompt engineering for PDF tasks.

Mind the privacy question

Trimming a document before uploading also reduces what you share. If only three pages matter, only three pages need to leave your device. For sensitive files, read risks of using AI on confidential PDFs first.

Takeaway

Bigger context windows made it possible to upload whole documents, not wise. Trim, split, check the text, and pull tables out separately, and AI answers about long PDFs become dramatically more reliable.

Try Docento's free PDF editor

No sign-up, 100% private — sign, annotate, and stamp PDFs in your browser.

Open the editor

Related Posts