← BlogAI

Twenty Prompts That Get Real Answers Out of a PDF

PDFalot Editorial Team·May 5, 2026·12 min read
AI illustration for the article: Twenty Prompts That Get Real Answers Out of a PDF

Most people ask a document to 'summarise this'. Here are the prompts that actually extract usable information.

Chatting with a document is only as good as the question. 'Summarise this' produces a bland paragraph you could have guessed. Specific, structured prompts produce output you can act on. Here are the ones that consistently earn their keep, grouped by what you are trying to do.

Understanding a document fast

  • 'In five bullets, what is this document asking me to do or decide?'
  • 'What would a sceptical reader push back on, and where in the document is the weakest support?'
  • 'List every claim in this document that is stated without evidence, with page numbers.'
  • 'Explain the core argument to someone with no background in this field, in under 150 words.'
  • 'What is deliberately not addressed in this document?'

Contracts and agreements

  • 'List every obligation that falls on me, with the clause number and deadline.'
  • 'Which clauses shift risk, cap liability, or create automatic renewal? Quote them exactly.'
  • 'What happens if I terminate early? Quote the relevant text.'
  • 'Compare the termination rights of each party — is the arrangement symmetric?'
  • 'Identify any defined term that is used but never defined.'

Studying

  • 'Write ten exam-style questions a lecturer would ask on this chapter, then answer them with page citations.'
  • 'What are the three concepts here that students most commonly confuse, and how do they differ?'
  • 'Turn the key definitions into flashcard pairs: term on one side, one-sentence definition on the other.'
  • 'Which parts of this chapter depend on material from earlier chapters I should revise first?'

Reports and data

  • 'Extract every number that appears with a unit or currency, and say what it measures.'
  • 'What changed compared to the previous period, and by how much?'
  • 'List every forward-looking statement and flag which are quantified.'
  • 'Where does the document use a percentage without stating the base?'

Two prompt habits that improve everything

Always ask for citations

Adding 'quote the exact text and give the page number' to any prompt changes the output materially. It forces the answer to be grounded in the document rather than in the model's general knowledge, and it gives you a two-second way to verify each claim.

Ask for the shape of the answer

'As a table with columns X, Y, Z' or 'as a numbered list, one line each' produces output you can paste straight into a spreadsheet or a memo. Unstructured prose is almost always the wrong format for a working document.

What to stay sceptical about

Ask for something that is not in the document and you will often get a confident answer anyway. Test any new document with one question you know it does not answer. If the model invents something, tighten your prompts with 'if the document does not say, reply: not stated'. That single instruction removes most of the risk.

A worked example: turning a 40-page vendor contract into a decision

Instead of asking a single vague question, run a short chain of four prompts in sequence and read each answer before sending the next. First: 'List every clause where I owe the vendor money, with the clause number, amount or formula, and trigger condition, quoted exactly.' Second, once you have that list: 'Of these, which are triggered automatically versus which require the vendor to take an action first?' Third: 'List every clause that lets either party terminate, the notice period required, and any penalty for early termination.' Fourth: 'Given the above, in three sentences, what is the actual financial exposure if we terminate at month six?' Each answer becomes the grounding for the next question, which keeps the model anchored to what it already found rather than re-reading the whole document from scratch and drifting toward a generic answer.

Prompt patterns worth copying verbatim

  • 'Quote the exact sentence, then explain it in plain English on the next line, for every clause in Section 4.'
  • 'Act as a hostile auditor. What in this report would you flag as understated or overstated, and why?'
  • 'Produce a table with columns: Clause, Obligation, Owed By, Deadline, Quoted Text.'
  • 'Rewrite the answer as if explaining to someone who will act on it in the next ten minutes and has not read the document.'
  • 'List every date mentioned in the document and what happens on that date.'
  • 'If two sections of this document appear to contradict each other, quote both and explain the contradiction.'

Diagnosing a bad or vague answer

When an answer feels thin, the fault is almost always the prompt, not the document. A generic answer to 'what are the risks in this document' usually means the model defaulted to generic risk language instead of reading closely. Rephrase as 'quote the three sentences in this document that a lawyer reviewing it for risk would highlight first, and explain what specifically concerns each one' — forcing a quote and a specific concern per item consistently produces a sharper answer than asking for 'the risks' in the abstract.

Common failure modes and how they show up

  1. Fabricated numbers: a total or percentage appears in the answer but does not exist anywhere in the source text — always cross-check any number against the document before acting on it.
  2. Section bleed: an answer describes what a nearby but different clause says, because two topics are visually close in the layout — ask for the exact clause number to catch this.
  3. Silent scope narrowing: on a long document, an answer only covers the first section unless you explicitly ask it to check the whole document, including appendices.
  4. Overconfident synthesis: asked to 'summarise the risk profile', the model may blend the document's actual content with generic domain knowledge about similar documents — the citation requirement is the fix.
  5. Table collapse: asking for structured output on a long list sometimes truncates after a plausible-looking but incomplete set of rows — always ask 'how many rows in total should this table have' as a follow-up.

A short checklist before trusting an answer for a real decision

  1. Did you require a quote and page or clause reference for every substantive claim?
  2. Did you test at least one question the document does not answer, to confirm it says 'not stated' rather than guessing?
  3. Did you ask it to check the whole document explicitly, rather than assume it did?
  4. Did you verify any number in the answer against the visible text yourself?
  5. For anything you will sign, pay, or file: did a human read the quoted source text, not just the model's paraphrase of it?
"A prompt that cannot be wrong is not a useful prompt. Ask questions specific enough that a wrong answer would be obviously wrong."

Multi-document comparison prompts

Comparing two versions of a contract, two competing vendor proposals, or a report against its prior-year equivalent needs a different prompt shape than single-document analysis. 'Compare document A and document B and list every numerical difference' tends to work poorly because the model has no forced structure to anchor the comparison to. A more reliable version: 'For each of the following fields — price, term length, termination notice period, liability cap — state the value in document A, the value in document B, and whether the difference favours us or the counterparty.' Naming the fields explicitly, rather than asking the model to find 'the differences' on its own, is what keeps the comparison from drifting into a vague prose summary.

Prompts for long or scanned documents

On a 200-page report, a single broad question tends to answer from whichever section the model attended to most, usually the first few pages. Break the task into passes: 'List every section heading and its page range' first, so you and the model both have a map of the document, then ask targeted questions naming the specific section by title. For scanned or OCR'd documents, add 'if any number looks like it may be an OCR misread, such as a 5 that could be a 6 or an S, flag it explicitly rather than presenting it as certain' — this single instruction catches a meaningful share of digit-level OCR errors that would otherwise pass through into a confident-sounding answer.

Try it on your own PDF

Upload a document and put these ideas to work in under a minute.

Open PDFalot →

Keep reading

Try AI Now