← All guides·Guide 3 of 11
A PDF that is perfectly legible to you can be partly or wholly unreadable to a language model, and nothing in the answer will say so. This guide explains why, shows how to convert a scanned bundle into text that can be trusted, and sets out the checks that find the pages the conversion got wrong.
_hopper/, connected in Cowork, and an inventory that says which PDFs are scans. A copy of Adobe Acrobat Pro is useful for large bundles but not essential.What a PDF is to a computer
Open a PDF and you see an affidavit. The computer opens the same file and sees one of two things, and the two look identical on screen.

A born-digital PDF, made by saving a Word document or downloading a judgment from SAFLII, carries its text as text. A program can pull the words out in the right order, with the paragraph numbers, and a language model can read them.
A scanned PDF, made by a photocopier or a phone, is a picture of a page. Unless the scanner also ran OCR (some do), it contains no words at all, only dots. To you it is a page of an affidavit. To the computer it is ink at coordinates. A language model given this file has two options: look at the picture and read it as a person would, which it can do but not reliably across hundreds of pages, or find no text and carry on regardless.
Most court files are mixtures. The notice of motion was typed and saved; the annexures were scanned; the record came from the transcribers as text, with the exhibits photocopied in. The inventory from Guide 2 tells you which is which.
What OCR does, and what it does not tell you
OCR, optical character recognition, turns the picture of the words into words. It looks at the grid of dots, matches shapes to letters, and writes an invisible text layer behind the picture. Afterwards the page looks the same, but the text can be searched, copied and read by a program. OCR is old, and on a clean scan of a typed page it is very good.

It goes wrong in predictable places.

A court stamp over the text. A watermark (COPY, DRAFT, PRIVILEGED) printed across the paragraph that matters. Handwriting, including a second applicant added in the margin of a form. A faint fax or a bad photocopy. A page scanned sideways. A table, whose columns are read across instead of down. Characters that look alike: a 0 read as an O, a 1 as an I or an l, a 5 as an S, in a case number or an account number.
The important point is what OCR does not do. It does not tell you which pages it got wrong. A page it read well and a page it read badly both come back as text. That is why the conversion is only half of this guide; the other half is finding the pages that need a human look.
What Claude does on its own
It helps to know what happens when nothing special is done.
If you attach a born-digital PDF to a chat, Claude reads its text layer. If you attach a scanned PDF, Claude looks at the pages as pictures and reads what it can. For a short document this works well. For a long one it does not: there is a limit to how much a single conversation can hold, and a long scan reaches it without warning, so the answer comes back as a summary of the part that was read, presented as a summary of the whole.
In Cowork, with the folder connected, Claude has more tools. Its working environment includes programs for pulling the text layer out of a PDF and, at the time of writing, an OCR engine (Tesseract) that it can run on scanned pages when asked. It can also turn each page into an image and look at it. None of this happens automatically. If you ask for a summary of a scanned bundle without asking for the conversion first, you may get a summary of the pages that happened to be readable.
So the sequence is: convert every scan to text; find the pages the conversion probably got wrong; look at those pages; and only then ask questions about the documents.
Step 1 — Find the scans
From the inventory in Guide 2, list the PDFs with no text layer on some or all of their pages. If you did not make the inventory, ask now:
For every PDF in _hopper/, tell me whether it has a text layer on every page, on some pages, or on none. Give me a table: file, pages, pages with text, pages without. Do not change any file.
A PDF with text on some pages and not on others is common (a typed pleading with scanned annexures). It goes through Step 2 with the scans.
Step 2 — Convert the scans to text
There are two practical routes. Use the first for a large bundle and the second for anything else.
Adobe Acrobat Pro. Open the PDF, use the Scan & OCR tool and choose Recognize Text. Acrobat writes a text layer into the PDF and saves it. It can be set to process a whole folder at once, which is the fastest way through a several-hundred-page record. Save the OCR’d copies into record/ (or the appropriate subfolder) and leave the originals in the hopper untouched.
Claude, in Cowork. Ask Claude to do the conversion in the folder. Copy this prompt, replacing the words in square brackets.
For every PDF in [_hopper/] that has no text layer, or an unreliable one, on any of its pages (whole scans and mixed files alike): make an image of every page; use the text layer for pages where it exists and matches the page image, and run OCR on every other page; and save the result as a text file in [record/] with the same name as the PDF, with one heading per page ("## Page 3") so that page references survive. Keep the page images in a folder next to the text file. You may read the original PDF; do not change, rename, move or delete it. At the top of each text file, list which pages came from the text layer and which from OCR, and the pages that have stamps, handwriting, watermarks, rotated text or tables. Mark anything you could not read as [?]. When you are done, give me one table: file, pages, pages OCR'd, pages with problems, number of [?] marks.
Whichever route you use, the originals stay in the hopper and the readable versions go in the working subfolders. If a conversion goes wrong, you can always start again from the original.
Step 3 — Find the pages that need a human look
This is the step that OCR cannot do for itself. Ask for a list of the risky pages:
For each converted document in [record/], go through the page images and flag every page that has: rotated or sideways text; a table; more than one column; a stamp, signature or handwriting; a watermark; a faint or dark scan; or identifiers that are easy to misread (0/O, 1/I, 5/S). Give me a table: file, page, flag, and one line on what is on the page. Do not change any file.
Then, for the flagged pages, ask Claude to compare the text with the picture:
For every flagged page, open the page image next to the extracted text and compare them line by line. Where they differ, correct the text from the image and record the correction. Where you cannot read the image either, mark the spot [?]. Give me a table of every correction and every [?], with the file and page.
For a typical bundle this turns a few hundred pages into a short list of pages that someone must look at with the original open. That list is the point. The worked example PDF intake shows what this looks like on a made-up forensic report, and Guide 4 turns the routine into a skill so that it runs the same way every time.
Step 4 — Settle what could not be read
The [?] marks are not a nuisance to be tidied away. Each one is a word or number that nobody has read yet. Some can be settled from a cleaner copy elsewhere in the bundle: a smudged case number on a stamped cover page is typed clearly in the heading of the founding affidavit. Some can only be settled by a person looking at the original.
Ask Claude to go through them one at a time, showing you the spot:
Go through every [?] in [record/]. For each one, show me a crop of the page image at that spot, your best reading, how confident you are, and why the word matters. Ask me one at a time and do not write anything into the text until I answer. Where a clearer copy of the same information exists elsewhere in the bundle, tell me where and what it says.
The worked example the legibility check shows this exchange on a made-up referral form. The rules that matter: no silent guessing, no silent correction of the original’s typing errors, and an honest [unreadable] where nothing can be done.
Step 5 — Sort the hopper
Only now, with every document readable, ask Claude to sort. Copies, not originals:
Using the converted documents in [record/] and the originals in [_hopper/], copy each document into the right subfolder: pleadings/, record/, correspondence/, or research/ for judgments and legislation. Do not move or change anything in _hopper/. Give me a table of what went where, and list anything you were not sure about instead of guessing.
Check that nothing went missing between the hopper and the subfolders: the count of documents should match the inventory.
Check yourself
At the end of this guide every scanned PDF in the matter should have a readable text version in the working folders, with its page images kept next to it; every document should have a short note of which pages were checked against the picture and which readings are still uncertain; the [?] list should be settled or assigned to someone to settle; the originals in the hopper should be untouched; and the working subfolders should hold every document, sorted.
The rule that comes out of this guide is simple to state: text pulled out of a PDF is not trusted until someone, or Claude on your instruction, has compared it with the picture of the page. Everything Claude does with the matter afterwards rests on that.
Sources: Anthropic documentation, Get started with Claude Cowork (as at 24 September 2026). Adobe, Recognize text in scanned documents (Acrobat help).