Scanned and image-only PDFs (OCR)
RemakeCV runs OCR automatically when a PDF has no usable text layer. Results take longer and warrant closer review than text-based files.
When a PDF has no usable text layer, RemakeCV runs optical character recognition over the pages automatically. You can also force it with the Image Process option, which is worth doing when the candidate's name sits in a header graphic. OCR takes longer and is more prone to character errors, so review those CVs more carefully.
When does OCR run?
Automatically, in three situations:
- The PDF has no text layer at all — a scan or a photograph of a printed CV. OCR happens transparently during extraction.
- Fewer than 50 words are extracted from a PDF, which indicates a text layer that is present but broken.
- The extracted text is in a jumbled reading order — see jumbled and multi-column CVs.
Checks 2 and 3 apply to PDFs only. A DOCX is never re-examined and never falls back to OCR.
Can I force OCR myself?
Yes, and sometimes you should. Alongside the default Standard Process, there is an Image Process option — settable for a whole upload or per file — plus a reprocess as image action on a CV you have already run.
Force Image Process when a CV's header is a graphic. The text layer of a designed PDF often omits header content entirely, which means the candidate's name can be missing from the extraction while the rest of the CV parses perfectly. OCR reads the header as a human sees it.
How do I know OCR was used?
Less reliably than you would hope, so do not build a process that depends on it.
The processing_method field reports text or image. Two caveats matter:
- An image-only PDF is often still labelled
text. When the text layer is empty, OCR runs inside the extraction step and returns plenty of words — so the fallback threshold never trips and the CV is recorded astext. - In the web app and CV history, the value stored is the mode that was requested, not the one the server resolved.
If you are integrating via the API, do not rely on processing_method alone to route CVs into human review. It under-reports OCR. Treat any PDF from an unknown source as potentially OCR'd.
What should I check on an OCR'd CV?
In order of how often they go wrong:
The candidate's name
Names frequently sit inside a header graphic or stylised banner. This is the single most common casualty — on the text path the name can be missing entirely rather than merely wrong.
Email addresses and phone numbers
A misread digit or a dot read as a comma makes the contact details useless — and you will not find out until the client cannot reach the candidate.
Dates
Digits are where OCR errors concentrate.
2013and2018are one character apart.Registration and licence numbers
Anything with legal or safety weight. See custom extraction fields.
Run-together words
OCR sometimes loses word boundaries. RemakeCV corrects obvious cases, but check anything that reads oddly.
Why is a scanned CV slower?
OCR is a separate pass over every page image before parsing can begin, and it is the dominant cost — roughly four seconds on a typical document, more on long ones. A batch of scanned CVs takes substantially longer than the same number of Word files.
Very tall pages
CVs exported from Canva, Notion and web page-builders are sometimes a single continuous page metres long rather than a set of A4 pages. Rendered whole, such a page comes out too small to read.
RemakeCV detects unusually tall pages and slices them into overlapping windows before OCR, then stitches the results. This handles the great majority of such documents, but an extremely long page can exceed the slice limit and lose content from the end.
If a continuous-scroll CV comes back with the later sections missing, that is the cause. Ask for a paginated PDF or a Word version.
How do I get better results?
| Source quality | Advice |
|---|---|
| Word version available | Ask for it. This beats every other option |
| Clean flatbed scan | Usually fine |
| Phone photo of a printed CV | Ask for a proper scan or the original file |
| Skewed, shadowed or low-contrast scan | Reject it — correcting the output takes longer than requesting a better copy |
| Scanned handwriting | Not reliably recoverable |
The most effective fix is not technical. Asking the candidate for the Word version of their CV moves the document from the least reliable input category to the most reliable one, in one email.
Frequently asked questions
- Do I need to do anything to enable OCR?
- Usually not — RemakeCV switches to OCR automatically when a PDF has no usable text layer. You can also force it with the Image Process option, which helps when the candidate's name sits in a header graphic.
- Can I tell whether OCR was used?
- Not reliably. The processing_method field under-reports OCR, because an image-only PDF is often recorded as 'text'. Treat any PDF from an unknown source as potentially OCR'd.
- Why is the candidate's name missing or wrong?
- Names often sit inside header graphics. The text layer of a designed PDF frequently omits them entirely — reprocessing with Image Process usually recovers the name.
Related articles
Last updated . Still stuck? Email support@remakecv.com or book a call.