SardineCon SF/2026

Learn More
Identity verification4 分で読めます

Optical character recognition (OCR)とは?

SUBSCRIBE

Optical character recognition pulls printed text off an ID image to auto-fill and validate fields like name, birth date, and document number. It speeds onboarding and feeds cross-checks against the strip, the chip, and the application data, but a field that reads cleanly is not proof it was not altered.

What is OCR, in plain English?

Optical character recognition reads the printed text on a document image and turns it into structured data. On an ID, it lifts the name, date of birth, document number, and expiry off the photographed page so the system can auto-fill the application and run checks without the user typing everything by hand. It is the step that converts a picture of a passport into fields a computer can work with.

Its value is speed and consistency. By extracting the data automatically, OCR cuts onboarding friction and reduces typos, and it feeds the cross-checks that catch fraud: comparing the printed page against the MRZ strip, against the chip, and against what the applicant entered. When those sources agree, confidence rises; when they disagree, you have a lead.

The catch is that OCR only reports what it sees. Errors from poor image quality or clever edits can seed mistakes further down the line, and a cleanly read field is not evidence it was genuine. In fraud and AML, OCR is a data-extraction tool, not a verdict on authenticity, so its output has to be reconciled against the other zones and sources rather than trusted on its own.

Where OCR fits in a document check

  1. Capture — Photograph the document. The user snaps the ID, and image quality sets how reliable the read will be.
  2. Extract — Read the printed fields. OCR lifts name, date of birth, and number off the visual page into structured data.
  3. Reconcile — Compare the sources. The extracted data is checked against the MRZ, the chip, and the application entry.
  4. Decide — Flag any disagreement. Agreement builds confidence; a mismatch between sources points to error or tampering.

Who is involved?

Who

Their role

The applicant

Submits the document image that OCR reads to auto-fill the application.

The IDV vendor

Runs OCR, structures the fields, and reconciles them with the strip, chip, and entry.

The fraud analyst

Reviews mismatches and low-confidence reads that OCR alone cannot resolve.

The forger

Edits a printed field, hoping OCR reads the altered value as if it were genuine.

What it looks like in practice

In practice

An applicant photographs their ID and OCR auto-fills the form: name, date of birth, and document number appear instantly, and onboarding feels effortless. The read is clean and confident, so the fields look trustworthy at a glance.

But when the system reconciles the OCR against the MRZ strip, the document number differs by one digit. OCR faithfully read a printed page that had been edited, so a clean read masked a forged field. Because the flow cross-checked the extracted data against the strip and chip rather than trusting the read alone, the mismatch surfaced and the case went to review instead of approval.

Why it matters to operators

OCR is the quiet workhorse of digital onboarding. It removes typing, cuts drop-off, and produces the structured data that every downstream check depends on. Without it, comparing the page to the strip, the chip, and the application would be slow and manual, so it is what makes automated document verification practical at scale.

The risk is trusting the read as if it were the truth. OCR reports what is printed, including a value a forger changed, and poor image quality can introduce its own errors that ripple downstream. Reconcile the output against the other zones and sources: a field that reads cleanly is not proof it was not altered before capture, so the cross-check, not the read, is where the assurance comes from.

What to watch in the data

  • OCR versus strip mismatch. A printed field that disagrees with the MRZ or chip suggests an edit, not a misread, and warrants review.
  • Low-confidence reads. Fields flagged as uncertain from blur or glare can seed downstream errors if accepted blindly.
  • Too-clean edits. A single field that reads perfectly while the surrounding print is noisy can hint at a pasted alteration.
  • Entry that overrides OCR. An applicant manually changing an auto-filled field away from the document value deserves scrutiny.
  • Repeated document data. The same extracted number or name across many applications points to reuse rather than genuine new customers.

Quick questions

Does OCR check whether a document is genuine?

No. OCR only extracts the printed text. Judging authenticity is the job of document authentication and cross-checks against the strip and chip. A clean read says nothing about whether the field was altered.

Why reconcile OCR against the MRZ and chip?

Because those sources should carry the same data. If the printed page was edited but the strip or chip was not, the OCR value will disagree with them, and that mismatch is a strong tampering signal a read alone would miss.

What causes OCR errors?

Mostly poor image quality: blur, glare, low resolution, or bad cropping. Deliberate edits can also produce a value OCR reads faithfully. Both can seed mistakes downstream, which is why low-confidence reads should be flagged, not accepted blindly.

How does OCR speed up onboarding?

By auto-filling the application from the document image, it removes manual typing and reduces typos. That lowers friction and drop-off while producing structured data the rest of the verification pipeline can check.

Can OCR be fooled by an edited document?

Yes. OCR reads whatever is printed, so an altered field is extracted as if it were real. The defense is not better reading but reconciliation against the strip, chip, and application data to catch the disagreement.

Go deeper

  • NIST Digital Identity Guidelines (SP 800-63) ↗ — The US standard for identity proofing and authentication assurance levels.
  • FATF ↗ — The global standard-setter for AML, counter-terrorist-financing, and counter-proliferation. Recommendations, guidance, and jurisdiction lists.

Optical character recognition (OCR)と併せて知っておきたい用語