, ,

OCR Is Not the Law: How to Verify 7,765 Scanned State Acts Safely

A practical verification guide for lawyers, researchers and legal-AI builders using 7,765 State Acts recovered from scans. Learn what OCR can find, what it cannot prove and how to cite the source safely.

·

·

eCourtsIndia Knowledgebase

OCR is not the law: how to verify 7,765 scanned State Acts safely

OCR text can help you find a State-law provision, but it should not be quoted as authoritative law until it has been checked against the source scan. IndiaCode has recovered 185,941 searchable sections from 7,765 State Acts that government repositories supplied as scanned documents. That unlocks discovery at national scale—and creates a duty to preserve the warning, scan and verification status wherever the text travels.

Verification guide checked 25 September 2026. This article concerns discovery and citation hygiene, not the legal effect or amendment status of any particular State enactment.

The short rule

Use OCR to locate the law. Use the Gazette or official scan to verify the words. Use amendment and commencement sources to decide whether those words operated at the relevant time.

Key takeaways

  • OCR is a transcription process, not a source of legal authority.
  • A correct-looking sentence can still contain a missing proviso, wrong number, broken table or heading from the next section.
  • The safest citation identifies the Act, provision, jurisdiction, Gazette or scan, page and verification date.
  • “As on” or “page updated” does not prove that every amendment has been incorporated.
  • AI systems must carry the OCR warning and scan provenance in JSON, Markdown, XML and structured data—not only on the HTML page.
  • If verification is impossible, say that clearly and cite the document rather than asserting the extracted text.
Seven-step workflow for verifying OCR State Acts against official or preserved source scans
Use OCR to find the law, then verify the source. Do not quote unverified OCR as settled statutory wording.

What OCR actually does

Optical character recognition turns page images into characters. A legal-document pipeline then tries to identify the title, chapter, section number, heading, body, provisos, explanations, schedules and footnotes. This works well on clean digital scans and less well on old Gazettes, skewed pages, bilingual columns, stamps, faint type and complex tables.

That distinction matters because a PDF can be public and still be practically unsearchable. Without OCR, a researcher must know which document to open and then inspect it page by page. With OCR, a phrase, section number or subject can lead directly to a candidate provision. The retrieval gain is enormous; the legal status of the underlying scan does not change.

Six common OCR failures in legislation

  1. Character substitutions: “1” becomes “l,” “0” becomes “O,” or a currency symbol disappears.
  2. Boundary spillover: the next section heading is appended to the current section.
  3. Missing provisos: an indented proviso or explanation is skipped because the layout differs from the main paragraph.
  4. Broken numbering: clauses, sub-clauses and Roman numerals lose their hierarchy.
  5. Table collapse: rows and columns are read in the wrong order, changing the relationship between a rate and its description.
  6. Marginal-note confusion: headings, page furniture or amendment notes are inserted into the statutory body.

These are not hypothetical. On a sampled Assam Goods and Services Tax Act, 2017, section 4 page, the extracted body ends with “Powers of officers,” which is the heading of section 5. An independently published copy does not include that phrase in section 4. The error is easy for a human to recognize when the neighbouring section and scan are available; it is easy for a retrieval model to quote without context.

Four separate questions that researchers often collapse

QuestionWhat answers itWhat OCR can establish
Was this text published?Gazette or competent official publicationNothing by itself
What does the source page appear to say?Human reading of the scanA candidate transcription
Is this the current consolidated wording?Original Act plus amendments and commencement materialNot reliably, unless consolidation is independently verified
Did this provision apply on a particular date?Commencement, amendment, repeal, savings and factual periodNot by text extraction alone

A page can therefore be an accurate transcription of an old Gazette and still not represent the law in force today. Conversely, an OCR defect does not invalidate the official source document; it only makes the transcription unreliable.

A safe seven-step verification workflow

  1. Identify the exact instrument. Record the short title, Act number and year, jurisdiction and provision.
  2. Check the text-source label. Determine whether the page is source text, OCR text or listed-only.
  3. Open the linked scan. Use the official URL where it remains available; a preserved mirror is useful when a government link has moved.
  4. Locate the page. Compare the heading, subsection numbering, provisos, explanations, punctuation and schedules.
  5. Check amendments and commencement. Search the State Gazette and relevant department for later instruments.
  6. Record the verification. Note the scan, page, date checked and amendment cutoff.
  7. Report discrepancies. Send the exact URL and corrected reading to hi@ecourtsindia.com.

A citation format for an OCR-derived provision

The pleading, opinion or article should cite the enactment—not the OCR process—as the legal source. A useful research note can then disclose the transcription:

[Act title], [Act number] of [year], section [number], [State Gazette issue/date/page]. Source scan: [official or preserved URL]. Searchable transcription: [IndiaCode URL], labelled OCR and verified against scan page [x] on [date]. Amendments incorporated through: [date or “not independently verified”].

If the Gazette details or page cannot be confirmed, do not manufacture them. State exactly what is known: “Searchable OCR transcription from a linked government scan; wording and amendment currency not independently verified.”

What a machine-readable record must carry

A warning banner in HTML protects only the human who opens that page. Legal AI may consume JSON, Markdown, XML, a bulk CSV or a vector index. Provenance must travel with the text in every format.

  • text_source: source text, OCR, manual transcription or unknown.
  • verification_status: unverified, machine-checked, human-checked or verified against Gazette.
  • Official source URL and preserved scan URL.
  • Gazette number, date and scan page.
  • Document checksum so the cited scan can be identified later.
  • OCR engine/version and confidence where meaningful.
  • Corpus retrieval date and a separate amendment-consolidation cutoff.
  • Known defects and correction history.

A live sample currently shows why this is urgent: the HTML page discloses OCR provenance, while the corresponding JSON response does not return the OCR flag, scan or page. Markdown, XML and JSON-LD also need equivalent qualification before unsupervised agents are encouraged to quote the text.

How a legal-AI answer should behave

When the source is verified text

The answer can quote the relevant words, identify the Act and section, link the authoritative source and state the applicable date. It should still avoid personalized advice unless the necessary facts are known.

When the source is unverified OCR

The answer should summarize cautiously, label the text as OCR-derived, link the scan and ask the user to verify the wording before relying on it. For high-stakes drafting, the agent should stop at retrieval and require a scan check.

When only a listed scan exists

The answer should report that no section text is published, link the document and avoid inferring that the requested provision is absent. Missing from an index is not the same as nonexistent in law.

When not to rely on OCR without expert review

  • Criminal offences, penalties, limitation periods or jurisdiction clauses.
  • Tax rates, tariff tables, thresholds and exemption conditions.
  • Provisos, explanations, exceptions and deeming provisions.
  • Repeal, savings, commencement and transitional clauses.
  • Forms, schedules and multi-column tables.
  • Any text being quoted in a pleading, legal opinion, compliance decision or automated eligibility system.

Why publishing the limitation builds authority

A legal database does not become more trustworthy by hiding uncertainty. Explicitly distinguishing source text, OCR and listed-only documents lets lawyers decide what verification is proportionate. It lets developers exclude unverified text from high-risk workflows. It also gives answer engines a reason to preserve the caveat when citing the page.

The national scale is described in our companion analysis, India’s State-Law Digitisation Gap: What 9,175 Acts and 224,367 Sections Reveal. That article measures the corpus; this guide explains how to use one recovered provision responsibly.

Read next and test the source chain

Frequently Asked Questions

Is OCR text legally valid?

OCR has no independent legal authority. It is a machine transcription that can help locate words in a scan. Verify the wording against the Gazette or competent official publication before citing or relying on it.

How do I know whether an IndiaCode State section is OCR-derived?

Check the text-source warning on the page and follow the linked source scan. For high-stakes use, confirm that the same provenance is present in any JSON, Markdown or XML record you consume; current machine-format coverage is being improved.

Does a recent page date mean the State Act is fully amended?

No. A corpus or page update date does not prove that every amendment has been consolidated. Look for a specific amendment cutoff and verify later Acts, notifications and commencement material.

Can I cite an IndiaCode OCR URL in a pleading?

Cite the Act, provision and authoritative Gazette or official scan. You may add the IndiaCode URL as a convenient searchable transcription, clearly labelled OCR, but it should not replace the authoritative source.

What should an AI do when only OCR text is available?

It should disclose that the text is OCR-derived, link the scan, avoid definitive quotation until verification and preserve the source, page, verification status and amendment cutoff with the answer.

eCourtsIndia is a private legal-technology platform. It is not affiliated with, associated with, or endorsed by the Government of India, the Supreme Court of India or its e-Committee, or any court. Official case information is published on ecourts.gov.in. Always verify details against official court records or certified copies. This article is general information, not legal advice. Spotted an error? Write to support@ecourtsindia.com.

Search 32 crore+ Indian court case records, free

One search across the Supreme Court, all 25 High Courts, district courts and 18 tribunal and commission types. Hearing alerts, AI summaries and an API for developers.