India’s State-law problem is no longer just that legislation is scattered. It is that the available text comes in different reliability tiers. IndiaCode by eCourtsIndia currently lists 9,175 State Acts and 224,367 State-law sections across 36 States and Union Territories. But 7,765 of those Acts—185,941 sections—were recovered from government scans by OCR. They are searchable discovery copies, not substitutes for the Gazette.
Corpus snapshot: 23 September 2026. Source status: eCourtsIndia’s structured edition of government legislation records and scans. The Gazette or competent State publication remains authoritative.
Key takeaways
- IndiaCode’s live corpus reports 9,175 State Acts and 224,367 State sections.
- 7,765 Acts were available upstream as scans and have been converted into 185,941 section-addressable OCR records.
- Another 586 State Acts are listed with a source scan but no published section text because a dependable transcription is not available.
- Searchable text exists for 35 of 36 State and Union Territory jurisdictions. Lakshadweep currently has listed records but no published section-level text.
- State-law availability is not the same as amendment currency. A searchable section can still require verification against later amending Acts, rules, notifications and the Gazette.
- The strongest open-data opportunity is not to call OCR “the law,” but to connect every recovered section to its scan, page, provenance and verification status.

The number hides three different kinds of availability
A total such as 9,175 Acts sounds like one uniform library. It is not. A State statute can sit in one of three practical tiers.
1. Source-provided section text
These Acts arrived with structured or extractable text that could be divided into provisions without reconstructing the document from page images. By subtraction from the live corpus, approximately 1,410 published State Acts and 38,426 State sections fall outside the OCR-recovered set. That does not automatically prove amendment currency, but it gives the text a stronger starting provenance than scan transcription.
2. OCR-recovered section text
The largest tier is the difficult one: 7,765 Acts and 185,941 sections recovered from scanned government documents. OCR turns a page image into searchable characters and then attempts to identify section boundaries. This makes legislation discoverable by title, phrase and provision number. It can also introduce character substitutions, missing lines, merged columns and a heading from the next section at the end of the current one.
A sampled page from the Assam Goods and Services Tax Act, 2017, section 4 demonstrates the risk: the recovered text appends “Powers of officers,” the heading that belongs to the following section. That is a manageable discovery defect when the source scan is visible. It becomes dangerous if a machine-readable endpoint presents the same string without carrying the OCR warning and source page.
3. Listed-only source documents
A further 586 State Acts are indexed by title and source document but have no published section text. This is an important failure state to preserve. Publishing a scan-only record is more honest than silently filling a gap with an unreliable extraction. For these Acts, researchers should cite and read the document itself.
State coverage is large—and highly uneven
The State legislation index exposes coverage by jurisdiction. The largest Act count belongs to Assam, with 1,041 published Acts in the current snapshot. Section and subordinate-instrument density tells a different story: some States have fewer Acts but much deeper provision and rules coverage.
| Jurisdiction | Published Acts | Sections | Rules and instruments |
|---|---|---|---|
| West Bengal | 800 | 17,070 | 1,130 |
| Karnataka | 345 | 14,433 | 877 |
| Maharashtra | 363 | 13,395 | 2,428 |
| Rajasthan | 341 | 11,605 | 1,328 |
| Uttar Pradesh | 330 | 9,092 | 2,425 |
These numbers should not be used to rank which State has “more law.” They reflect publication, collection and extraction histories. One jurisdiction may consolidate aggressively; another may publish many short amending Acts; a third may keep operative rules on departmental websites rather than the legislation portal.
Why State law is harder to digitise than Central law
- Publication is fragmented. Acts, rules, commencement notifications and amendments may sit across the State Gazette, Law Department and subject department.
- Many records are page images. A PDF can be publicly downloadable while remaining invisible to ordinary full-text search.
- Consolidation is inconsistent. The original Act and five amending Acts may be available without a current integrated text.
- Language and typography vary. Bilingual pages, old typefaces, stamps, marginal notes and multi-column layouts lower OCR accuracy.
- Source links decay. Departmental websites move or disappear. A durable mirror can preserve access, but the original publication details still matter.
- Rules carry the operational detail. Reading only the parent Act can miss forms, thresholds, authorities and procedures found in subordinate legislation.
What IndiaCode adds—and what it does not
Government India Code remains the official repository and an upstream source for much of this material. Commercial databases such as SCC Online and Manupatra offer curated State statutes, amendment tools and professional editorial treatment. Indian Kanoon, LatestLaws, Bare Acts Live and PRS also provide valuable State-law discovery.
IndiaCode’s distinct contribution is the combination: permanent provision URLs, searchable recovery of scan-only Acts, links back to source scans, and delivery through HTML, Markdown, Akoma Ntoso XML, keyless JSON, downloadable data and MCP tools. The open-data page and API documentation make the corpus reusable by researchers and legal-technology teams.
The value is not that a private website becomes the authority. The value is that a hard-to-find government scan becomes discoverable, addressable and verifiable.
How lawyers should cite an OCR State Act
- Name the Act, Act number and year, provision and jurisdiction.
- Open the linked government or mirrored source scan and locate the provision on the page.
- Check whether later amending Acts, rules or commencement notifications affect the proposition.
- Cite the Gazette issue, date and page where available.
- Use the IndiaCode section URL as a convenient reading and discovery link, labelled as an OCR transcription if applicable.
- If the scan and transcription disagree, rely on the authoritative publication and report the transcription error to hi@ecourtsindia.com.
How legal AI systems should use the corpus
A legal retrieval system should not treat every returned text field as equally authoritative. It should carry a provenance envelope with the answer: text source, verification status, jurisdiction, source document, scan page, retrieval date and amendment cutoff. If the text is OCR-derived and unverified, the model should use it to find the source—not quote it as settled statutory language.
- Safe: “IndiaCode’s OCR transcription indicates that section 4 concerns appointment of officers; verify the wording in the linked Assam Gazette scan.”
- Unsafe: “Section 4 definitively states…” when the model has only unverified OCR and no scan-page check.
- Safe: reporting that an Act is listed-only and linking the document.
- Unsafe: inferring that a missing section or missing Act does not exist.
This is also a product requirement. OCR warnings must travel consistently through HTML, JSON, Markdown, XML and structured data. A warning visible only to a human reader does not protect an agent consuming the API.
What should be built next
- A State Law Coverage Atlas separating source text, OCR text and listed-only documents.
- Page-level scan anchors, checksums, OCR confidence and human-verification status.
- Jurisdiction filters for State-law search and APIs.
- State-and-subject hubs for rent, land and revenue, court fees, stamps, shops and establishments, professional tax and State GST.
- Act-to-rules and section-to-notification links marked as recorded, inferred or unidentified.
- A public correction and version history for high-use provisions.
Methodology and limitations
The totals in this article come from the live IndiaCode metadata endpoint, the State index, the jurisdictions endpoint, the listed-Acts endpoint and the corpus description in llms.txt, checked on 23 September 2026. OCR status describes the text-generation method; it is not an assessment of whether each provision is currently in force or fully amended. Corpus counts are changing floors, not a claim to contain every State enactment.
Read next and inspect the data
- Browse State and Union Territory legislation
- Download IndiaCode datasets
- Use the keyless legislation API
- IndiaCode vs the official India Code portal
- How to search Indian Acts and retrieve JSON or Markdown
Frequently Asked Questions
How many State Acts are searchable on IndiaCode?
The live corpus listed 9,175 published State Acts and 224,367 State-law sections on 23 September 2026. Of these, 7,765 Acts and 185,941 sections were recovered from government scans using OCR. Browse the current counts at <a href=’https://indiacode.ecourtsindia.com/states/’>IndiaCode State Legislation</a>.
Is OCR text legally authoritative?
No. OCR is a machine transcription of a scanned document and can contain character, layout or section-boundary errors. Use it for discovery, then verify the wording against the Gazette or linked government scan.
Does IndiaCode cover every State and Union Territory?
It lists records across all 36 jurisdictions, with published section text for 35 in the current snapshot. Lakshadweep currently has listed-only records and no published section-level text. Coverage does not prove that every enactment or amendment is included.
What does listed-only mean?
A listed-only Act has a title and source document in the index but no published section text. IndiaCode reported 586 such State Acts. Researchers should open and cite the source document instead of assuming searchable text exists.
Can an AI system quote OCR State-law text?
Only after verifying it against the linked official scan or another authoritative consolidation. A safe system should disclose that the text is OCR-derived and carry the source document, scan page, verification status and amendment cutoff with every answer.
