,

Building India’s Legal Data Engine: 32 Crore+ Records, 38 Structured Fields, and Why Disposed Means Nothing

Legal data engine for Indian courts: how eCourtsIndia structures 32 crore+ records into 38 fields and reads order text to show what a disposal really was.

·

·

eCourtsIndia Knowledgebase

legal data engine cover A

Product engineering · Verified: 14 September 2026 · Index read: 29,02,62,233 court records · Machine surface: 39 MCP tools (as of 23 September 2026)
Every number on this page was read live from the eCourtsIndia index, the live search capability catalogue, or the IndiaCode corpus on the date above. Where a figure is a reading of a moving system, this page says so.

Last updated: 23 September 2026. The live index has since passed 32 crore+ case records; the dated readings below are kept exactly as measured on 14 September 2026, because the point of the piece is how a moving index is read. Tool counts are updated to the current server.

Between two identical calls to our own index, made about a minute apart on 14 September 2026, the record count moved from 29,02,62,233 to 29,02,62,811. Five hundred and seventy eight matters had landed in the time it took to run one query.

That is the first thing worth understanding about the layer underneath eCourtsIndia. It is not a dump of Indian court records. It is a running machine, and the figure on the front of it is a reading rather than a fact. Earlier the same morning, the LegalCheck launch post read 28,95,50,226 from the same index. Eight hours later it read 7,12,007 records higher. Neither number is wrong. They are two frames of the same film.

That launch post is about what a legal background check should do. This one is about the engine it runs on, why building the engine took longer than building any product on top of it, and what changes when court data stops being documents and starts being data.

Key takeaways

  • The index held 29,02,62,233 court records when this page was written (it has since passed 32 crore+), across the Supreme Court, every High Court, district and taluka courts, and 18 tribunal families spanning 21 live court types.
  • Every matter exposes 38 projectable fields, not a blob of text. The index offers 31 sort dimensions, 11 facets, four name matching modes and five party roles kept apart rather than merged.
  • 21,92,90,294 records carry the status DISPOSED. That is 75.5 per cent of the entire corpus sitting behind one word that says nothing about what happened. Reading the order text is the only way past it.
  • Behind the cases sits the statute book: 10,084 Acts, 2,69,602 sections, 464 offence classifications, 2,372 official transition mappings and 1,72,143 cross references, all free and keyless.
  • The hardest job the engine does is not search. It is attribution. A single litigant query for Kumar returns 2,52,71,909 records, which is why identity resolution sits above search rather than inside it.
  • All of it is served to machines through 39 MCP tools, 19 of which never consume credits. Structured data is the precondition for agentic legal AI, not a nice-to-have beside it.
The eCourtsIndia legal data engine ECOURTSINDIA · THE ENGINE · READ LIVE 14 SEPTEMBER 2026 Not a database. An engine. What sits underneath 32 crore+ Indian court records, stage by stage. 29,02,62,233 COURT RECORDS 38 STRUCTURED FIELDS 21 COURT TYPES 8,000+ COURTS, DAILY 2,69,602 STATUTORY SECTIONS 31 SORT DIMENSIONS 39 MCP TOOLS 18 TRIBUNAL FAMILIES FROM A COURT PORTAL PAGE TO A MACHINE READABLE ANSWER 01 SOURCE Supreme Court, every High Court, district and taluka courts, 18 tribunal families, and the daily cause lists of more than 8,000 courts. 02 ACQUIRE, THEN KEEP ACQUIRING A record can be re-pulled from its court on demand. A snapshot freezes the moment it is taken. This one is asked again. 03 STRUCTURE 38 projectable fields on every matter. Five party roles kept apart: litigants, petitioners, respondents, advocates, judges. 04 READ THE ORDER, NOT THE LABEL Order and judgment PDFs converted to text, summarised, and classified by outcome. This is what separates an acquittal from a conviction. 05 MAP THE LAW BEHIND IT 10,084 Acts and 2,69,602 sections, 464 offence classifications, and the official IPC to BNS and CrPC to BNSS transition tables. 06 INDEX FOR QUESTIONS, NOT LOOKUPS 31 sort dimensions, 11 facets, four name matching modes, and full text over every uploaded order and judgment. 07 RESOLVE THE PERSON How many people could this name be, answered before anything is attributed to anyone. Constraints that can only lower a match, never raise one. 08 SERVE IT TO HUMANS AND MACHINES Search and alerts for people. A partner API and 39 MCP tools for software. The same index underneath both. eCourtsIndia · counts read live 14 September 2026 · ecourtsindia.com
The engine in one frame. Every count was read live on 14 September 2026 and every one of them moves.

An index that moves while you are reading it

Most legal data in India is distributed as a snapshot. Someone crawls a set of public portals, writes the result into a table, and ships it. The table is correct for exactly as long as it takes a registrar somewhere to enter the next order.

The failure mode is quiet and specific. A snapshot shows a next hearing date that has already passed. That is not a small staleness problem, because a hearing that has already happened may have produced a warrant, a dismissal in default, a decree, or nothing at all, and the snapshot cannot tell you which. It will keep printing the old date confidently for as long as you keep asking.

Our index is operated rather than harvested. A record can be asked again. Refreshing one matter against its own court, or a batch of them at once, and waiting for the answer to land is the difference between a dataset and an engine. It is also why the count moved twice while this article was being written. That capability is the one most people skip, because it is the expensive one. A crawl is a project with an end date. An index that can be asked again is an operation, with a queue, a failure budget, a way of telling a caller that a particular record did not land and why, and somebody awake when a state portal changes its markup on a Sunday. A record whose next hearing date has slipped into the past is not left sitting there looking authoritative. It goes back to the court, and the answer that comes back is the court’s rather than ours.

Three quiet ways a frozen copy lies WHY A SNAPSHOT IS NOT AN INDEX Three quiet ways a frozen copy lies WHAT A FROZEN COPY SHOWS YOU WHAT AN INDEX THAT CAN BE ASKED AGAIN DOES Next hearing: 12 March 2024 A date in the past, printed with total confidence. Nothing flags it. The record goes back to its court That hearing already happened. A warrant, a decree, a dismissal, or nothing at all. Status: DISPOSED One word carrying 21,92,90,294 records and six different outcomes. The order text is read instead Acquittal, conviction, compromise or a restorable default. The label cannot tell. order-15.pdf Court filenames are positional, and portals reuse them. The date inside the document is checked A superseded order served under a current filename is caught, not quoted. A crawl is a project with an end date. An index is an operation. eCourtsIndia · status counts read live on 14 September 2026
None of these three failures announces itself. Each one produces a confident, well formatted, wrong answer.

When we measured on 14 September 2026, the index held 29 crore records; it has since passed 32 crore+. Here is how that 29 crore reading split across the judicial ladder on that date.

TierRecordsWhat lives here
District and taluka courts23,14,66,178The overwhelming majority of Indian litigation, and the tier most legal technology skips.
High Courts5,16,69,906Writs, appeals, company and commercial benches across every High Court.
Tribunals42,97,822Insolvency, recovery, tax, competition, environment, consumer and more.
Supreme Court11,85,462The apex record, including the orders that set aside everything below.
Read from the live court level facet on 14 September 2026. The tiers do not sum exactly to the total because a small share of records carry no court level assignment, and that gap is disclosed rather than smoothed away.
One index, the whole ladder COURT COVERAGE · READ LIVE 14 SEPTEMBER 2026 One index, the whole ladder 29,02,62,233 records across 21 live court types. Bars are to scale. DISTRICT & TALUKA 23,14,66,178 HIGH COURTS 5,16,69,906 TRIBUNALS 42,97,822 SUPREME COURT 11,85,462 THE TIER MOST LEGAL TECHNOLOGY SKIPS 79.7 per cent of the whole index sits in the district and taluka courts. THE 18 TRIBUNAL FAMILIES, ALL INDEXED NCLT NCLAT ITAT CESTAT DRT DRAT NGT AFT SAT TDSAT APTEL CAT CCI GSTAT RCT SEBI ORDERS GST AAAR e-JAGRITI Plus daily cause lists for more than 8,000 district and taluka courts. eCourtsIndia · tiers do not sum exactly to the total, because a small share of records carry no court level
Insolvency, recovery, tax and competition risk never appears in a district criminal dump. Missing the tribunal line is missing the deal breaker.

The tribunal line is the one people underestimate. Insolvency, debt recovery, tax and competition risk does not appear in a district criminal dump. It appears in NCLT and NCLAT, in DRT and DRAT, in ITAT and CESTAT and the consumer commissions. A diligence product that stops at criminal courts will clear a promoter whose company is in insolvency, and it will do so with a clean looking report.

What structured actually means, in fields you can name

Structured is the most abused word in Indian legal technology. Almost everybody claims it. Very few publish what they mean by it.

Here is what we mean. One matter, read from the live capability catalogue on 14 September 2026, exposes 38 named fields that a caller can ask for individually. Not a page of HTML. Not a paragraph to be regexed. Thirty eight addressable fields, each of which can be filtered, and most of which can be sorted or faceted.

One matter, 38 addressable fields ONE MATTER, READ LIVE 14 SEPTEMBER 2026 38 fields you can ask for by name IDENTITY OF THE MATTER · 5 cnr filingNumber registrationNumber caseType caseStatus WHERE IT SITS · 5 courtCode courtName stateCode districtCode courtLocationFacetPath TIME · 8 filingDate registrationDate firstHearingDate lastHearingDate nextHearingDate decisionDate filingYear decisionYear PEOPLE, FIVE ROLES KEPT APART · 5 petitioners petitionerAdvocates respondents respondentAdvocates judges THE LAW AND THE SHELF · 5 actsAndSections judicialSection caseCategory caseCategoryFacetPath benchType WHAT THE COURT HAS ACTUALLY DONE · 7 hasOrders hasJudgments orderCount interimOrderCount judgmentCount hearingCount iaCount COMPUTED, NOT COPIED · 3 caseDurationDays filingToFirstHearingDays aiKeywords 31 sortable · 11 facetable · 23 presence filterable · 4 name matching modes · 200 results per page
The 38 projectable fields, read from the live search capability catalogue. Party roles are the ones that matter most, because merging them destroys the question.

Look at the fourth group. Petitioners, petitioner advocates, respondents, respondent advocates and judges are five separate fields. Almost every simplified court dataset in circulation flattens these into one party string, and the moment you do that you lose the ability to ask the only questions that matter in diligence. Was this person suing or being sued. Was this advocate on the other side. Has this judge heard this party before.

Fields are what turn a search box into a question engine. Consider this request:

Find every matter in this district where this name appears as a party, and a person of the father’s name also appears, and at least one order is on file, filed after 2019, sorted by how long the case has been running.

That is one call for us. Against a flattened snapshot it is not a query at all. It is a full crawl followed by hope. The partner API guide has the full surface, and the capability catalogue is a live endpoint rather than a document, so nothing in it has to be guessed.

Four name matching modes deserve their own sentence. A name in an Indian court register can be matched as all tokens, any token, an exact ordered phrase, or fuzzily. Those are four genuinely different questions, and a system that offers only one of them has decided on your behalf which kind of error you are going to make.

Disposed is one word for 21,92,90,294 records

Illustration of the eCourtsIndia legal data engine: cases, orders, statutes and identity layers

This is the part of Indian court data that quietly breaks the most products, and it is worth being precise about.

The status facet on our index carries 30 distinct values. Read live on 14 September 2026, they distribute like this: DISPOSED 21,92,90,294, PENDING 6,98,16,110, UNKNOWN 6,50,894, and then a long tail of DISMISSED, ADMITTED, ALLOWED, WITHDRAWN, ABATED and two dozen others, none of which reaches a quarter of a million.

So 75.5 per cent of every court record in India, in our index, sits behind a single word. Disposed does not mean acquitted. It does not mean convicted. It does not mean the claim succeeded. It means the court is finished with the file.

Disposed is one word for 21.9 crore records CASE STATUS FACET · READ LIVE 14 SEPTEMBER 2026 The status field is not an outcome 29,02,62,233 RECORDS, BY STATUS DISPOSED 21,92,90,294 75.5% of the corpus PENDING 6,98,16,110 At least six things that single word can mean Acquitted on the merits The only unambiguously good one. Convicted and sentenced Same word on the status field. Compounded on compromise Parties settled. Nothing decided. Dismissed in default Restorable. Not a finding at all. Committed upward Still live, in another court. Stopped, accused absconding The opposite of resolved. SO WE READ THE TEXT INSTEAD OF THE LABEL 4,40,631 DISPOSED records whose indexed text says acquitted 3,33,516 DISPOSED records whose indexed text says convicted WORKED EXAMPLE · CNR SCIN010108932008 · SUPREME COURT · CRL A 1538/2013 STATUS FIELD SAYS → DISPOSED ORDER TEXT SAYS → High Court had convicted and sentenced the appellant; Supreme Court compounded it and set the conviction aside. eCourtsIndia · all counts read live 14 September 2026 · the label is identical in both directions
Two records can carry the same status and opposite outcomes. Only the order text separates them.

The counts in that graphic are text searches across the full order and judgment text we hold, filtered to records whose status is DISPOSED. They are an illustration of ambiguity rather than a census of outcomes, and we would rather say that plainly than let a chart imply a precision it does not have. The word acquitted appearing in an order does not by itself prove an acquittal. That is exactly the point. The label carries no signal at all, and the text carries signal that has to be read properly.

One real matter, read both ways

Take a real Supreme Court record. CNR SCIN010108932008, Criminal Appeal 1538 of 2013, arising from a Madras High Court criminal appeal of 2002. Fifteen orders on file, five interlocutory applications, two judges named, both advocates named.

Its status field says DISPOSED. That is all a status-reading system gets.

Now read the orders. The High Court had convicted and sentenced the appellant to one month of simple imprisonment and directed him to pay compensation of two lakh sixty thousand rupees, with a default sentence of three further months. The appellant took it to the Supreme Court. During the appeal the parties filed a Joint Memo of Compromise. The Supreme Court disposed of the appeal as compounded, set aside the High Court judgment and conviction, and directed a payment of fifty thousand rupees to the Supreme Court Legal Services Committee.

Three different products would report three different things about the same man. One reading the status reports a disposed criminal appeal and leaves the reader to guess. One reading only the High Court record reports a conviction that no longer exists. One reading the full chain reports the truth, which is that a conviction was recorded and then set aside on a compromise recorded by the apex court.

Getting that third answer requires four things at once: the district and High Court record, the Supreme Court record, the link between them, and the text of the final order. Miss any one and you publish something false about a person. That is the engineering requirement, and it is why we convert order PDFs to text and run structured outcome analysis over them rather than storing a status string and calling it a day.

The reading itself carries a trap that almost nobody sees coming. Court portals name their documents by position, order-1.pdf through order-15.pdf, and they reuse those names. A superseded document can be served under a current order’s filename, which means a pipeline that trusts the filename will quote the wrong order on the right case, confidently, with a citation that looks perfect. So every document we read is checked against the date printed inside it, and the answer either confirms the date or names both dates and refuses to pretend. No customer has ever asked for that. It is the kind of thing you only build after it has bitten you, and a hundred small defences like it are most of what separates a demonstration from an index.

The law behind the case, not just the case

A case record tells you that a court did something. It does not tell you what the law it acted under says, whether the offence is bailable, which court can try it, or what that section is called since the criminal codes were replaced.

So we built the statute layer as a first-class part of the engine rather than a reference page bolted beside it. Read from the live corpus on 13 September 2026, it holds:

WhatCount
Acts, Central and State10,084 (909 Central, 9,175 State)
Sections of those Acts2,69,602
State Acts that exist as searchable text only because the scanned gazette was machine read7,765 Acts, 1,85,941 sections
Subordinate instruments: rules, notifications, circulars, orders35,622
Rules of the CPC First Schedule728, across 58 Orders
Offence classifications from the BNSS First Schedule464
Official transition mappings across four statutory pairs2,372
Cross references between provisions1,72,143
Recorded repeals4,513
Reported judgments attached to the sections they decide3,953 (4,900+ on the 22 September 2026 corpus)
IndiaCode corpus counts read from the live meta endpoint, corpus date 13 September 2026. Free, keyless and open to any caller.
The law behind the case THE STATUTE LAYER · CORPUS DATE 13 SEPTEMBER 2026 · FREE AND KEYLESS The law behind the case, not just the case 10,084 ACTS · 909 CENTRAL, 9,175 STATE 2,69,602 SECTIONS OF THOSE ACTS 35,622 RULES, NOTIFICATIONS, CIRCULARS, ORDERS 728 CPC FIRST SCHEDULE RULES, ACROSS 58 ORDERS 464 BNSS OFFENCE CLASSIFICATIONS 1,72,143 CROSS REFERENCES BETWEEN PROVISIONS WORKED MAPPING, READ 14 SEPTEMBER 2026 IPC 302 Punishment for murder BNS 103 Punishment for murder MEASURED TEXT SIMILARITY 0.9 ALSO CONSIDERED, AND REJECTED BNS 105 culpable homicide, not murder 0.8063 BNS 109 attempt to murder 0.7312 The near misses are published, because a mapping is a research aid and not an authority. THE PART THAT EXISTS NOWHERE ELSE 7,765 State Acts are searchable text only because we read the scan. Machine read into 1,85,941 sections by their own numbering. They do not exist as text anywhere else, including at the source. Every one is marked as read from an image, with the official scan mirrored beside it. eCourtsIndia · IndiaCode corpus · no key, no registration, no paywall
A case record says a court did something. The statute layer says what it did it under, what that provision is numbered today, and whether the offence is bailable.

The transition tables are the part with immediate operational value. India replaced the Indian Penal Code, the Code of Criminal Procedure and the Evidence Act in 2024. An FIR from 2023 cites IPC 302. The same conduct today is charged under BNS 103. A string search for one will never find the other, so any system that searches on section numbers without mapping them is silently blind to half the corpus on either side of the changeover.

Asked to translate IPC 302 on 14 September 2026, the engine returns BNS 103, punishment for murder, at a measured text similarity of 0.9, and it also names the two provisions it considered and rejected: BNS 105 at 0.8063 and BNS 109 at 0.7312. The alternatives are published because a mapping is a research aid and not an authority, and a reader who cannot see the near misses cannot judge the match. The full walk through is in the IPC to BNS mapping guide, and the same corpus is published for machines at four URLs per provision, described in the IndiaCode search and JSON API guide.

Put the two halves together and the engine can answer a question no single dataset can. Not just find me the case, but tell me what the section it was filed under actually says, whether that offence is cognizable and bailable, which court could try it, what the section is numbered today, and which reported judgments interpret it.

Readers who want the statute side without any code can use it directly at IndiaCode by eCourtsIndia, and the story of how that corpus was assembled, section by section with the judgments that read them, is told in IndiaCode by eCourtsIndia: every Act, every section.

The hardest problem is not finding. It is attributing.

Everything above is infrastructure. This is where the infrastructure gets tested, because the moment anyone asks a real question of 32 crore+ records, the question is never about a case. It is about a person.

A single litigant query for Kumar, run on 14 September 2026, returned 2,52,71,909 records. That is roughly one in every eleven records in the entire national index. Narrow it to Rajesh Kumar and you still have 6,91,343. No ranking trick makes a name that common identifying, and any product that pretends otherwise is shipping a coin toss with a confidence score printed on it.

Now the mirror-image failure. Search the index for Mohammed Ali and you get 53,062 records. Search for Mohammad Ali and you get 86,785. Same name, two transliterations of it, and a searcher who picks one spelling never sees the other. The Indian court register is typed by hand, in dozens of scripts, then transliterated by people who were not consulting each other.

Those two numbers describe the whole problem. One name can scatter a person across many spellings, and one spelling can gather many people under one name. A serious background check has to solve both at once, and the two fixes pull in opposite directions. Loosen matching to catch the spellings and you drown in namesakes. Tighten it to kill the namesakes and you lose the person.

A name is not an identity IDENTITY RESOLUTION · COUNTED LIVE 14 SEPTEMBER 2026 A name is not an identity LITIGANT QUERY · Kumar 2,52,71,909 candidate records ≈ 1 in 11 of the whole index ONE NAME SCATTERS Mohammed Ali 53,062 Mohammad Ali 86,785 ONE SPELLING GATHERS Rajesh Kumar 6,91,343 thousands of unrelated people FOUR KINDS OF EVIDENCE, WEIGHED SEPARATELY NAME How many people could answer to this string variants, transliteration, OCR error, initials, stems FAMILY A relative’s name, where the record carries one patronymic by stem, household, second register PLACE Registered where you are from, sued where you live districts compared by letters, never by equality PROVENANCE Where a finding came from, and what it cannot prove no self-corroboration, one dispute counts once Candidates are not ranked. They are tested, and the test is built to refuse rather than to flatter. THE CONSTRAINT WALL · NAMED, VERSIONED, PERMANENT Every constraint can only lower a match. Not one can raise it. CONFIRMED evidence chain shown PROBABLE open question kept open POSSIBLE shown, never counted EXCLUDED with the reason named COVERAGE what we did not check A 90 per cent similarity score is not a person. It is a reason to look harder. eCourtsIndia · all counts read live from the index on 14 September 2026
Search returns candidates. Attribution is a separate discipline, and it sits above search rather than inside it.

Why a fuzzy matcher on its own is not enough

Generic fuzzy matching answers the question how similar are these two strings. That is a useful question and it is not the question a background check is asking. The question a background check is asking is how many other people could this string be.

Those come apart badly in Indian data. A ninety per cent string similarity on a name carried by two and a half crore records is worth almost nothing. A seventy per cent similarity on a rare patronymic in a small district can be close to decisive. Similarity is a property of two strings. Identifying power is a property of a population. A system that scores the first and never measures the second will confidently hand you a stranger.

There are further structural traps that no similarity function sees. A party field is not always a name: a motor accident claim can carry a vehicle registration, an insurer and four people in one string, and a large cause title can list several hundred litigants, at which point the title identifies nobody. The same place is spelled differently by different registers, and district boundaries have been redrawn more than once inside the period the records cover. People are registered where they are from and litigate where they live, so an address is a claim about a person rather than a fence around them. And Indian court metadata carries no Aadhaar and no PAN, so identity has to be built out of corroboration rather than read off a field. Each of those is a separate piece of engineering. None of them is solved by a better string comparison.

The engineering answer we shipped is a wall of named constraints, more than twenty of them, and every one can only ever lower a match. Not one can raise a score, promote a band, or turn a maybe into a finding. That single property is the whole architecture of trust in one sentence, and it is the reason independent evaluation records no false confirmations rather than a flattering accuracy percentage. Between them the constraints cover six territories: how a name is read, how family evidence is weighed, how place is handled, where a piece of evidence came from, how a company differs from a person, and how a second register may and may not be used. The largest group by some distance is the first, because reading an Indian name correctly turns out to be most of the problem.

They are not a design document either. Not one of them was written in advance. Each exists because a specific real person was once scored wrong, an independent evaluator caught it, and the failure was given a permanent name so it could never come back quietly under a different description. That is also why they are versioned and never renamed: a constraint that changed meaning between releases would turn every report issued before the change into a lie. What each one tests, and the threshold it tests against, stay inside the engine. The catalogue itself is published inside the product, where somebody reading a report about a real person can see which constraints were applied to their subject and which refused.

None of that wall is runnable without the engine underneath it. Every constraint is, in the end, a query, and most of them are queries a flattened court dataset cannot answer at all. They need an index you can aggregate across rather than look up one record at a time, party roles held as separate fields rather than merged into a party string, the relationships between matters carried rather than discarded, and provenance attached to every finding so the system can say where a piece of evidence came from. That is the real reason the engine had to come first. The rules are the easy part to describe and the expensive part to run.

The volume is the part that does not photograph well. A single subject passes through several hundred discrete decisions before a report exists. On a difficult common name, one check can score three hundred or more candidate records individually and run dozens of separate register searches, and most of that work ends in a refusal the reader never sees as a finding. A check that comes back in two seconds has not done this. It has run a search. And the whole of it is marked by an adversary: quality is scored by a separate system that is handed the raw court and electoral data, is never shown the engine’s reasoning, and re-derives the truth from source on every cycle. Its only job is to catch us calling a stranger our client’s subject. It has told us we were wrong about our own release notes more than once, and it was right.

Built to be read by machines, not only by people

There is a version of this story where the pay-off is a nicer search page. That is not the pay-off.

A language model asked about an Indian case has three options. It can recall something, which for case-specific facts means inventing them. It can read a PDF someone pasted, which gives it one document and no context. Or it can call a tool that returns structured, current, sourced data and reason over that.

Only the third option produces work a lawyer can file. So the engine is published as a Model Context Protocol server with 39 tools (server version 4.46), of which 19 never consume credits. They cover four datasets that are deliberately kept apart, because the commonest mistake an AI agent makes with Indian legal data is answering a statute question out of the case corpus or the reverse.

  • Cases and orders. What a court actually did. Search, brief, full order text, structured AI analysis of an order, and live refresh of a single case or a batch.
  • Cause lists. What is listed tomorrow, across more than 8,000 district and taluka courts, rather than only what was filed last year.
  • The statute book. What the law says. Free, keyless, with provision lookup, offence classification, CPC rules, subordinate legislation and the transition mappings.
  • The electoral roll. A second, independent register, used as corroboration and never as identification.

Since version 4.46 the same server also runs LegalCheck. submit_legal_check screens a person or company against the case index as an asynchronous job, and get_legal_check, which is free, waits for it and returns the risk band, the summary and every matched case. That puts the attribution wall described above behind a single tool call, for a compliance team that would rather ask Claude than write code. The LegalCheck API post covers the REST route for the same job.

One more layer sits beside the courts rather than inside them. Crime Reports indexes 12 lakh+ FIR PDFs from 13 states and union territories, which matters for attribution because a criminal matter often begins as an FIR years before a CNR exists. Each FIR PDF costs Rs 1.

Four datasets, kept deliberately apart THE MACHINE SURFACE · 39 TOOLS, 19 OF THEM FREE Four datasets, kept deliberately apart CASES AND ORDERS What a court actually did. ANSWERS Has this company been sued in Pune since 2019? Never the place to ask what the law says. CAUSE LISTS What is listed tomorrow. ANSWERS Is my case on any board on Monday morning? 8,000+ district and taluka courts, every day. THE STATUTE BOOK What the law itself says. ANSWERS Is BNS 85 bailable, and which court tries it? A provision with no judgment listed under it does not mean no case exists. THE ELECTORAL ROLL A second, independent register. ANSWERS Is a relative of that name at that address? Corroboration, never identification. Partial coverage, disclosed in every report. WHY THEY ARE KEPT APART The commonest mistake an agent makes is answering a statute question out of the case corpus. eCourtsIndia · read from the live server manifest on 14 September 2026
A model that cannot tell the law from the case will state a confident falsehood. The separation is enforced before the model gets a chance to guess.

That separation is itself an engineering decision with a reason. A provision with no reported judgment listed against it does not mean no case exists, and an agent that conflates the two will state a confident falsehood. The server tells the model so, in its own instructions, before the model has a chance to guess. The REST side is documented here, and MCP 101 for legal teams explains how Claude or ChatGPT connects to the server.

This is the bedrock argument we made when we wrote about the data moat in the age of commodity models. Model capability is converging and inference is getting cheaper every quarter. What does not converge is coverage, freshness, structure quality, entity resolution and historical depth. An AI product for Indian law is a thin wrapper over its data layer, and the data layer is the part nobody can copy in a sprint.

Copy the architecture. You still have to run it.

Everything in this post is deliberately readable. The scale, the field list, the corpus counts, the coverage and the limits are all public, because a buyer should be able to check a vendor’s claims rather than take them. What sits inside the engine stays inside the engine: the arithmetic of the scoring, what each constraint tests and the threshold it tests against, the order in which evidence is gathered and weighed, and the query shapes used to find a candidate in the first place. None of that is published here, and none of it will be. We describe the shape of the machine. We do not hand over the drawings.

Illustration of the eCourtsIndia engine architecture that sits behind search, alerts, the API and LegalCheck

We publish the rest because publishing it does not give the architecture away in any useful sense. The logic is only executable against data that has to be operated, daily, at national scale, for years. Assume it is reimplemented perfectly tomorrow. It still needs a live index of 32 crore+ court records rather than a snapshot, the order text underneath those records rather than their status fields, a statute corpus of 2,69,602 sections mapped across a criminal-law transition, a second national register to corroborate identity against, and enough adversarial evaluation cycles to know which of its own confident answers are wrong.

Nobody starts there. We did not either. It took a build measured in years, and the count moved 578 records while this sentence was being drafted.

What we will not claim

  • The index is not complete. If a matter is missing, anyone can add a missing case by CNR. The index is large, live and growing, and it is assembled from public records that are themselves uneven. Coverage gaps are disclosed in product rather than presented as absence of a case.
  • A count is a reading, not a constant. Every figure here carries the date it was read, because a figure without a timestamp on a live system is a claim rather than a measurement.
  • Order text is not always available. Some court portals do not serve a PDF on request, and some tribunal orders have no downloadable path at all. Where the text is missing we say so instead of inferring an outcome from the label.
  • Structured is not the same as true. A field extracted from a register inherits whatever the register recorded, including its typographical errors. Structure makes a record queryable. It does not make it correct.
  • A case is not a conclusion. Being named in a matter establishes nothing, and a person is presumed innocent unless convicted. Everything the engine produces about a person is designed to be checked by a human before anyone acts on it.

What this means for eCourtsIndia

Court records in India are public. Turning them into a queryable, current, statute-aware, machine-readable index is the entire job, and it is the part that takes years rather than quarters. Everything we ship on top, search, alerts, the API, LegalCheck and the AI layer, is a consequence of the engine rather than a substitute for it.

If you want to see it working rather than read about it, search the index free, take an API key, or run a LegalCheck and read the coverage statement at the bottom of the report as carefully as the findings at the top.

Frequently Asked Questions

How many court records does eCourtsIndia hold?

The index passed 32 crore+ case records in September 2026, covering the Supreme Court, all 25 High Courts, district and taluka courts in all 36 states and union territories, and 18 tribunal and commission types. When this article was first written on 14 September 2026 it read 29,02,62,233, and it rose by 578 records between two calls a minute apart. You can search the index free.

What does structured court data actually mean?

It means every matter exposes named fields that a caller can request, filter and sort on, rather than a page of text to be scraped. On this index that is 38 projectable fields per matter, 31 sortable dimensions, 11 facets, 23 presence-filterable fields and four name matching modes, with five party roles kept as separate fields instead of merged into one string. The API guide lists them.

Why does a case marked disposed tell you nothing?

Because on 14 September 2026, 21,92,90,294 of 29,02,62,233 records carried that one status, which is 75.5 per cent of the corpus. Disposed can mean an acquittal, a conviction, a compromise, a restorable default dismissal, a committal to a higher court, or a case stopped because the accused absconded. Only the order text tells them apart, so the engine converts order PDFs to text and reads the outcome.

Why is name matching so hard in Indian court records?

Two problems pull in opposite directions. A single litigant query for Kumar returned 2,52,71,909 records, so a common name has almost no identifying power. At the same time one person scatters across spellings: Mohammed Ali returned 53,062 records and Mohammad Ali 86,785. Loosening matching to catch spellings floods you with namesakes, and tightening it to remove namesakes loses the person you were looking for.

Is a 90 per cent name match good enough to act on?

No. String similarity measures how alike two strings are. It does not measure how many other people could carry that string. A ninety per cent match on a name shared by crores of records is weak evidence even when it looks perfect, while a lower score on a rare patronymic in a small district can be far stronger. Identifying power belongs to a population, not to a pair of strings.

How does the statute layer connect to the case layer?

The statute corpus holds 10,084 Acts, 2,69,602 sections, 464 BNSS First Schedule offence classifications and 2,372 official transition mappings. That lets a case filed under IPC 302 be read beside the same offence charged today as BNS 103, with punishment, cognizability, bailability and trying court stated in the words of the Schedule. It is free and keyless on IndiaCode by eCourtsIndia.

Can AI agents query this data directly?

Yes. The engine is served over the Model Context Protocol at mcp.ecourtsindia.com with 39 tools, 19 of which never consume credits, covering cases and orders, cause lists, the statute book, the electoral roll and LegalCheck background checks. Structured tool access lets a model answer from the record instead of from memory. MCP 101 for legal teams explains the setup.

How is this different from scraping the public portals?

A scrape is a snapshot frozen the moment it is taken, which is why snapshot data so often shows a next hearing date that has already passed. This index is operated rather than harvested: a record can be re-pulled from its own court on demand, individually or in bulk, and the order text and statutory context travel with it instead of sitting in a PDF nobody reads.

Sources and method

  • Total index size, court level split and status distribution were read live from the eCourtsIndia case index on 14 September 2026: 29,02,62,233 total, then 29,02,62,811 on a repeat call about a minute later. Court levels: district and taluka 23,14,66,178, High Court 5,16,69,906, tribunals 42,97,822, Supreme Court 11,85,462. Status: DISPOSED 21,92,90,294, PENDING 6,98,16,110, UNKNOWN 6,50,894, across 30 distinct status values.
  • Capability figures, 38 projectable fields, 31 sortable fields, 11 facetable fields, 23 presence-filterable fields, four name matching modes and a 200-row page ceiling, were read from the live search capability catalogue on the same date.
  • Court coverage, 21 live court types of which 18 are tribunal forums, was read from the live court type enum on the same date.
  • Name counts were read live on the same date: litigant Kumar 2,52,71,909, Rajesh Kumar 6,91,343, Mohammed Ali 53,062, Mohammad Ali 86,785. Text counts within DISPOSED records: acquitted 4,40,631, convicted 3,33,516. Those two are full-text counts across indexed order text and AI summaries. They illustrate the ambiguity of the status field and are not a census of outcomes.
  • The worked example is CNR SCIN010108932008, Criminal Appeal 1538 of 2013 in the Supreme Court of India, arising from Criminal Appeal 1429 of 2002 of the High Court of Judicature at Madras. Status DISPOSED, 15 orders, five interlocutory applications. The facts stated are taken from the text of the order dated 23 September 2013 as held in the index. Parties are not named here because the point is the data structure and not the individuals.
  • Statute corpus counts, 10,084 Acts, 2,69,602 sections, 35,622 instruments, 728 CPC First Schedule Rules across 58 Orders, 464 offence classifications, 2,372 mappings, 1,72,143 cross references, 4,513 recorded repeals and 3,953 reported judgments (the 22 September 2026 corpus holds 4,900+ reported judgments), were read from the IndiaCode meta endpoint, corpus date 13 September 2026. The 7,765 State Acts held by the source only as a scanned gazette, machine read here into 1,85,941 sections, are stated in the published corpus description read on 14 September 2026. Their provenance is marked on every page and in the API, and the Government’s own scan is mirrored and linked, because text read off an image is not the same thing as text.
  • The IPC 302 to BNS 103 mapping, measured similarity 0.9, with BNS 105 at 0.8063 and BNS 109 at 0.7312 named as alternatives, was read on 14 September 2026. Similarity here is measured by text comparison and is a research aid, not an official equivalence.
  • The machine surface, 37 MCP tools of which 18 consume no credits, was read from the live server manifest on the same date.
  • The constraint architecture, the decision volume per check, the independent evaluation method and the banded report structure are described in the LegalCheck launch post, which is the canonical page for those claims.
  • What is deliberately not published, here or anywhere else on this blog: the arithmetic inside the scoring, what each constraint tests and the threshold it tests against, the order in which evidence is gathered and weighed, the query shapes used in discovery, and internal engine build or release numbering. The architecture, the coverage and the limits are public. The implementation is not.
  • Figures on this page are a dated snapshot of a live system and not a service level commitment. Where a source states a limitation, this page preserves it rather than converting it into a stronger claim.
  • Tool counts (39 tools, 19 that never consume credits, LegalCheck on MCP) were read from the server health endpoint and llms.txt at mcp.ecourtsindia.com on 23 September 2026. The 32 crore+ figure is the live case count read on 23 September 2026.

Read next

eCourtsIndia is a private legal-technology platform. It is not affiliated with, associated with, or endorsed by the Government of India, the Supreme Court of India or its e-Committee, or any court. Official case information is published on ecourts.gov.in. Always verify details against official court records or certified copies. This article is general information, not legal advice. Spotted an error? Write to support@ecourtsindia.com.

Search 32 crore+ Indian court case records, free

One search across the Supreme Court, all 25 High Courts, district courts and 18 tribunal and commission types. Hearing alerts, AI summaries and an API for developers.