,

Mapping India’s Court Data Stack: From NJDG to APIs to AI Agents

India’s court data stack mapped: from NJDG and eCourts to aggregation platforms, research tools and AI agents, and which layers still have whitespace.

·

·

eCourtsIndia Knowledgebase

Mapping India's court data stack from NJDG to APIs to AI agents, cover design variant A for the eCourtsIndia blog

Most Indian legaltech market maps in circulation skip the layer that every other layer depends on. They show you contract lifecycle management, online dispute resolution, AI copilots, litigation finance, RegTech, legal marketplaces. They do not show you where the data comes from.

Last updated: 23 September 2026

Mapping India's court data stack from NJDG to APIs to AI agents, portrait cover image for the eCourtsIndia blog

Court data is the substrate for most of Indian legaltech. An AI drafting tool needs judgment text. A litigation intelligence product needs case status. A background verification check needs litigation history. A case management app needs hearing dates. None of it works without a reliable court data pipeline underneath.

This post maps that stack end to end. Five layers, from the public source at the bottom to the AI agent at the top, with the companies operating at each layer, what the infrastructure layer actually has to do, and why the market has historically under-built it.

Mapping India's court data stack from NJDG to APIs to AI agents, cover design variant B for the eCourtsIndia blog

Key takeaways

  • India’s court data stack has five layers: the public source, aggregation and structuring, research databases, AI agents, and applications.
  • Layer 1 (eCourts, NJDG, High Court, Supreme Court and tribunal portals) is authoritative and free, but built for single-case lookup.
  • Layer 2, aggregation, is where scattered records become a dataset. It is hard, slow to build and the layer most market maps leave out.
  • Layers 3 to 5 inherit the strengths and gaps of Layer 2. District courts and Indian languages are the two biggest gaps.
  • eCourtsIndia sits at Layer 2, with 32 crore+ case records, an API, an MCP server, and products such as LegalCheck built on its own data.

Layer 1: The public source

India’s court data stack has five layers: a public government source at the bottom, an aggregation layer that structures it, legal research databases, AI agents, and applications at the top. Layer 1, the public source, is where every court record originates, and every layer above it depends on this foundation being reliable, complete, and current.

  • eCourts Services (services.ecourts.gov.in). District and subordinate court records across 36 states and union territories.
  • National Judicial Data Grid (njdg.ecourts.gov.in). Live pendency, disposal and institution statistics. It launched for district courts in September 2015 and for High Courts on 3 July 2020. We cover its history in our primer on how India digitised 18,000+ courts under eCourts.
  • High Court portals. 25 separate High Court websites, each with its own interface, case information system and cause-list format.
  • Supreme Court of India. Registry, cause lists, judgments and case status on its own site.
  • Tribunals and commissions. NCLT, NCLAT, ITAT, CESTAT, CAT, DRT and DRAT, NGT, AFT, SAT, TDSAT, APTEL, consumer commissions and more, each on its own portal with its own numbering.

This layer is comprehensive, authoritative and free. It is also, by design, built for single-case lookup rather than bulk or programmatic access. Captchas, fragmented interfaces and the absence of a third-party API are deliberate. They protect the public source from scraping abuse. The government is still investing heavily here: Phase III of eCourts carries a Rs 7,210 crore outlay, which we break down in what Rs 7,210 crore will build.

Layer 2: Aggregation and structuring

This is the layer that turns scattered public records into a queryable dataset, the document-to-data journey that moves Indian court records from PDFs to APIs. It is also the layer most investors miss.

What “legal data infrastructure” actually means

Legal data infrastructure is the layer that converts raw, scattered court records into clean, structured data that any application can query. It takes case records, orders, cause lists, party names, advocate details and judge assignments, and turns them into fields a machine can filter, join and count.

That is harder than it sounds. The data arrives in different formats and different languages from thousands of court complexes, 25 High Courts, the Supreme Court and the tribunals. A single litigant might appear as “M/s XYZ Pvt Ltd” in one court and “XYZ Private Limited” in another. The same advocate’s name can be spelled three ways in three districts. An operator has to ingest from all of it, normalise it, resolve entities, detect updates and keep the whole pipeline fresh every day. The work is operational and engineering-heavy, it does not produce shiny demos, and it takes years to get right. We describe the fields we extract and why a “disposed” status alone tells you little in Building India’s legal data engine.

Who operates here

  • eCourtsIndia. 32 crore+ case records and 125 crore+ orders and judgments across the Supreme Court, all 25 High Courts, district and taluka courts in all 36 states and union territories, and 18 tribunal and commission types. A directory of 34 lakh+ advocates and 82,000+ judges. Developers get a REST API with 23 endpoints and Rs 200 of free credits on signup, and a hosted MCP server at mcp.ecourtsindia.com with 39 tools covering cases, cause lists, a statute book of 10,084 Acts and the electoral roll. Our developer quickstart for the eCourtsIndia API is the fastest way in.
  • IndianKanoon. Free search over a large judgment corpus, bootstrapped, and a much-loved utility for Indian lawyers for well over a decade. It is a search tool rather than a structured data platform.
  • LegitQuest. Research plus background verification, with partial district court coverage.
Court tierCase records indexedShare of corpus
District and taluka courts26.0 croreAbout 81%
High Courts5.43 croreAbout 17%
Tribunals and commissions44 lakh+About 1.4%
Supreme Court11.9 lakhUnder 0.5%
Total32 crore+100%
Where Indian court records sit, by tier. eCourtsIndia index, read live on 23 September 2026. District courts hold about four in five records, which is why Layer 2 coverage of that tier matters most.

Layer 2 is what makes everything above it possible. Every research platform, every AI agent, every due-diligence product is only as good as the aggregation layer it depends on.

Layer 3: Legal research databases

This is the oldest commercial layer in Indian legaltech. Built for Supreme Court and High Court case research, typically sold to urban firms and chambers on annual per-seat subscriptions.

  • Manupatra. Curated Indian case law, statutes and commentary. Legacy incumbent.
  • SCC Online. Supreme Court Cases Online. The other legacy incumbent.
  • Casemine. Research with AI-assisted discovery.
  • Smaller research databases such as Nearlaw and LawAtYourFingertips.

Strong at what they were built for. Largely absent from the district-court tier, which is where most Indian advocates practise, as we argue in The district court lawyer is ninety percent of Indian law. For a feature-by-feature view, see our comparison of eCourtsIndia, IndianKanoon and SCC Online.

Layer 4: AI agents and copilots

This is the most visible layer in 2026. Funded, press-covered and full of experimentation.

  • Jhana. AI paralegal and research copilot. Raised a $1.6 million seed round led by Together Fund, announced in September 2024.
  • Nyayanidhi. A litigation operating system for drafting, translation and filings. Raised a $2 million seed round led by 3one4 Capital in November 2025.
  • Lucio. An AI-native workspace for lawyers. Raised a $5 million seed round led by DeVC in October 2025.
  • SpotDraft, SimpliContract. AI-assisted contract lifecycle management (contract data rather than court data).

The honest observation about Layer 4 is that it depends entirely on Layers 2 and 3. Without structured court data beneath it, an agent has nothing to reason over. Several Layer 4 products rely on a mix of IndianKanoon, Manupatra, scraped public sources and custom pipelines, which is why the best of them are moving toward structured-data partnerships. The MCP protocol makes that partnership almost trivial: an agent that speaks MCP can query our server directly. We look at what Harvey, Hebbia and the Indian copilots tell us about this in The AI agent layer for Indian law.

Layer 5: Applications and workflows

The top of the stack. Products that consume the lower layers and package them for a specific workflow.

  • Case management and practice tools. Tools that help advocates manage matters, hearings and clients. On eCourtsIndia, AI Clerk does this with WhatsApp and email alerts and AI order summaries, on a free tier of 50 credits a month and paid plans from Rs 250 + GST a month.
  • Background verification. BGV vendors running court-record checks for hiring, lending and onboarding. LegalCheck, our identity-first check, is one; see the LegalCheck launch post. Partners can call it through the LegalCheck API at Rs 99 per check pay-as-you-go, or Rs 33 with a subscription.
  • Online dispute resolution. Presolv360, Jupitice, Sama and others. Arbitration and mediation platforms.
  • Litigation finance. LegalPay, FightRight. Capital for claimants.
  • Compliance and RegTech. IDfy, RegisterKaro and others. Compliance automation for companies.
  • Legal marketplaces. Vakilsearch, LegalKart, Lawyered. Consumer and SME legal services.

The stack, at a glance

LayerWhat it doesRepresentative players
5. ApplicationsWorkflow products: case management, ODR, litigation finance, BGV, marketplacesPresolv360, LegalPay, IDfy, Vakilsearch, SpotDraft, LegalCheck
4. AI agentsResearch copilots, drafting assistants, AI workflowsJhana, Nyayanidhi, Lucio, Casemine
3. Research DBsCurated case law for SC and HC researchManupatra, SCC Online, Casemine
2. AggregationStructuring public court data into queryable datasets and APIseCourtsIndia, IndianKanoon, LegitQuest
1. Public sourceCanonical government dataeCourts, NJDG, SC, HC portals, tribunals
The five layers of the Indian court data stack, with representative players at each layer.

Why the infrastructure layer is the valuable position

In most technology markets, the infrastructure layer produces some of the largest outcomes. Three reasons explain why.

It earns from everyone building above it. A legaltech startup that needs court data can build its own pipeline, which takes a long time and a team that does nothing else, or it can call an API. Each product that builds on the data layer is a customer that does not compete for the same end user.

Switching costs are structural. Once a company wires a data API into its product, changing providers means rewriting core logic, re-testing entity matching and re-validating coverage. This is not like exporting a CSV from one SaaS tool into another.

The moat compounds. Every day of ingestion, cleaning and entity resolution adds to the gap a new entrant has to close. Usage adds more: error reports, missing-case submissions and corrections make the dataset better the more it is used. We go deeper on this in The data moat in the age of commodity LLMs.

The global comparables

The largest legal information businesses in the world are, at their core, data infrastructure with products on top. Thomson Reuters reported Legal Professionals revenue of about $2.77 billion for 2025. RELX reported revenue of about £1.8 billion for LexisNexis Legal & Professional in the same year. Outside law, Nasdaq agreed to buy the anti-financial-crime data platform Verafin for $2.75 billion in November 2020, closing the deal in early 2021. None of these numbers says what an Indian company will be worth. They do show that owning the reference data layer in a regulated domain is a durable business.

What gets built on top

  • Legal research tools that need case records, orders and judgments across all courts, not only the Supreme Court and High Courts.
  • Litigation management platforms that need hearing dates, status changes and new orders the day they appear.
  • Due diligence and compliance tools that check a person or company against court records, for banks, insurers, employers and investors.
  • AI legal assistants that need to ground answers in real case records instead of a general-purpose model’s memory.
  • RegTech platforms that watch for enforcement and litigation risk across suppliers and portfolios.

All of them need the same underlying data. The operator that supplies it reliably, at scale, captures a slice of each category.

Where the Indian market is underbuilt

Looking at the stack honestly, three patterns stand out.

Layer 2 is the bottleneck. Most visible Indian legaltech funding has gone to contract software, AI copilots and applications. Very little has been directed at the aggregation layer, despite it being the choke point. We look at the funding record in our analysis of Indian legaltech funding. This is the opposite of how the US stack evolved, where PACER access and projects like CourtListener built a mature data layer before the application wave.

District courts are under-represented at every layer above 1. Research databases cover the Supreme Court and High Courts well and district courts barely at all, even though district and taluka courts hold about four in five of the case records we index. AI agents inherit that bias.

Vernacular is a green field. The stack, from Layer 2 upward, is built mostly in English. India’s trial courts produce orders and cause lists in many scheduled languages. An AI agent that reads Hindi, Tamil, Marathi and Bengali court documents does not yet exist at scale.

What this means for investors and operators

If you are investing in Layer 4 or Layer 5, the due-diligence question that matters most is where the data comes from. An AI drafting product with no structured court data behind it is a thin wrapper. A case management tool without hearing alerts is a contact book. The moat in Indian legaltech, for the next several years at least, sits in Layer 2.

If you are operating at Layer 4 or 5, the decision you will face is whether to build your own ingestion pipeline or partner with a Layer 2 provider. The companies that decide early will look very different in 2028 from the companies that do not.

Mapping India's court data stack from NJDG to APIs to AI agents, cover design variant C for the eCourtsIndia blog

What this means for eCourtsIndia

We sit at Layer 2, deliberately. Our job is to turn the public source into the queryable, API-accessible dataset that every layer above us depends on. That is why we built the REST API and the MCP server, why we index tribunals as well as courts, and why we mint CNRs for 14 tribunals that do not issue their own. It is also why our own products, LegalCheck, AI Clerk and Crime Reports with 12 lakh+ FIR PDFs, sit on exactly the same data we sell to developers.

Over time, we expect the Indian court data stack to look less like a pyramid and more like a pipeline: the public source feeding the aggregation layer, the aggregation layer powering research, AI agents and applications, and every layer getting smarter as the ones below it get richer. For the longer version of that argument, read The operating system for Indian law.


Building at Layer 4 or Layer 5 and need structured court data? Start with the eCourtsIndia API or connect an agent to mcp.ecourtsindia.com/mcp. Search any case free at ecourtsindia.com/search.

Sources and further reading

  • Department of Justice: eCourts Mission Mode Project overview (PIB, September 2023); eCommittee, Supreme Court of India: launch of NJDG for High Courts, 3 July 2020.
  • Entrackr and Analytics India Magazine, September 2024: Jhana seed round. Inc42 and Entrackr, October 2025: Lucio seed round. Entrepreneur India, November 2025: Nyayanidhi seed round.
  • Thomson Reuters: fourth-quarter and full-year 2025 results. RELX: Annual Report 2025, Legal segment.
  • Nasdaq: agreement to acquire Verafin, November 2020.
  • Harshith Viswanath, The LegalTech Thesis: Mapping India’s LegalTech Ecosystem, March 2026.
  • eCourtsIndia index counts, read live on 23 September 2026.

Frequently Asked Questions

What are the five layers of India’s court data stack?

Mapping India's court data stack from NJDG to APIs to AI agents, square social cover for the eCourtsIndia blog

India’s court data stack runs from the public government source at the bottom, through an aggregation layer that structures records, up to legal research databases, AI agents and applications at the top. Each layer depends on the one beneath it, so gaps in the data layer show up everywhere above it. Explore the structured data layer at eCourtsIndia search.

What is legal data infrastructure?

Legal data infrastructure is the layer that turns raw court records, orders, cause lists, party names, advocates and judges into clean, structured data that any application can query. It means ingesting from thousands of courts, resolving entities spelled many ways, and refreshing daily. Developers can use it through the eCourtsIndia API, which has 23 endpoints and Rs 200 of free credits.

Which layer of Indian legaltech is most underbuilt?

The aggregation layer, Layer 2, is the bottleneck. Most visible legaltech funding in India has gone to contract software, AI copilots and applications, while the layer that structures raw court data into queryable datasets saw little directed capital. Read the funding record in our analysis of Indian legaltech funding.

What is the National Judicial Data Grid (NJDG)?

The National Judicial Data Grid (njdg.ecourts.gov.in) is part of the public source layer. It publishes pendency, disposal and case institution statistics across court tiers. It launched for district courts in September 2015 and for High Courts on 3 July 2020. You can look up case status, orders and parties on eCourtsIndia.

Why are district court records under-represented in Indian legaltech?

Legal research databases were built for Supreme Court and High Court research and barely cover the district tier, even though district and taluka courts hold about four in five Indian case records. AI agents trained on those databases inherit the gap. Read more in The district court lawyer is ninety percent of Indian law.

Where does eCourtsIndia sit in the court data stack?

Mapping India's court data stack from NJDG to APIs to AI agents, X share card for the eCourtsIndia blog

eCourtsIndia operates at Layer 2, the aggregation layer. It structures 32 crore+ case records from the Supreme Court, all 25 High Courts, district courts in all 36 states and union territories, and 18 tribunal types, and serves them through a REST API and an MCP server. Products such as LegalCheck run on the same data.

eCourtsIndia is a private legal-technology platform. It is not affiliated with, associated with, or endorsed by the Government of India, the Supreme Court of India or its e-Committee, or any court. Official case information is published on ecourts.gov.in. Always verify details against official court records or certified copies. This article is general information, not legal advice. Spotted an error? Write to support@ecourtsindia.com.

Search 32 crore+ Indian court case records, free

One search across the Supreme Court, all 25 High Courts, district courts and 18 tribunal and commission types. Hearing alerts, AI summaries and an API for developers.