,

The Data Moat in the Age of Commodity LLMs: Why Court Data Wins Indian Legal AI

Data moat for Indian legal AI: as LLMs commoditise, durable value sits in court data coverage, freshness, structure, entity resolution and history.

·

·

eCourtsIndia Knowledgebase

The data moat in the age of commodity LLMs, cover design variant A for the eCourtsIndia blog

Last updated: 23 September 2026. Corpus and directory figures refreshed, Harvey and Legora updated to their 2026 rounds, and two earlier posts folded in: our framework of five real and three fake legal data moats, and our piece on why legal software builds Salesforce-like switching costs.

The durable moat in Indian legal AI is not the model, it is the data. As frontier models converge and inference gets cheaper, the lasting advantage shifts to whoever owns the structured, refreshed, entity-resolved court data that every model has to call. The raw eCourts data is public, so the moat is not access. It is coverage, freshness, entity resolution, developer distribution and institutional trust, built together and kept running. In India that substrate had to be built from scratch, and it compounds with every use.

The data moat in the age of commodity LLMs, portrait cover image for the eCourtsIndia blog

The model layer is commoditising in slow motion, and most people building AI products are watching the wrong race. The story was supposed to be Claude versus GPT versus Gemini. The real story is that frontier capability is consolidating into a small number of foundation models that hit roughly the same ceiling on roughly the same tasks. The model wars matter a little less every quarter. What matters more is what those models can call out to.

In vertical AI, and especially in a regulated, jurisdiction-specific market like Indian law, that shift moves the durable moat toward two things: proprietary structured data the model has to call, and distribution that makes the model worth opening. This post is about the first. It sets out what a real data moat looks like in court data, which moats are fake, how switching costs build on top of them, and a rubric founders and investors can use to tell the difference.

Key takeaways

  • Frontier models are converging and routing layers treat them as interchangeable. A non-frontier company cannot build a lasting moat at the model layer.
  • Five moats are real in Indian legal data: coverage completeness, entity resolution, freshness, developer and AI distribution, and institutional trust. Three are fake: exclusive access to public data, celebrity endorsements and a pretty UI.
  • The real moats reinforce each other in a flywheel, and they compound with use. Every API call and MCP query sharpens the substrate. A model does not get better because more people use it.
  • On top of the data layer sits a second advantage: switching costs across workflow, data and habit. Done right, that is not lock-in. It is a tool that keeps paying the user back.
  • Global legal AI is scaling fast. Harvey was valued at $15.5 billion in September 2026 with more than $400 million in ARR. None of that removes the need for an Indian court-data substrate.

Why the model layer is commoditising

Three signals point the same way.

Frontier capability is converging. The gap between the top models on standard reasoning, coding and instruction-following benchmarks has narrowed sharply. New releases feel incremental rather than generational.

Inference is getting cheaper. Price per million tokens has fallen steeply across the major providers, and open-weight models are catching up on the tasks that matter for vertical applications: long-context reading, citation work and structured extraction. Harvey itself released a post-trained open-weight model, Tenet, in 2026, which says a lot about where the model layer is heading.

Switching between models is routine. Most serious AI applications now run on routing layers that pick Claude, GPT or Gemini per task on cost, latency or quality. The infrastructure treats models as interchangeable.

None of this means models have stopped improving. It means the differentiation per rupee is shrinking. For a product company sitting on top of the model layer, the moat has to live somewhere else.

What “data moat” actually means in this market

A data moat is the gap between what a platform can do because of its accumulated data and what a well-funded new entrant can do on day one. In consumer products the moat comes from user-generated content: reviews on a map, career histories on a professional network. In legal data it comes from years of normalised, cross-linked records of cases, judges, counsel and orders.

The Indian version has two unusual properties. The raw data is public: anyone can technically crawl the eCourts ecosystem. And the normalisation cost is huge, because schemas are inconsistent and formats change without notice. So the moat is not about access. It is about the engineering, quality checks and user feedback that make the data reliably usable. That does not show in a pitch deck. It shows the first time a lawyer tries an alternative and half the searches come back empty.

Break it into components:

Coverage. Every court that matters: the Supreme Court, all 25 High Courts, district and taluka courts in all 36 states and union territories, and 18 tribunal and commission types. On eCourtsIndia that layer now holds 32 crore+ case records, 125 crore+ orders and judgments, 79 crore+ litigant records, 34 lakh+ advocate profiles and 82,000+ judge profiles, linked and refreshed as new orders publish. A product that covers only the appellate tier is of little use to the large majority of the bar that argues in district courts.

Freshness. Daily refresh at minimum, and next-day cause lists that are current the evening before. Cause-list data goes stale within hours. Order uploads happen all day. The team with the pipeline to keep everything current owns this moat. The team that scrapes weekly does not.

Structure quality. HTML pages and PDFs are not data; they are signal that has to be turned into data. Structure means consistent field schemas across courts, parsed bench composition, parsed section citations and mapped case types. When one High Court calls an application CM(M) and another calls it MA, the user should see one concept. On eCourtsIndia that mapping runs to 269 case type codes and 71 status codes. Without it, the language model has to do the parsing itself, slowly and with errors.

Entity resolution. One advocate is one entity, not three. One company is one entity, not eight. Across courts, spellings, name changes and mergers. It is hard, never finished, and the difference between a useful API and a noisy one. A good recent example is how eCourtsIndia mints a 16-character CNR for 14 tribunals that never issued one, so a tribunal case can be tracked like any court case.

Vernacular OCR. Hindi, Marathi, Tamil, Telugu, Bengali, Kannada, Gujarati, Odia and Punjabi, in scanned PDFs going back years. A platform that reads these at production grade serves the whole bar. One that does not, serves the metros.

Historical depth and adjacent layers. A counterparty’s litigation history matters more than its latest case, and a judge’s pattern over five years matters more than today’s listing. Depth also means the layers around the case record: the statute book (10,084 Acts and 2,69,602 provisions on IndiaCode), and the MCP server’s 39 tools that let an AI agent move from a case to the section it cites without leaving the conversation.

A real moat has all six. We set out how they fit together in The Operating System for Indian Law, and mapped the wider stack in Mapping India’s Court Data Stack.

The data moat in the age of commodity LLMs, cover design variant B for the eCourtsIndia blog

Five real moats, and three fake ones

Components describe what a data layer is made of. Moats describe what actually stops a competitor. In Indian legal data, five hold up under pressure.

Real moatWhat it meansWhy it is hard to copy
Coverage completenessNational coverage, including taluka courts and tribunals, not a handful of statesExpensive to reach and more expensive to keep; every portal change breaks something
Entity resolution at scaleClean, de-duplicated links between parties, judges, advocates and casesGraph-level data work that is never finished
FreshnessDaily crawls, change detection and alerts the moment a tracked case movesEnterprise buyers care about this more than anything once they are in production
Developer and AI distributionAPI quality, documentation, MCP support and products built on topCompounds: each integration makes the next one easier
Institutional trustSecurity posture, audit trails, data handling and a record with procurement teamsSlow to earn, and often the deciding factor in a buying committee
The five real moats in Indian legal data. Adapted from our earlier framework on legal data moats.

Three moats sound impressive in a deck and do not survive contact with a buyer.

  • Exclusive access to public data. Court data is public by design. Claiming exclusivity over it is not a moat, and it is not a position a serious data company should want.
  • Celebrity lawyer endorsements. Useful for early credibility. They do not win enterprise procurement or developer adoption.
  • A pretty UI. Necessary, not sufficient. A beautiful interface on unreliable data becomes a liability the first time a lawyer misses a hearing.

Nor are these moats: swapping in a different foundation model, clever prompt engineering (an early edge that commoditises fast), a single-jurisdiction app (a great tool for one High Court is an app, not a platform), or an exclusive government tender, which changes hands every cycle.

How the real moats compound

The five real moats are not independent. They feed each other.

  • Coverage improves entity resolution, because more records give the resolver more evidence.
  • Entity resolution makes freshness more valuable, because a fresh update on a well-linked entity is worth more than one on a fragmented record.
  • Freshness drives developer adoption, because builders trust sources that stay current.
  • Developer adoption builds institutional trust, because procurement committees ask who else builds on it.
  • Trust brings the revenue that funds more coverage, which closes the loop.

The same pattern shows up wherever someone owns a data layer: Bloomberg in market data, Plaid in bank data, CIBIL in credit. We looked at those outcomes in What Happens When Someone Owns the Data Layer. The shape differs by market. The flywheel does not.

India’s legal data moats and where value accrues, cover design for the eCourtsIndia blog

Why the moat is harder to copy than people think

The first reaction from anyone hearing this is that scraping is easy and AI makes extraction trivial. Neither holds at scale in Indian law.

Scrapers do not survive captchas, rate limits and schema changes across dozens of court portals for long. Anyone who has tried to run one across the Indian court ecosystem for a quarter has learned that maintenance grows faster than the data. Most teams give up after the second portal change.

AI extraction works at the level of a document. It does not solve fragmentation. You still need a pipeline that reaches the document, retrieves it, and turns it into a record linked to the case, the bench, the parties, the advocates, the earlier orders and the next listing. AI helps with the reading. The rest is engineering, and there is no shortcut.

Licensing the data from an intermediary does not work either. The intermediary usually covers a slice, so the product breaks the moment a user asks about a court outside it. Its schema never quite fits, so a fragile translation layer grows in between. And it can change prices at will. The lesson from Indian broking, where the winners built their own order routing and risk systems, applies here too; we made that argument in The Zerodha Playbook for Indian Legaltech. Real moats sit on owned infrastructure, not on data leases.

A second team starting from zero would also be chasing a moving target. By the time it has matched today’s coverage, the first team has another year of depth, paying contracts, agent traffic and vernacular coverage.

Why the moat compounds with use

This is the part the model-layer conversation misses. A data moat is not a static asset. Every API call an enterprise buyer makes, every MCP (Model Context Protocol) query that Claude or another assistant runs, and every case file a lawyer uploads to the AI Clerk surfaces edge cases that improve the schema. Every new judge name learned, advocate spelling resolved or tribunal added is permanent value the next user inherits. When a user finds a case the index is missing, Add a Missing Case closes the gap for everyone after them.

The model layer does not work that way. A newer model is better; a model with more users is not better because of them. The data layer is the reverse: more use means more coverage, more freshness signals and better resolution. That is a network effect at the level of the substrate.

We showed what this looks like for an in-house counsel running portfolio monitoring through Claude in Litigation Portfolio Monitoring for General Counsel. She never sees the data layer; the chat window is the only surface. But every query she runs makes the substrate sharper for the next user.

Switching costs: the tool that keeps paying you back

A data moat protects the platform from competitors. Switching costs are what the user experiences on top of it. Legal software builds them the way Salesforce, QuickBooks and Tally did, across three layers that compound.

Economists have a name for this. Paul Klemperer’s work on markets with switching costs (1995) and Carl Shapiro and Hal Varian’s Information Rules (1999) describe how a customer facing high switching costs stays with a product for years, absorbs feature gaps and recommends it, because moving costs more than staying. In professional software the cost usually comes in three layers: workflow, data and habit. Legal practice has all three.

Workflow. Salesforce’s hold was never about features; plenty of CRMs matched it. Sales teams had organised their day around its views, and switching meant replacing a shared vocabulary. The legal equivalent is the morning routine: what is listed today, what orders came in yesterday, which deadlines fall this week. Once that cadence runs through one tool, replacing the tool means replacing the cadence.

Data. QuickBooks became hard to leave because years of a small business’s ledger lived inside it. Legal files are richer: briefs, written statements, annexures, witness lists, cross-examination notes, drafts and research memos. Once those sit in a workspace that can search, summarise and cross-reference them, moving a decade of work is not a weekend job.

Habit. Tally is the cleanest Indian case. CA offices train juniors and clerks on it from day one, so a business that wants to switch has to switch its accountant’s habits too. A chamber works the same way. Juniors, clerks and instructing advocates learn the senior’s tools, and changing stack means retraining everyone in the chain.

LayerBefore AIWith AI
WorkflowTool shows cause lists and hearing datesTool drafts, summarises and reminds on its own
DataFiles on a shared drive; the tool indexes metadataFiles become a private retrieval corpus the assistant reads
HabitTeam trained on menu pathsTeam trained on prompts, shortcuts and how the assistant behaves
How AI deepens each switching-cost layer in legal software.

AI compounds all three, because an assistant that has read a lawyer’s last five years of drafts and orders gives better answers than one starting cold. The honest framing is that this is value the user builds, not a trap. Case search, cause lists and directories on eCourtsIndia stay free, and the court data underneath stays public, as we committed to in Public Data, Private Experience. What a user builds on top, their matters, notes and alerts, is what keeps paying them back.

Why legal software has Salesforce-like switching costs, cover design for the eCourtsIndia blog

Why this matters more in India than in the US

American legal AI can underweight the data layer for one reason: the US already has Westlaw, LexisNexis and PACER. The structured legal corpus exists, and application companies plug into it.

The scale of that application layer is now large. Harvey announced a $550 million round at a $15.5 billion valuation on 9 September 2026, up from $11 billion in March, and said it had crossed $400 million in annual recurring revenue with more than 3,000 customers (TechCrunch). Legora reached a $5.6 billion valuation in April 2026. Harvey also said in January 2026, when it bought Hexus, that it planned a Bangalore engineering office. Global players are paying attention to India. We covered the comparables in India’s Legal AI and the Ten Billion Dollar question.

But an engineering office is not a court-data substrate. India has no Westlaw-plus-PACER equivalent for district courts, taluka courts and tribunals. Building one is a precondition for the application layer, not a feature of it. Any application company, global or Indian, that serves Indian litigation without owning or reliably accessing that substrate depends on someone else to build it, maintain it and price it. That is the reasoning behind The AI Agent Layer for Indian Law.

A rubric for founders and investors

When you evaluate an Indian legaltech company, five questions separate signal from noise. Any infrastructure company worth its price answers them specifically. One that dodges is telling you something.

QuestionWhat a good answer looks likeHow eCourtsIndia answers it
What is your actual coverage?Named courts and tiers, not “supported”Supreme Court, all 25 High Courts, district and taluka courts in all 36 states and UTs, 18 tribunal and commission types; 32 crore+ records
What is your freshness SLA?In hours, with how state-level lags are handledDaily refresh, next-day cause lists (7 crore+ entries), WhatsApp and email alerts on tracked cases
How do you resolve entities?Show two records for the same party and how they were linked34 lakh+ advocate profiles with bar-card verification, 82,000+ judges, CNRs minted for 14 tribunals
Who builds on you?A public API reference, SDKs or MCP, and real integrationsREST API with 23 endpoints and Rs 200 free credits; MCP server with 39 tools
What is your security and data-handling posture?Specific certifications held, or a dated plan; data residency answerAsk for the current documentation during procurement; do not accept a slide
Five questions to test an Indian legal data company, with eCourtsIndia’s own answers as of 23 September 2026.

What the moat looks like in five years

Project forward. The Indian legal AI category will likely have one or two platform leaders and a long tail of agent-layer products built on a small number of structured data providers. The agent layer will be fragmented and competitive. The substrate layer will be concentrated.

India has seen this shape before. NPCI sits under PhonePe, Paytm and Google Pay. The Aadhaar stack sits under fintech and insurance KYC. GSTN sits under every accounting product. CIBIL sits under Indian lending, which is the parallel we develop in Why India Needs a CIBIL for Litigation. The platform layer wins through scale and consistency. The application layer wins through experience and distribution. Both can be good businesses. Only one captures the network effect.

As the eCourts system itself matures, with faster uploads and cleaner exports, part of today’s engineering gap will become public utility. The serious response is to keep moving up: better analytics, better summaries, integrations with calendars, document systems and billing. A platform that rests on its crawl alone will be overtaken as the raw data gets easier for everyone.

The data moat in the age of commodity LLMs, cover design variant C for the eCourtsIndia blog

What this means for eCourtsIndia

The model layer is commoditising. In Indian legal AI, the moat that survives is at the data layer: coverage, freshness, structure, entity resolution, vernacular and depth, held together by developer distribution and institutional trust. We are building that substrate. Every agent built on top of it, from Claude to in-house counsel tools to lender risk systems, makes it better for the next user.

If you are building, investing or researching, start with the data itself. Search is free on ecourtsindia.com, developers can begin with the eCourtsIndia API, and AI teams can connect an assistant through the MCP server at mcp.ecourtsindia.com.

TL;DR

  • Frontier models are converging, inference is getting cheaper and routing layers treat models as interchangeable. A model-layer moat is not realistic for a non-frontier company.
  • In Indian legal AI the durable moat is the data layer: coverage, freshness, structure, entity resolution, vernacular OCR and depth.
  • Five moats are real (coverage, entity resolution, freshness, developer distribution, institutional trust). Three are fake (exclusive public data, celebrity endorsements, a pretty UI).
  • Data moats grow with use. Switching costs grow on top of them, and the honest version is a tool that keeps paying its user back.
  • Harvey at $15.5 billion shows how large legal AI can get. In India, the application layer still needs a court-data substrate underneath.

Sources

  • TechCrunch, “Harvey hits $15.5B valuation, months after reaching $11B”, 9 September 2026
  • TechCrunch, “Legal AI startup Legora hits $5.6B valuation”, 30 April 2026
  • TechCrunch, “Legal AI giant Harvey acquires Hexus”, 23 January 2026
  • Klemperer, P. (1995), “Competition when Consumers have Switching Costs”, Review of Economic Studies
  • Shapiro, C. and Varian, H. R. (1999), Information Rules: A Strategic Guide to the Network Economy
  • eCourtsIndia index and MCP server, corpus and directory counts verified 23 September 2026

Read next: The Operating System for Indian Law and The AI Agent Layer for Indian Law.

Frequently Asked Questions

Why is the AI model layer becoming commoditised?

The data moat in the age of commodity LLMs, square social cover for the eCourtsIndia blog

Frontier models are converging on the same benchmarks, inference prices keep falling, and routing layers now swap Claude, GPT and Gemini per task. Differentiation per rupee is shrinking, so a company that does not train frontier models cannot build a lasting moat at the model layer. We explain where the moat moves in The Operating System for Indian Law.

What does a data moat mean in Indian legal AI?

It is the gap between what a platform can do with years of normalised, cross-linked court records and what a new entrant can do on day one. It rests on coverage of every court, daily freshness, structure quality, entity resolution, vernacular OCR and historical depth. You can explore that structured corpus through eCourtsIndia case search.

What are the five real moats in Indian legal data?

Coverage completeness across all states and tribunals, entity resolution at scale, freshness with alerts when a case moves, developer and AI distribution through APIs and MCP, and trust with institutional buyers. They reinforce each other as a flywheel. Developers can test the data layer with the eCourtsIndia API, which comes with Rs 200 in free credits.

Why is exclusive access to public court data not a moat?

Court data is public by design, so anyone can technically crawl it and no one can own it. The durable advantage is the engineering, quality checks and user feedback that make the data reliably usable. Fake moats also include celebrity endorsements and a pretty interface. Our Public Data, Private Experience manifesto sets out the position.

Why does legal software develop Salesforce-like switching costs?

Because it builds three layers at once: the daily workflow of checking listings and orders, years of case files stored in the tool, and the trained habits of a whole chamber. AI deepens each layer, since an assistant that has read your files answers better. Done well, that is value the user keeps. See how the AI Clerk works.

Does Harvey’s growth change the case for an Indian data moat?

The data moat in the age of commodity LLMs, X share card for the eCourtsIndia blog

No. Harvey was valued at $15.5 billion in September 2026 with more than $400 million in ARR, and it plans a Bangalore office. But US legal AI sits on Westlaw, LexisNexis and PACER. India has no equivalent for district courts and tribunals, so a court-data substrate is still a precondition. More in our comparables analysis.

eCourtsIndia is a private legal-technology platform. It is not affiliated with, associated with, or endorsed by the Government of India, the Supreme Court of India or its e-Committee, or any court. Official case information is published on ecourts.gov.in. Always verify details against official court records or certified copies. This article is general information, not legal advice. Spotted an error? Write to support@ecourtsindia.com.

Search 32 crore+ Indian court case records, free

One search across the Supreme Court, all 25 High Courts, district courts and 18 tribunal and commission types. Hearing alerts, AI summaries and an API for developers.