Skip to content

Open Legal Data Initiative

The public data layer for Indian legal AI

Case law, statutes, rules, and notifications — structured, licensed, and ready for search, RAG, benchmarks, and production systems. Released on Hugging Face under Apache 2.0.

Judgment records
17.1M
Legal documents
35,851
Courts
SC + 25 HCs
License
Apache 2.0

Coverage

Case law

Supreme Court and High Court judgments with parties, citations, benches, and PDF provenance.

Statutes & instruments

Acts, Rules, Notifications, Circulars, Orders, Regulations, and guidelines — Central and state.

Legal metadata

Court, jurisdiction, authority, dates, disposition, and quality flags as first-class columns.

Structured corpora

Parquet tables for analytics, streaming loaders, and Dataset Viewer — not a pile of PDFs.

Public datasets

Two live releases on Hugging Face. More annotated corpora ship through Pro and Enterprise.

Case lawApache-2.0Parquet

Indian Case Laws

Structured metadata for Indian court judgments, with provenance to source JSON and original PDFs. Built for search, retrieval, ranking, RAG, and legal NLP evaluation.

Records
17.1 million
Size
10.9 GB
Timeframe
1950–2026
Coverage
Supreme Court of India + 25 High Courts
  • 17.1M judgment records with parties, docket, CNR, bench, and disposition
  • Neutral citations and law-report citations where available
  • Parser diagnostics and quality flags for safe filtering
  • Links to original eCourts judgment PDFs
case_title, parties, docket_number, cnr_number
neutral_citation, law_report_citation
court_name, bench_name, presiding_judge, coram
decision_date, disposition_text
source_json_s3_url, source_pdf_s3_url
indexable_text, headnote_text, quality_json
Open on Hugging FaceKanoonGPT/indian-case-laws

How access works

Open data for training and research. Annotated inference in Pro. Licensed training data for enterprises.

Research & training

Open data

Free

  • Apache-2.0 datasets on Hugging Face
  • Case law metadata and statutory documents
  • Parquet + Dataset Viewer + streaming loaders
Browse on Hugging Face

Product inference

Pro

From ₹129

  • Annotated data for inference in KanoonGPT Pro
  • Section-wise Bare Acts, chat, and KanoonFM
  • Same account on web and mobile
Get Pro access

Training & licensing

Enterprise

Custom

  • Annotated corpora for commercial training
  • Licensing and volume terms
  • Support for production legal AI systems
Talk to us

Built for teams who ship

Legal search & RAG

Retrieve the right judgment or statute, then cite section and court context.

Benchmarks & eval

Score Indian legal NLP models on real courts, acts, and notifications.

Classification

Type, jurisdiction, bench, and disposition labels for analytics pipelines.

Copilots & agents

Give agents a public, structured Indian-law layer instead of scraping gazettes.

Responsible use

  • Verify important legal facts against official gazettes, eCourts PDFs, and court records.
  • Do not treat dataset rows or model outputs as legal advice.
  • Metadata is machine-extracted and may be incomplete or noisy — filter before high-stakes use.
  • Add citations and human review in any product that answers legal questions.

Need annotated data or a commercial license?

Open datasets stay public. Talk to us for enterprise training data, or start Pro for product inference.