Using India's e-courts data to answer questions around access to justice

The e-courts database has become an important repository of India's judicial data, and has opened up possibilities of deep empirical research. We describe the dataset, explore what can be done using the data, and note the problems with the data that pose barriers to its wider use.

Understanding how the law works in practice requires access to legal data. For years, journalists and researchers interested in questions around what happens inside courts have had to rely on either the government's own aggregation of legal data, such as National Crime Records Bureau's (NCRB) reports and in occasional parliamentary answers, or by navigating the severely gatekept and tedious process of accessing physical court record rooms and archives.

The annual Crime in India reports by the NCRB have been the most widely used source of legal data. However, this too is an imperfect source. By the principal-offence rule that the report follows, a crime is recorded in NCRB statistics only under the offence carrying the highest sentence cited in the First Information Report (FIR). In a case of rape and murder, for instance, the offence of rape is essential to understanding the true nature of the crime, but since murder carries a higher possible sentence, the offence would be counted in NCRB statistics under the number of incidents of murder and not of rape.[1]

Though NCRB reports have some aggregated data with respect to the court process, including number of cases for trial and conviction rates, they do not reveal more granular details around the legal process. The gender and age profile of the complainant and accused, number of witnesses, production of evidence, time between hearings, reasoning offered by the court, and combinations of these indicators, for instance, can offer much greater insight into how justice is truly delivered, or not.

The NCRB or any other data source, also does not say anything about the number and kinds of cases under civil laws, such as contractual disputes, that are before the courts.

In this context, the launch of the e-courts Mission Mode Project is a significant step in expanding the universe of legal indicators available for research, shedding greater light on the legal process.

The e-courts project allows litigants, lawyers and the larger public to access legal data, including case status, its history, future hearing dates, and case orders and judgments, among other things, from the country's district and subordinate courts, High Courts and the Supreme Court of India.[2]

However, many gaps remain, severely restricting the possibilities of working with this dataset. In this piece, we describe the dataset, outline some important research that has used it, and flag some key problems with the accuracy and completeness of the data.

E-courts database and its creation

The e-courts project's origins lie in the National Policy and Action Plan for the Implementation of Information and Communication Technology in the Indian Judiciary, 2005, that was drafted by the newly formed e-committee of the Supreme Court of India. It was drafted with the final goal of improving judicial productivity and making the judicial system more affordable, accessible and transparent.

Following this plan, the judiciary and the central government went on to install digital infrastructure in all courts of the country, put in place an integrated case management system, and created a web portal to make judicial data available in a single window.

This e-courts national portal was launched in August 2013 with Phase-I of the project (2007-2015) making records available for about 25 million cases in district courts of the country. Phase-II (2015-2023) involved installing video-conferencing facilities, developing a mobile application and a centralised online payment system, and connecting courts with prisons, police and prosecution through an interoperable data system. During this phase, a unified e-court services portal for High Courts was also launched.

One of the most important developments in the second phase was the launch of the National Judicial Data Grid (NJDG) and the addition of all levels of courts on it in phases. The NJDG is a centralised database which displays summary statistics with respect to pendency, disposal and filings for all district courts, High Courts and the Supreme Court making it possible to track courts on these indicators.

Currently in its third phase (2023-2027), the e-courts system offers case information on 300 million pending and disposed cases, and nearly 400 million orders/judgments.[3]

Details of a new case and subsequent updates are entered into this database using the Case Information System (CIS) software by the court staff.[4] This includes information about the court, basic information about the petitioner or the plaintiff and the respondent or the defendant, details of the subordinate court, the Act and section involved, the FIR number and the police station if it is a criminal case, and details of the dispute.

After each court proceeding, the staff also updates this information for a case and uploads orders. The CIS also allows the court to receive documents such as chargesheets from other pillars of the criminal justice system - police stations, prisons and the prosecutor through the Interoperable Criminal Justice System (ICJS), a data system that connects all four pillars.

The CIS has dashboards that allow the judge and the court staff to see the day's cases, cases pending in the courtroom, cases that are dormant, cases for which they are yet to upload judgments, those disposed of that month, as well as information about undertrials.

Accessing the e-courts database

At present, the e-courts portal offers a number of services to both litigants and advocates, including accessing case status information, order and judgment search, cause-lists and online case filing ("e-filing").

Since the e-courts and the case status data and orders on it are public, anyone, including researchers, can access these using any of the available search fields. The same data can also be accessed on the individual websites of each district court or the high court.[5]

Regardless of the portal, a case - whether pending or disposed of - can be located with the help of the name of the court using any of the following search parameters: Case Number Record (CNR), a unique 16-digit ID assigned to every case; case number; filing number; the name of either party or litigant; name of the advocate; the FIR number and name of the police station in case of a criminal case; or the Act under which the case has been filed.[6]

The case record obtained has over 30 distinct output variables that help make sense of the case:

  1. Basic details: case type, filing number and filing date, e-filing number and e-filing date, registration number, registration date, and CNR.
  2. Case status: first and next hearing dates, current stage of the case, court number and designation of the judge, or if disposed of, the nature of the disposal and the date of the disposal.
  3. Litigant details: petitioner and their advocate's names, and respondent and their advocate's name.
  4. Law and FIR details: Act and section involved; and if it is a criminal case, the FIR number and the police station.
  5. Case history: date and purpose of past hearings and the judge before whom it was heard.
  6. Other details: information about transfer of case including the previous judge's designation and the history of the case if it was before a subordinate court.
  7. Orders and judgments: Downloadable orders and judgments passed in the case.

Many researchers and firms have scraped this portal to create databases of this data. The Development Data Lab (DDL), a research portal, created a public dataset of 81 million case records from 2010-2018 extracted from the e-courts portal for district and subordinate courts that we utilised for our analysis of the kinds of cases that go before the courts.

Possibilities of empirical research through e-courts

Access to case data in a single window has opened up many possibilities for researchers interested in understanding the operation of law and the judicial system, but with several challenges.

Counting pendency, filings, disposals and caseload

One of the key macro-indicators that has been made easier to access and even easier to collate has been simply the count of pending cases. Prior to the launch of the NJDG for High Courts in July 2020, this was collected at each High Court and compiled and reported quarterly by the Supreme Court in its Court News newsletter. After the launch of the NJDG, data on pending cases, filings and disposals across all courts is now available daily. This can be further filtered by district, down to individual courtrooms.

The availability of this data is a substantial improvement on the data that was earlier available from the Supreme Court. Pendency in the Rajasthan High Court appeared to nearly double between 2018 and 2019 with no apparent cause. Following the launch of the NJDG in 2021, pendency in the Bombay and Madras High Courts appeared to nearly double, indicating that the earlier Supreme Court data drawn from the High Courts undercounted the number of cases in the two courts. The data has also been produced in Parliament with no explanation or clarification ever offered. Since 2020, however, the NJDG data has been more consistent year-on-year.

The usage of laws, judicial performance and outcomes

Besides macro indicators on the judiciary, researchers have used the e-courts dataset to ask deeper questions about certain laws: how frequently they are invoked, their geographical spread, the nature of cases filed under them and their outcomes. One such use is Enfold Proactive Health Trust's research on the functioning of the law on child labour, the Child and Adolescent Labour (Prohibition and Regulation) Act, 1986 (CALPRA).[7] In their research using the e-courts database in 2024, Enfold found long delays, high instances of the child victims not appearing for their testimony or appearing but not testifying against the accused, and a high rate of acquittal. It made a larger point about the absence of support for the victims and their families by the child welfare system.

Enfold found that using e-courts meta-data allowed it to find data on the actual number of criminal trials under a legislation, which NCRB data was unable to do.[8] According to Enfold's study mentioned earlier, while NCRB reported 1329 cases under the CALPRA for six selected states between 2015 and 2022, the researchers were able to find about seven times more cases in the e-courts dataset where the Act was involved in these years.

The e-courts dataset also let Enfold ascertain the number of cases where another offence in addition to CALPRA was invoked, which is not possible with NCRB data. For example, data on cases in which the trafficking provision under the Indian Penal Code or a provision under the law on bonded labour was invoked in addition to CALPRA provided a window into cases of child labour upon their trafficking, which is perhaps more serious than employing them near their home.[9]

Enfold combined this meta-data analysis with analysis of judgments in some of these cases that are also available on the e-courts portal. This effectively turned qualitative data into quantitative.[10] Analysis of the text on these variables offered deeper understanding of the process and evidence used in the courts, profile of the victims and the accused, and importantly factors that led to a certain outcome.

The e-courts database is also the only source of data on India's civil laws. For an analysis of the Indian Divorce Act, for instance, there is no other source to indicate the number of divorce cases filed across the country annually. E-courts data can tell us how many such cases were filed, whether they were filed by the parties together cooperatively or by one of them with the other contesting it, which states most of these were filed in, how long they took and what the outcomes were.

Building on e-courts data using qualitative and other kinds of data

Another category of studies is one in which researchers use the e-courts data as a tool to find and identify cases, and use other sources to add qualitative details.

Researchers used e-courts data to identify 1,125 cases involving 925 accused persons under Unlawful Activities (Prevention) Act, 1967 in the state of Karnataka between 2005 and 2015.[11] They then examined FIRs, court orders, and other documents, and interviewed some of the accused persons to build a more detailed picture of these cases.[12]

Researchers at the Development Data Lab used a dataset that they had extracted from e-courts, and combined it with a machine learning algorithm that labelled judges and litigants based on their gender and religion.They used this to test the effect of the judge's and the litigant's identity on the outcomes of cases.[13]

The availability of this data in public does raise data privacy concerns. Information given under litigant details and contained in the orders passed can sometimes include phone numbers and addresses. There is also a special interest in protecting the confidentiality of children and other vulnerable groups that requires attention and careful balancing with the interests of accessibility and transparency.[14]

Data for India's work using e-courts data

Data from e-courts also allows researchers to capture the full landscape of civil and criminal cases in India, which cannot be captured from any other data sources. Using a set of 44 million cases disposed between 2010 and September 2020, we found that litigations in India's subordinate courts are dominated by criminal cases under the Indian Penal Code, Motor Vehicles Act, Negotiable Instruments Act, Excise and Police Acts.

Using data on pendency, filing and disposals directly from NJDG, we looked at the extent of the judicial backlog. We found that the total pendency at the end of 2025 had gone up by 80% over the previous decade and that, one in three cases before the subordinate courts and one in two cases before the High Courts were older than 5 years.

Some challenges with using e-courts data

Although it has opened up the gates to analyse court processes, especially at the level of district courts, working with the e-courts system has not been without challenges.

Inaccurate count

At the outset, e-courts data cannot be taken on face value to make any quantitative claim and policy decisions about judges caseload and other pendency indicators. When the NJDG dashboard shows a pendency of 50 million cases before district courts, this does not suggest 50 million legal disputes. Most legal disputes spur more than one 'case': a main case and sub-case. For example, a criminal case will have a main case that consists of a full criminal trial and each hearing for a bail application is a sub-case. A civil suit - a claim made by a private party against another - may have several applications moved by parties during the course of the life of the main case (called 'interlocutory applications'), each recorded as a sub-case.

Researchers from Cross Disciplinary Knowledge Data Research (XKDR) estimated that between 2017 and 2024, on an average only 60% 'cases' filed before the Bombay High Court every year were main cases.[15] This inaccuracy, they said, exaggerated the appearance of pendency.

Some of these discrepancies arise from the creation of the e-courts database as primarily an administrative dataset for court officers and litigants. However in the absence of other data, the Parliament, the Courts and other institutions are now using e-courts data to determine policies around ideal case timelines, case-loads and judges' strength. The distinction between main cases and sub-cases highlights the pitfalls of this approach.

Even when main cases can be isolated, there are concerns that the case count is inflated by record duplications. In their book, lawyers Prashant Reddy T. and Chitrakshi Jain suggest that cases are also possibly double or triple-counted.[16] In a limited study of one courtroom in Ernakulam city in Kerala, they found about a fourth of the pending cases had been counted more than once.

Unreliable records

A researcher who wants to study cases more deeply using the data from the case records would come across another set of challenges that call into question the quality of the data. These insights come from our work using e-courts data, as well as similar findings by researchers from the National Institute of Public Finance and Policy on contractual disputes and others like Enfold and Article 14 mentioned earlier.

These problems can be understood under three categories:

1. Improper labelling or mislabelling

Two of the most basic indicators in the e-courts database are the name of the Act and the name of the section under the Act for a given case. However, we found that the Act name variable often listed only a procedural law- the Code of Criminal Procedure (CrPC) or the Code of Civil Procedure (CPC) - and no other Act. Procedural law, by definition, is involved in cases which come before a court. As a result, the mere mention of the procedural law reveals little about the nature of the case, or the specific offence that a person is facing trial for. From a sample set drawn from the larger e-courts dataset, we found that a quarter of all cases are tagged only under a procedural law.[17]

Further, there were records in which the Act and section name was either missing or with clear typographical errors.

2. Lack of standardisation

One major problem with using the e-courts database is the lack of standard practice with respect to referring to an Act. For instance, we found 220 different spellings of the Indian Penal Code.

Others have found similar anomalies. While working on a dataset of contractual dispute cases, NIPFP found that cases under the Arbitration and Conciliation Act, 1996, were filed under 58 unique names, while those under the Negotiable Instruments Act were filed under 21 unique names."[18]

The lack of standardisation is a defect not just for Act names but almost all other entries. NIPFP found that the raw data on e-courts portal records a total of 844 'case types', which are the categories under which the case is filed, even though the NJDG, which provides a numerical summary from the e-courts, reported only 31 case types.[19]

The problem magnifies considerably when we move from Act names to section numbers, where the variations in the way the sections are recorded are too numerous to be recorded comprehensively. In a small sample of cases NIPFP looked at, the section was incorrectly reported in nearly 30% of the cases. The field was either left blank or mentioned the case type or Act name instead of the section number.

These defects make analyses of the number of cases recorded under an Act or section challenging, requiring standardisation using machine learning, time-consuming and prone to subjective calls taken by individual researchers.

The same defects also make any kind of analysis of the outcomes of a disposed case, without having to go through the judgment of the case next to impossible. Researchers including us would like to analyse the share of cases that end on convictions, acquittals or similar outcomes of interest. However in our sample set, we find over 4,000 distinct ways that the 'nature of disposal' has been described, with not just multiple variations in spellings but also broad descriptions such as "disposed" and "decided" that do not indicate the precise outcome.

While the kinds of problems outlined here can be detected and remedied both manually or using an automated process- a machine can be taught to identify expected errors and standardise spellings- this type of error is so far impossible to rectify without manual check against original case records.

3. Mislabelling of cases

Some cases have clearly been erroneously labelled, and these errors are only discovered by chance.

We found that many cases that were recorded under Registration of Births and Deaths Act, 1969, especially in Karnataka, were cases under the Negotiable Instruments Act, with presumably Section 13(3) of the former getting mixed up with Section 138 of the latter.

Similarly, we found numerous cases under the Prevention of Illicit Traffic in Narcotic Drugs and Psychotropic Substances Act, 1988 (PIT-NDPS) in Maharashtra that were actually cases under The Immoral Traffic (Prevention) Act, 1956, which is colloquially abbreviated as PITA.

Looking ahead

The judicial establishment is not unaware of the issues with e-courts data. The Supreme Court of India's National Court Management Systems (NCMS) committee tasked with this challenge itself noted in 2024, "While the availability and tracking of judicial data with some granularity has been made possible now with the NJDG, several challenges persist with data analysis in a manner so as to inform policy making on judicial case load management. There are vast differences in case categorization, nomenclature, and methodology in each State/ High Court, and there is presently no uniformity on these aspects, to enable meaningful data analysis at a national level."[20]

In 2020, the Supreme Court's e-committee constituted an expert group drawn from organisations Daksh, Vidhi Centre for Legal Policy and Agami to draft the report on the vision document of e-court project's Phase-III.[21] The report noted poor design of implementation processes to have been a challenge, especially mentioning the lack of uniformity in case nomenclature to be a problem. The report recommended simplified procedures, creating new digital infrastructure and putting in place new institutional and governance frameworks to be the important building blocks of Phase-III.

There have been improvements in the quality of data from Phase-I of the e-courts project to the present Phase-III. The Case Information System 4.0, the fourth version of the backend software and the administrative manuals for the recent versions shows the efforts to address some of these problems. In particular, courts are attempting to standardise the way that data can be entered, using more drop-down menus. This would help in cases where text instead of numbers has been entered for section numbers, for instance - the problem would not have arisen if the field allowed only numeric values for section numbers.

However, whether the software updates will wipe out these problems remains to be seen, as some of these problems persist even now. For example, we noticed that subordinate courts in Haryana, Karnataka and Kerala were continuing to allow new record creation under non-standardised variations of Indian Penal Code such as "IPC" and "I.P.C.(Police)", while a vast majority of states also have new cases improperly labelled under the Bharatiya Nagarik Suraksha Sanhita, the new criminal procedural law, instead of the specific criminal offence that is committed.

Even if the data labelling process were to improve and fewer errors of this kind were to happen in the future, wrongly recorded legislations in past data are unlikely to be revisited and corrected. This suggests that e-courts data, particularly for historical analysis, will continue to need to be used with caution.


[1] Since 2017, the NCRB has additionally disclosed the number of incidents of rape with murder separately.

[2] While individual case data can be accessed, researchers like the Development Data Lab have scraped the underlying database by writing code to create public datasets.

[3] Ministry of Law and Justice, Government Accelerates Digital Transformation of Judiciary Under eCourts Phase III to Reduce Case Pendency, Press Bureau of India (26 July 2026).

[4] E-Committee, the Supreme Court of India, Case Management through Case Information System (CIS) 4.0 (May 2025).

[5] As an example, the Bengaluru Urban District's website is https://bengaluru.dcourts.gov.in/.

[6] A key difference between the district court website and the other two is the absence of section number as a search field alongside Act name.

[7] Enfold Proactive Health Trust, Decoding data on Implementation of the Child and Adolescent Labour (Prohibition and Regulation) Act, 1986 (2024).

[8] Enfold Proactive Health Trust, The Possibilities of eCourts Data for Advancing Research on Law Implementation Experiences of using Case Metadata to understand the Child and Adolescent Labour (Prohibition and Regulation) Act, 1986 (Mar 2025).

[9] Other things Enfold was able to find out was the proportion of convictions, acquittals and settlements between parties, outcomes of bail applications, the median time of disposal and pendency in each state, and some combination of these indicators, for instance, the median time of disposal if the outcome was an acquittal.

[10] Many other researchers have also studied texts of the judgment using only qualitative research methods.

[11] Lubhyathi Rangarajan, India's Terrorism Law Weaponised Against Muslims Under BJP In Karnataka, New Research Shows, Article 14 (2 February, 2026).

[12] Article 14 researchers also documented their methodology, the process of narrowing down on Karnataka and the challenges of working with the e-courts database in a research paper. Lubhyathi Rangarajan, Sakshi Rai and Nikita Bansal, What Counts as Data? Empirical Legal Research with India's eCourts Portal, Socio-Legal Review, Vol. 22(1) 2026.

[13] Elliott Ash et al, In-Group Bias in the Indian Judiciary: Evidence from 5 Million Criminal Cases, The Review of Economics and Statistics (2025).

[14] Haq Centre for Child Rights and Civic Data Lab, #Data4Justice - Unpacking Judicial Data to Track Implementation of the POCSO Act in Assam, Delhi & Haryana(2012 to April 2020) (2021).

[15] Pavithra Manivannan et al, Beyond Pendency: Counting Cases Correctly, The Law, Economics and Policy (LEAP) Blog (5 October, 2025).

[16] Prashant Reddy T. and Chitrakshi Jain, Chapter 6: The Missing Judicial Statistics, Tareekh Pe Justice: Reforms for India's District Courts, Simon and Schuster India (2025).

[17] The sample set we created consisted of all cases in all of India's district courts, that had been filed on or after 1 January 2018, and disposed of by September 2020. This set had 18 million individual cases. We used the data for the three latest years on the presumption that the quality of data would have improved as time passed.

[18] Devendra Damle and Tushar Anand, Problems with the E-courts Data (NIPFP Working Paper) (July 2020).

[19] The 'case type' typically describes the kind of relief sought from the court or the Act under which a case is filed. The case type along with a numeric code and the year of the filing forms the 'case number'. For example, where the case numbers are Original Suit No. 100/2026 or Criminal Miscellaneous Bail Application No. 1234 of 2025, 'Original Suit' and 'Criminal Miscellaneous Bail Application' are the case types, respectively.

[20] The Supreme Court of India, National Court Management Systems (NCMS) - Policy & Action Plan, 2024 (2024). Estimating ideal case-load on judges based on such data has its own challenge as different case-types and subjects will have different average and ideal time frames for completion, and successive committees and other bodies have encountered the problem of not having data that is granular enough to do this calculation.

[21] E-committee, the Supreme Court of India, Digital Courts Vision & Roadmap e-Courts Project Phase-III (2022).

In this article list close
Footnotes chevron_forward

    To cite this article:

    Using India's e-courts data to answer questions around access to justice by Ameya Bokil, Data For India (September 2026): https://www.dataforindia.com/e-courts-data/

    Read next