eDiscovery is the process of recovering evidence in a legal case from electronically stored files, documents, and databases. The process of eDiscovery starts with a request from an attorney to produce all documents related to specific keywords that were produced over a specific time period. Evidence is then extracted from those electronic documents produced, analyzed using digital forensics, and used as part of a legal presentation in court.
The traditional discovery process of paper, email and electronic office documents involves the manual identification, review and classification of documents that are potentially related to the legal request. The discovery classification process of documents, sometimes referred to as the coding of the documents, involves the grouping of relevant documents by date, keyword and custodian. Historically the coding process has been very tedious and coding often accounts for 60-70 percent of the total discovery costs.
The manual coding of documents is often referred to as ‘linear coding’ — the process of examining one document at a time. But as the number of documents in organizations are growing rapidly, the costs of linear coding have become unsustainable. The manual review process is often subject to error.
Maura R. Grossman, counsel at Wachtell, Lipton, Rosen & Katz in New York City, reported by
Joe Dysart of the
ABAJournal, that “There has been a long-standing myth in the legal field that exhaustive manual review is the gold standard, or nearly perfect, but that has been shown to be a fallacy. Humans manually reviewing large numbers of documents for responsiveness make errors.”
Software-based tools are increasingly being used to speed up coding process by automatically identifying documents that have a high probability of relevance.
eDiscovery has been aided with the increased sophistication of electronic search technology. Using electronic search, many documents can be quickly identified which contain the keywords included in the discovery request. But even with search, the number of documents identified can be huge, and very often many of the candidate documents when reviewed turn out to be ‘false positives’ — documents that included the use of a particular keyword used in the search but which turned out to be non-relevant to the case at hand.
Software which combines standard search with more sophisticated analytics and artificial intelligence is being designed to specifically aid in the eDiscovery process. This type of eDiscovery software is being dubbed ‘predictive coding’. It’s another example for how analytics is dramatically changing the capabilities of software applications.
‘Predictive Coding’ isn’t totally new — in fact, it’s fairly well known. A number of companies like
Recommind,
Equivio and
FTI are using technologies which they label ‘predictive coding’ or ‘predictive analytics’ software which speeds the eDiscovery identification and review process.
Dysert explains the concept of predictive coding as follows: “Essentially, the software works by delivering results based on a barrage of keyword inputs, which are tweaked by a seasoned attorney who then continually refeeds the best resulting documents back into the system as examples until the software ‘learns’ what the attorney is really looking for.”
But despite the promise of the technology and success so far, defensibility in court remains a sticking issue. Symantec found that 97 percent of compliance and records managers were familiar with the concept of predictive coding, 69 percent say that they have not adopted the technology. Matthew Nelson, e-discovery counsel with Symantec, said that “The survey results were pretty telling because they gave us some good indicators of what the problems are in terms of lackluster adoption. I think it boils down to their concerns about accuracy, easy of use, and cost. I think accuracy and ease-of-use concerns breed a larger concern around defensibility. For attorneys it’s tough to defend a technology they don’t understand.”