The most popular and comprehensive Open Source ECM platform
Open Data: Poor Quality Threaten Usefulness of Data Sets
Open Data, data published by governments, organizations and businesses for use in the public domain, have touted economic and social benefits.
A big problem though with Open Data is that it’s difficult to know the quality of the data that’s being published, and unfortunately often the accuracy or completeness of the data sets that are released are questionable.
Kathryn Stack, U.S. Office of Management and Budget’s adviser for evidence-based innovation, said that “I think we need a really honest conversation about data, and sometimes it is hard for people to admit how much bad data we have that we collect and move around and push around.”
Harvey Lewis, head of data analytics at consulting firm Deloitte, said that the quality of government-released data has been “patchy. There’s some standardisation but it’s not complete, so cross-referencing is difficult. It’s of varying quality from different departments – that’s a challenge.”
Victoria Lemieux,Senior Public Sector Specialist said that “far from being a one-off problem, research suggests that this issue is ubiquitous and endemic. Some estimates indicate that as much as 80 percent of the time and cost of an analytics project is attributable to the need to clean up ‘dirty data'”.
Poor data quality, especially when used in conjunction with creating government policy or making business decisions can lead to misplaced funding, bad decisions and even fraud.
It’s hard to prepare and clean data and often the organization publishing the data isn’t aware all of the problems. A spokeswoman for the government-backed Open Data Institute (ODI) said that “[civil servants] just don’t have the skills. They don’t understand the difference between good and bad data.”













