The most popular and comprehensive Open Source ECM platform
Storage: Preserving the Web
Information on the web comes and goes. How many times have you clicked on a link of a web page only to find a non-existent page? Actually the average life of a web page is only around 44 to 77 days, or about a couple of months. Every six months, 10 percent of web pages are lost or recycled. Seems like not a big deal, but to people trying to chronicle the history and life of today, it is very disconcerting.
Google and internet search tools typically maintain a cache of old now defunct web pages. And the Internet Archive organizations keeps old stash of old web pages on their “Wayback Machine“.
In the UK, the British Library is teaming up with IBM to archive web content originating from the UK — web pages from the UK domain which currently has about 8 million sites. The project is being called BigSheets and is built using Open Source components like Hadoop, Nutch, Lucene and Pig.
The goal of BigSheets is not only to archive web content as it is created, it also will attempt to analyze, classify and provide tools that will allow research to be performed with the information. The British Library has been collecting data since 2004 and now has about 5 terabytes of data from 6000 web sites. Prior to BigSheets, the effort required 10 people to manually archive and maintain the data.
The British Library Chief Executive, Dame Lynne Brindley, said the website would aim to create a record of the “major cultural and social issues being discussed online”, and “avoid the creation of a digital black hole in the nation’s memory”.













