copy delete add this publication to your clipboard
community post
history of this post
URL
DOI
BibTeX
EndNote
APA
Chicago
DIN 1505
Harvard
MSOffice XML

Creating a Billion-scale Searchable Web Archive

D. Gomes, M. Costa, D. Cruz, J. Miranda, and S. Fontes. Proceedings of the 22Nd International Conference on World Wide Web, page 1059--1066. New York, NY, USA, ACM, (2013)
DOI: 10.1145/2487788.2488118

Abstract

Web information is ephemeral. Several organizations around the world are struggling to archive information from the web before it vanishes. However, users demand efficient and effective search mechanisms to access the already vast collections of historical information held by web archives. The Portuguese Web Archive is the largest full-text searchable web archive publicly available. It supports search over 1.2 billion files archived from the web since 1996. This study contributes with an overview of the lessons learned while developing the Portuguese Web Archive, focusing on web data acquisition, ranking search results and user interface design. The developed software is freely available as an open source project. We believe that sharing our experience obtained while developing and operating a running service will enable other organizations to start or improve their web archives.

Links and resources

BibTeX key: gomes2013creating
entry type: inproceedings
address: New York, NY, USA
booktitle: Proceedings of the 22Nd International Conference on World Wide Web
year: 2013
pages: 1059--1066
publisher: ACM
series: WWW '13 Companion
acmid: 2488118
isbn: 978-1-4503-2038-2
location: Rio de Janeiro, Brazil
numpages: 8
DOI: 10.1145/2487788.2488118
url: http://doi.acm.org/10.1145/2487788.2488118

@jaeschke's tags highlighted

Cite this publication

search on

Meta data

Last update 8 years ago
Created 8 years ago

Comments and Reviews
(0)

There is no review or comment yet. You can write one!

BibSonomy

copy delete add this publication to your clipboard
community post
history of this post
URL
DOI
BibTeX
EndNote
APA
Chicago
DIN 1505
Harvard
MSOffice XML

Creating a Billion-scale Searchable Web Archive

Abstract

Links and resources

Tags

community

Cite this publication

More citation styles

search on

Meta data

Comments and Reviews
(0)

BibSonomy

copydeleteadd this publication to your clipboardcommunity posthistory of this postURLDOIBibTeXEndNoteAPAChicagoDIN 1505HarvardMSOffice XML Creating a Billion-scale Searchable Web Archive

Abstract

Links and resources

Tags

community

Cite this publication

More citation styles

search on

Meta data

Comments and Reviews (0)

copy delete add this publication to your clipboard
community post
history of this post
URL
DOI
BibTeX
EndNote
APA
Chicago
DIN 1505
Harvard
MSOffice XML

Creating a Billion-scale Searchable Web Archive

Comments and Reviews
(0)