TU Darmstadt / ULB / TUbiblio

C4Corpus: Multilingual Web-size corpus with free license

Habernal, Ivan and Zayed, Omnia and Gurevych, Iryna
Calzolari, Nicoletta and Choukri, Khalid and Declerck, Thierry and Grobelnik, Marko and Maegaard, Bente and Mariani, Joseph and Moreno, Asuncion and Odijk, Jan and Piperidis, Stelios (eds.) (2016):
C4Corpus: Multilingual Web-size corpus with free license.
In: Proceedings of the 10th International Conference on Language Resources and Evaluation (LREC 2016), European Language Resources Association (ELRA), Portoroz, Slovenia, [Online-Edition: http://www.lrec-conf.org/proceedings/lrec2016/pdf/388_Paper....],
[Conference or Workshop Item]

Abstract

Large Web corpora containing full documents with permissive licenses are crucial for many NLP tasks. In this article we present the construction of 12 million-pages Web corpus (over 10 billion tokens) licensed under CreativeCommons license family in 50+ languages that has been extracted from CommonCrawl, the largest publicly available general Web crawl to date with about 2 billion crawled URLs. Our highly-scalable Hadoop-based framework is able to process the full CommonCrawl corpus on 2000+ CPU cluster on the Amazon Elastic Map/Reduce infrastructure. The processing pipeline includes license identification, state-of-the-art boilerplate removal, exact duplicate and near-duplicate document removal, and language detection. The construction of the corpus is highly configurable and fully reproducible, and we provide both the framework (DKPro C4CorpusTools) and the resulting data (C4Corpus) to the research community.

Item Type: Conference or Workshop Item
Erschienen: 2016
Editors: Calzolari, Nicoletta and Choukri, Khalid and Declerck, Thierry and Grobelnik, Marko and Maegaard, Bente and Mariani, Joseph and Moreno, Asuncion and Odijk, Jan and Piperidis, Stelios
Creators: Habernal, Ivan and Zayed, Omnia and Gurevych, Iryna
Title: C4Corpus: Multilingual Web-size corpus with free license
Language: English
Abstract:

Large Web corpora containing full documents with permissive licenses are crucial for many NLP tasks. In this article we present the construction of 12 million-pages Web corpus (over 10 billion tokens) licensed under CreativeCommons license family in 50+ languages that has been extracted from CommonCrawl, the largest publicly available general Web crawl to date with about 2 billion crawled URLs. Our highly-scalable Hadoop-based framework is able to process the full CommonCrawl corpus on 2000+ CPU cluster on the Amazon Elastic Map/Reduce infrastructure. The processing pipeline includes license identification, state-of-the-art boilerplate removal, exact duplicate and near-duplicate document removal, and language detection. The construction of the corpus is highly configurable and fully reproducible, and we provide both the framework (DKPro C4CorpusTools) and the resulting data (C4Corpus) to the research community.

Title of Book: Proceedings of the 10th International Conference on Language Resources and Evaluation (LREC 2016)
Publisher: European Language Resources Association (ELRA)
Uncontrolled Keywords: UKP_reviewed;AIPHES_corpus
Divisions: 20 Department of Computer Science
20 Department of Computer Science > Ubiquitous Knowledge Processing
DFG-Graduiertenkollegs
DFG-Graduiertenkollegs > Research Training Group 1994 Adaptive Preparation of Information from Heterogeneous Sources
Event Location: Portoroz, Slovenia
Date Deposited: 31 Dec 2016 14:29
Official URL: http://www.lrec-conf.org/proceedings/lrec2016/pdf/388_Paper....
Identification Number: TUD-CS-2016-0023
Related URLs:
Export:

Optionen (nur für Redakteure)

View Item View Item