Saved datasets
Last updated
Download format
Usage rights
License from data provider
Please review the applicable license to make sure your contemplated use is permitted.
Topic
Free
Cost to access
Described as free to access or have a license that allows redistribution.
2 datasets found
  1. Webis-Sentences-17

    • webis.de
    Updated Feb 27, 2017
  2. Webis-Simple-Sentences-17 Corpus

    • zenodo.org
    gz
    Updated Feb 27, 2017
  3. Not seeing a result you expected?
    Learn how you can add new datasets to our index.

Share
FacebookFacebook
TwitterTwitter
Email
Click to copy link
Link copied
Close
Cite
Kiesel, Johannes; Stein, Benno; Lucks, Stefan (2017). Webis-Sentences-17 [Dataset]. http://doi.org/10.5281/zenodo.205950
Organization logo

Webis-Sentences-17

Dataset updated Feb 27, 2017
Dataset provided by
Bauhaus-Universität Weimarhttp://www.uni-weimar.de/
The Web Technology & Information Systems Network
Authors
Kiesel, Johannes; Stein, Benno; Lucks, Stefan
License

Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically

Description

The Webis-Sentences-17 corpus is a collection of 3,369,618,811 sentences extracted from the ClueWeb12 web crawl. It is designed to allow for statistical analyses of human-written sentences. More details on the sentence extraction can be found in the associated publication. The Webis-Simple-Sentences-17 corpus contains 471,085,690 English sentences from the Webis-Sentences-17 corpus. The sentences were sampled to achieve a level of sentence complexity similar to the one of sentences that humans make up as a memory aid for remembering passwords. Sentence complexity was determined by syllables per word. Both corpora are split in training and test set as they are used in the associated publication. The test set is extracted from part 00 of the ClueWeb12, while the training set is extracted from the other parts.

Search
Clear search
Close search
Google apps
Main menu