Saved datasets
Last updated
Download format
Usage rights
License from data provider
Please review the applicable license to make sure your contemplated use is permitted.
Topic
Free
Cost to access
Described as free to access or have a license that allows redistribution.
3 datasets found
  1. Webis-QSeC-10

    • webis.de
    3256198
    Updated 2010
  2. Webis Query Segmentation Corpus 2010 (Webis-QSeC-10)

    • zenodo.org
    zip
    Updated Jul 23, 2010
  3. Webis Query Spelling Corpus 2017 (Webis-QSpell-17)

    • zenodo.org
    zip
    Updated Aug 11, 2017
  4. Not seeing a result you expected?
    Learn how you can add new datasets to our index.

Share
FacebookFacebook
TwitterTwitter
Email
Click to copy link
Link copied
Close
Cite
Matthias Hagen; Martin Potthast; Benno Stein (2010). Webis-QSeC-10 [Dataset]. http://doi.org/10.5281/zenodo.3256198
Organization logoOrganization logoOrganization logo

Webis-QSeC-10

10 scholarly articles cite this dataset (View in Google Scholar)
3256198Available download formats
Dataset updated
2010
Dataset provided by
Leipzig Universityhttp://www.uni-leipzig.de/
Martin-Luther-University Halle-Wittenberghttp://www.uchicago.edu/
Bauhaus-Universität Weimarhttp://www.uni-weimar.de/
The Web Technology & Information Systems Network
Authors
Matthias Hagen; Martin Potthast; Benno Stein
License

Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically

Description

The Webis Query Segmentation Corpus 2010 (Webis-QSeC-10) contains segmentations for 53,437 web queries obtained from Mechanical Turk crowdsourcing (4,850 used for training in our CIKM 2012 paper). For each query, at least 10 MTurk workers were asked to segment the query. The corpus represents the distribution of their decisions.

Search
Clear search
Close search
Google apps
Main menu