2 datasets found
  1. W

    Webis-QSeC-10

    • webis.de
    3256198
    Updated 2010
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Matthias Hagen; Martin Potthast; Benno Stein (2010). Webis-QSeC-10 [Dataset]. http://doi.org/10.5281/zenodo.3256198
    Explore at:
    3256198Available download formats
    Dataset updated
    2010
    Dataset provided by
    The Web Technology & Information Systems Network
    Friedrich Schiller University Jena
    University of Kassel, hessian.AI, and ScaDS.AI
    Bauhaus-Universit?t Weimar
    Authors
    Matthias Hagen; Martin Potthast; Benno Stein
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    The Webis Query Segmentation Corpus 2010 (Webis-QSeC-10) contains segmentations for 53,437 web queries obtained from Mechanical Turk crowdsourcing (4,850 used for training in our CIKM 2012 paper). For each query, at least 10 MTurk workers were asked to segment the query. The corpus represents the distribution of their decisions.

  2. E

    Webis Query Spelling Corpus 2017 (Webis-QSpell-17)

    • live.european-language-grid.eu
    • nde-dev.biothings.io
    • +2more
    csv
    Updated May 26, 2024
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    (2024). Webis Query Spelling Corpus 2017 (Webis-QSpell-17) [Dataset]. https://live.european-language-grid.eu/catalogue/corpus/7757
    Explore at:
    csvAvailable download formats
    Dataset updated
    May 26, 2024
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    The Webis Query Spelling Corpus 2017 (Webis-QSpell-17) contains 54,772 web queries that were manually spell-checked; for 9,171 queries alternative spelling variants are contained.As for segmentations of many of the queries (i.e., tagged concepts and phrases), please refer to the companion corpus Webis-QSeC-10.

  3. Not seeing a result you expected?
    Learn how you can add new datasets to our index.

Share
FacebookFacebook
TwitterTwitter
Email
Click to copy link
Link copied
Close
Cite
Matthias Hagen; Martin Potthast; Benno Stein (2010). Webis-QSeC-10 [Dataset]. http://doi.org/10.5281/zenodo.3256198

Webis-QSeC-10

Explore at:
9 scholarly articles cite this dataset (View in Google Scholar)
3256198Available download formats
Dataset updated
2010
Dataset provided by
The Web Technology & Information Systems Network
Friedrich Schiller University Jena
University of Kassel, hessian.AI, and ScaDS.AI
Bauhaus-Universit?t Weimar
Authors
Matthias Hagen; Martin Potthast; Benno Stein
License

Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically

Description

The Webis Query Segmentation Corpus 2010 (Webis-QSeC-10) contains segmentations for 53,437 web queries obtained from Mechanical Turk crowdsourcing (4,850 used for training in our CIKM 2012 paper). For each query, at least 10 MTurk workers were asked to segment the query. The corpus represents the distribution of their decisions.

Search
Clear search
Close search
Google apps
Main menu