100+ datasets found
  1. Plain Text Wikipedia 2020-11

    • kaggle.com
    zip
    Updated Nov 27, 2020
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    DavidShapiro (2020). Plain Text Wikipedia 2020-11 [Dataset]. https://www.kaggle.com/datasets/ltcmdrdata/plain-text-wikipedia-202011
    Explore at:
    zip(8291577009 bytes)Available download formats
    Dataset updated
    Nov 27, 2020
    Authors
    DavidShapiro
    License

    Attribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
    License information was derived automatically

    Description

    Context

    Wikipedia dumps contain a tremendous amount of markup. WikiMedia Text is a hybrid of markdown and HTML, making it very difficult to use. Wikipedia, however, is an extremely valuable dataset. I wanted a more usable format.

    Content

    This dataset includes ~40MB JSON files, each of which contains a collection of Wikipedia articles. Each article element in the JSON contains only 3 keys: an ID number, the title of the article, and the text of the article. Each article has been "flattened" to occupy a single plain text string. This makes it easier for humans to read, as opposed to the markup version. It also makes it easier for NLP tasks. You will have much less cleanup to do.

    Each file looks like this:

    [
     {
     "id": "17279752",
     "text": "Hawthorne Road was a cricket and football ground in Bootle in England...",
     "title": "Hawthorne Road"
     }
    ]
    

    Acknowledgements

    Absolutely all thanks goes to Wikipedia! Everyone who helped build and fund Wikipedia, and everyone who has contributed their time and expertise to the content of Wikipedia

    Inspiration

    I was inspired to create this dataset because I needed an offline encyclopedia for an Artificial General Intelligence project I'm working on.

  2. h

    simple-wikipedia

    • huggingface.co
    Updated Aug 28, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Rahul (2026). simple-wikipedia [Dataset]. https://huggingface.co/datasets/rahular/simple-wikipedia
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Aug 28, 2026
    Authors
    Rahul
    Description

    simple-wikipedia

    Processed, text-only dump of the Simple Wikipedia (English). Contains 23,886,673 words.

  3. h

    wikipedia-PT

    • huggingface.co
    Updated Apr 5, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Tucano (2025). wikipedia-PT [Dataset]. https://huggingface.co/datasets/TucanoBR/wikipedia-PT
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Apr 5, 2025
    Dataset authored and provided by
    Tucano
    License

    Attribution-ShareAlike 3.0 (CC BY-SA 3.0)https://creativecommons.org/licenses/by-sa/3.0/
    License information was derived automatically

    Description

    Wikipedia-PT

      Dataset Summary
    

    The Portuguese portion of the Wikipedia dataset.

      Supported Tasks and Leaderboards
    

    The dataset is generally used for Language Modeling.

      Languages
    

    Portuguese

      Dataset Structure
    
    
    
    
    
      Data Instances
    

    An example looks as follows: { 'text': 'Abril é o quarto mês...' }

      Data Fields
    

    text (str): Text content of the article.

      Data Splits
    

    All configurations contain a single train split.… See the full description on the dataset page: https://huggingface.co/datasets/TucanoBR/wikipedia-PT.

  4. Wikipedia Talk Corpus

    • figshare.com
    • datasetcatalog.nlm.nih.gov
    • +1more
    application/x-gzip
    Updated Jan 23, 2017
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Ellery Wulczyn; Nithum Thain; Lucas Dixon (2017). Wikipedia Talk Corpus [Dataset]. http://doi.org/10.6084/m9.figshare.4264973.v3
    Explore at:
    application/x-gzipAvailable download formats
    Dataset updated
    Jan 23, 2017
    Dataset provided by
    figshare
    Figsharehttp://figshare.com/
    Authors
    Ellery Wulczyn; Nithum Thain; Lucas Dixon
    License

    CC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
    License information was derived automatically

    Description

    We provide a corpus of discussion comments from English Wikipedia talk pages. Comments are grouped into different files by year. Comments are generated by computing diffs over the full revision history and extracting the content added for each revision. See our wiki for documentation of the schema and our research paper for documentation on the data collection and processing methodology.

  5. h

    wikipedia-ja-20230101

    • huggingface.co
    Updated Jan 9, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Range (2026). wikipedia-ja-20230101 [Dataset]. https://huggingface.co/datasets/range3/wikipedia-ja-20230101
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Jan 9, 2026
    Authors
    Range
    License

    Attribution-ShareAlike 3.0 (CC BY-SA 3.0)https://creativecommons.org/licenses/by-sa/3.0/
    License information was derived automatically

    Description

    range3/wikipedia-ja-20230101

    This dataset consists of a parquet file from the wikipedia dataset with only Japanese data extracted. It is generated by the following python code. このデータセットは、wikipediaデータセットの日本語データのみを抽出したparquetファイルで構成されます。以下のpythonコードによって生成しています。 import datasets dss = datasets.load_dataset( "wikipedia", language="ja", date="20230101", beam_runner="DirectRunner", )

    for split,ds in dss.items(): ds.to_parquet(f"wikipedia-ja-20230101/{split}.parquet")

  6. Wikipedia-dataset

    • kaggle.com
    zip
    Updated Aug 13, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Geetansh Patle (2026). Wikipedia-dataset [Dataset]. https://www.kaggle.com/datasets/geetanshpatle/wikipedia-dataset
    Explore at:
    zip(18976059 bytes)Available download formats
    Dataset updated
    Aug 13, 2026
    Authors
    Geetansh Patle
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    Dataset

    This dataset was created by Geetansh Patle

    Released under MIT

    Contents

  7. wikipedia

    • sophon.at
    Updated May 31, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Wikimedia Foundation (2026). wikipedia [Dataset]. https://sophon.at/evals/wikimedia-wikipedia
    Explore at:
    Dataset updated
    May 31, 2026
    Dataset authored and provided by
    Wikimedia Foundationhttp://www.wikimedia.org/
    License

    Attribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
    License information was derived automatically

    Description

    Dataset Card for Wikimedia Wikipedia

  8. Wiki-talk Datasets

    • zenodo.org
    • data.europa.eu
    gz
    Updated Jan 24, 2020
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Jun Sun; Jérôme Kunegis; Jun Sun; Jérôme Kunegis (2020). Wiki-talk Datasets [Dataset]. http://doi.org/10.5281/zenodo.49561
    Explore at:
    gzAvailable download formats
    Dataset updated
    Jan 24, 2020
    Dataset provided by
    Zenodohttp://zenodo.org/
    Authors
    Jun Sun; Jérôme Kunegis; Jun Sun; Jérôme Kunegis
    License

    Attribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
    License information was derived automatically

    Description

    User interaction networks of Wikipedia of 28 different languages. Nodes (orininal wikipedia user IDs) represent users of the Wikipedia, and an edge from user A to user B denotes that user A wrote a message on the talk page of user B at a certain timestamp.

    More info: http://yfiua.github.io/academic/2016/02/14/wiki-talk-datasets.html

  9. Wikipedia Cultural Diversity Dataset

    • figshare.com
    bz2
    Updated Jun 1, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Marc Miquel-Ribé; David Laniado (2023). Wikipedia Cultural Diversity Dataset [Dataset]. http://doi.org/10.6084/m9.figshare.7039514.v4
    Explore at:
    bz2Available download formats
    Dataset updated
    Jun 1, 2023
    Dataset provided by
    figshare
    Figsharehttp://figshare.com/
    Authors
    Marc Miquel-Ribé; David Laniado
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    For each existing Wikipedia language edition, the dataset contains a classification of the articles that represent its associated cultural context, i.e. all concepts and entities related to the language and to the territories where it is spoken (places, traditions, language, politics, agriculture, biographies, events, etcetera.).For each article, the dataset contains a rich set of context-related features, including geolocation, ISO codes, wikidata properties related to the language or to the corresponding country or territories, as well as related categories, among many other metadata. Other general article features are additional included, such as the number of edits and number of pageviews.The methodology employed to classify articles through machine learning is described in: Wikipedia Cultural Diversity Observatory: Cultural Context Content Methodology Miquel-Ribé, M., & Laniado, D. (2018). Wikipedia Culture Gap: Quantifying Content Imbalances Across 40 Language Editions. Frontiers in Physics.The uses of the dataset are several but we want to highlight three: 1) Wikipedia Culture Gap assessment and overall improvement of the cultural diversity, 2) Academic research in the Digital Humanities field, and 3) User-generated Content based technologies.You can read more at wcdo.wmflabs.org.

  10. m

    legacy-datasets/wikipedia

    • metatext.io
    • mlforge.in
    parquet
    Updated Sep 6, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    legacy-datasets (2026). legacy-datasets/wikipedia [Dataset]. https://metatext.io/datasets/legacy-datasets/wikipedia
    Explore at:
    parquetAvailable download formats
    Dataset updated
    Sep 6, 2026
    Dataset authored and provided by
    legacy-datasets
    Variables measured
    Fill Mask, Text Generation
    Description

    Wikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wikipedia dump (https://dumps.wikimedia.org/) with one split per language. Each example contains the content of one full Wikipedia article with cleanin...

  11. h

    ner-wikipedia-dataset

    • huggingface.co
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Stockmark Inc., ner-wikipedia-dataset [Dataset]. https://huggingface.co/datasets/stockmark/ner-wikipedia-dataset
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset authored and provided by
    Stockmark Inc.
    License

    Attribution-ShareAlike 3.0 (CC BY-SA 3.0)https://creativecommons.org/licenses/by-sa/3.0/
    License information was derived automatically

    Description

    Wikipediaを用いた日本語の固有表現抽出データセット

    GitHub: https://github.com/stockmarkteam/ner-wikipedia-dataset/ LICENSE: CC-BY-SA 3.0

    Developed by Stockmark Inc.

  12. 23 Million Wikipedia Topics and Categories

    • kaggle.com
    zip
    Updated Jul 6, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    NikBearBrown (2023). 23 Million Wikipedia Topics and Categories [Dataset]. https://www.kaggle.com/datasets/nikbearbrown/23-million-wikipedia-topics-and-categories
    Explore at:
    zip(611852937 bytes)Available download formats
    Dataset updated
    Jul 6, 2023
    Authors
    NikBearBrown
    Description

    Wikipedia page titles are great topic tags and hashtags for NLP. Wikipedia has the added benefit of annotating some of the title in categories.

    For example, if you want to create a tag list of film stars there is no better source than Wikipedia.

    For example the tag/title Anarchism has a lot of meta information associated with it.

    Anarchism ['Anarchism| ', 'Anti-capitalism', 'Anti-fascism', 'Economic ideologies', 'Far-left politics', 'Left-wing politics', 'Libertarian socialism', 'Libertarianism', 'Political culture', 'Political ideologies', 'Political movements', 'Social theories', 'Socialism']

  13. Wikimedia editor activity (monthly)

    • figshare.com
    bz2
    Updated Dec 17, 2019
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Aaron Halfaker (2019). Wikimedia editor activity (monthly) [Dataset]. http://doi.org/10.6084/m9.figshare.1553296.v1
    Explore at:
    bz2Available download formats
    Dataset updated
    Dec 17, 2019
    Dataset provided by
    figshare
    Figsharehttp://figshare.com/
    Authors
    Aaron Halfaker
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    This dataset contains a row for every (wiki, user, month) that contains a count of all 'revisions' saved and a count of those revisions that were 'archived' when the page was deleted. For more information, see https://meta.wikimedia.org/wiki/Research:Monthly_wikimedia_editor_activity_dataset Fields: · wiki -- The dbname of the wiki in question ("enwiki" == English Wikipedia, "commonswiki" == Commons) · month -- YYYYMM · user_id -- The user's identifier in the local wiki · user_name -- The user name in the local wiki (from the 'user' table) · user_registration -- The recorded registration date for the user in the 'user' table · archived -- The count of deleted revisions saved in this month by this user · revisions -- The count of all revisions saved in this month by this user (archived or not) · attached_method -- The method by which this user attached this account to their global account

  14. wikimedia/wikisource

    • metatext.io
    parquet
    Updated Sep 17, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    wikimedia (2026). wikimedia/wikisource [Dataset]. https://metatext.io/datasets/wikimedia/wikisource
    Explore at:
    parquetAvailable download formats
    Dataset updated
    Sep 17, 2026
    Dataset provided by
    Wikimedia Foundationhttp://www.wikimedia.org/
    Authors
    wikimedia
    Variables measured
    Fill Mask, Text Generation
    Description

    Dataset Card for Wikimedia Wikisource

    Dataset Summary

    Wikisource dataset containing cleaned articles of all languages. The dataset is built from the Wikisource dumps (https://dumps.wikimedia.org/) with one subset per language, each c...

  15. m

    olm/wikipedia

    • metatext.io
    parquet
    Updated Sep 5, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    olm (2026). olm/wikipedia [Dataset]. https://metatext.io/datasets/olm/wikipedia
    Explore at:
    parquetAvailable download formats
    Dataset updated
    Sep 5, 2026
    Dataset authored and provided by
    olm
    Variables measured
    Fill Mask, Text Generation
    Description

    Wikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wikipedia dump (https://dumps.wikimedia.org/) with one split per language. Each example contains the content of one full Wikipedia article with cleanin...

  16. a

    Wikidata PageRank

    • danker.s3.amazonaws.com
    application/vnd.hdt +3
    Updated Sep 10, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Andreas Thalhammer (2026). Wikidata PageRank [Dataset]. https://danker.s3.amazonaws.com/index.html
    Explore at:
    tsv, ttl, application/vnd.hdt, ntAvailable download formats
    Dataset updated
    Sep 10, 2026
    Authors
    Andreas Thalhammer
    License

    Attribution-ShareAlike 3.0 (CC BY-SA 3.0)https://creativecommons.org/licenses/by-sa/3.0/
    License information was derived automatically

    Description

    Regularly published dataset of PageRank scores for Wikidata entities. The underlying link graph is formed by a union of all links accross all Wikipedia language editions. Computation is performed by Andreas Thalhammer with 'danker' available at https://github.com/athalhammer/danker .

  17. m

    graelo/wikipedia

    • metatext.io
    parquet
    Updated Sep 17, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    graelo (2026). graelo/wikipedia [Dataset]. https://metatext.io/datasets/graelo/wikipedia
    Explore at:
    parquetAvailable download formats
    Dataset updated
    Sep 17, 2026
    Dataset authored and provided by
    graelo
    License

    Attribution-ShareAlike 3.0 (CC BY-SA 3.0)https://creativecommons.org/licenses/by-sa/3.0/
    License information was derived automatically

    Variables measured
    Fill Mask, Text Generation
    Description

    Wikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wikipedia dump (https://dumps.wikimedia.org/) with one split per language. Each example contains the content of one full Wikipedia article with cleanin...

  18. Wikidata Reference

    • figshare.com
    gz
    Updated Mar 17, 2025
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Sven Hertling; Nandana Mihindukulasooriya (2025). Wikidata Reference [Dataset]. http://doi.org/10.6084/m9.figshare.28602170.v2
    Explore at:
    gzAvailable download formats
    Dataset updated
    Mar 17, 2025
    Dataset provided by
    figshare
    Figsharehttp://figshare.com/
    Authors
    Sven Hertling; Nandana Mihindukulasooriya
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    Dataset SummaryThe Triple-to-Text Alignment dataset aligns Knowledge Graph (KG) triples from Wikidata with diverse, real-world textual sources extracted from the web. Unlike previous datasets that rely primarily on Wikipedia text, this dataset provides a broader range of writing styles, tones, and structures by leveraging Wikidata references from various sources such as news articles, government reports, and scientific literature. Large language models (LLMs) were used to extract and validate text spans corresponding to KG triples, ensuring high-quality alignments. The dataset can be used for training and evaluating relation extraction (RE) and knowledge graph construction systems.Data FieldsEach row in the dataset consists of the following fields:subject (str): The subject entity of the knowledge graph triple.rel (str): The relation that connects the subject and object.object (str): The object entity of the knowledge graph triple.text (str): A natural language sentence that entails the given triple.validation (str): LLM-based validation results, including:Fluent Sentence(s): TRUE/FALSESubject mentioned in Text: TRUE/FALSERelation mentioned in Text: TRUE/FALSEObject mentioned in Text: TRUE/FALSEFact Entailed By Text: TRUE/FALSEFinal Answer: TRUE/FALSEreference_url (str): URL of the web source from which the text was extracted.subj_qid (str): Wikidata QID for the subject entity.rel_id (str): Wikidata Property ID for the relation.obj_qid (str): Wikidata QID for the object entity.Dataset CreationThe dataset was created through the following process:1. Triple-Reference Sampling and ExtractionAll relations from Wikidata were extracted using SPARQL queries.A sample of KG triples with associated reference URLs was collected for each relation.2. Domain Analysis and Web ScrapingURLs were grouped by domain, and sampled pages were analyzed to determine their primary language.English-language web pages were scraped and processed to extract plaintext content.3. LLM-Based Text Span Selection and ValidationLLMs were used to identify text spans from web content that correspond to KG triples.A Chain-of-Thought (CoT) prompting method was applied to validate whether the extracted text entailed the triple.The validation process included checking for fluency, subject mention, relation mention, object mention, and final entailment.4. Final Dataset Statistics12.5K Wikidata relations were analyzed, leading to 3.3M triple-reference pairs.After filtering for English content, 458K triple-web content pairs were processed with LLMs.80.5K validated triple-text alignments were included in the final dataset.

  19. Plaintext Wikipedia (full English)

    • kaggle.com
    zip
    Updated May 19, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Ffatty (2024). Plaintext Wikipedia (full English) [Dataset]. https://www.kaggle.com/datasets/ffatty/plaintext-wikipedia-full-english
    Explore at:
    zip(4562950901 bytes)Available download formats
    Dataset updated
    May 19, 2024
    Authors
    Ffatty
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    Full English Wikipedia (plain text)

    Unsupervised text corpus of all 4,366,845 articles on the English Wikipedia.

    • ? M tokens
    • ? unique words
    • 22GB (uncompressed)

    Extracted from the Wikipedia dumps using an open source tool, using an AWS EC2 (cloud VM).

    Data is remarkably clean and uniform.

    There is also a dataset available for the Simple English Wikipedia, generated in the same manner as this dataset. It is much more compact at just 120MB, and easier to work with: https://www.kaggle.com/datasets/ffatty/plain-text-wikipedia-simpleenglish/

    Format:

    • Each article's title appears before the content.
    • Articles are plain text; they are stripped of all Wiki formatting syntax, including font styles, citations, links, etc.
    • Articles are concatenated into .txt files of ≤ 1MB each.


    Random example excerpt:

    (a portion of 1 file; 2 articles, concatenated in place)

    Peter Burroughs
    
    Peter Burroughs (born 27 January 1947) is a British television and film actor and the director of Willow Management.
    
    Burroughs initially ran a shop in his village at Yaxley, Cambridgeshire.
    
    His first dramatic role was that of the character ""Branic"" in the 1979 television series "The Legend of King Arthur". He played roles in Hollywood movies such as "Flash Gordon", "The Dark Crystal", "Labyrinth", and "The Hitchhiker's Guide to the Galaxy". He portrayed a bank goblin in the Harry Potter series ("Harry Potter and the Philosopher's Stone" and "Harry Potter and the Deathly Hallows – Part 2").
    
    
    Eddi McKee
    
    Eddi McKee is a fictional character from the BBC medical drama "Holby City", played by actress Sarah-Jane Potts. She first appeared in the thirteenth series episode "Rescue Me", broadcast on 7 June 2011. Eddi was a Senior Nurse at Holby City Hospital. 
    
    Eddi is portrayed as being straight-talking, no-nonsense, loyal and compassionate. Eddi believed the nurses did the most important job in the hospital and she had very forthright opinions. Exploration of the character's backstory began when her younger brother, Liam (Luke Tittensor), was introduced and it was revealed that Eddi had run away from home to escape their alcoholic mother. Potts pointed out that Eddi herself liked to drink to deal with her many issues. The character had a relationship with Luc Hemingway (Joseph Millson), which proved popular with viewers. Eddi also became involved with locum consultant Max Schneider (John Light) and developed an addiction to painkillers.
    
    


  20. h

    persian-wikipedia

    • huggingface.co
    Updated Nov 18, 2024
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    MaralGPT (2024). persian-wikipedia [Dataset]. https://huggingface.co/datasets/MaralGPT/persian-wikipedia
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Nov 18, 2024
    Dataset authored and provided by
    MaralGPT
    Description

    MaralGPT/persian-wikipedia dataset hosted on Hugging Face and contributed by the HF Datasets community

Share
FacebookFacebook
TwitterTwitter
Email
Click to copy link
Link copied
Close
Cite
DavidShapiro (2020). Plain Text Wikipedia 2020-11 [Dataset]. https://www.kaggle.com/datasets/ltcmdrdata/plain-text-wikipedia-202011
Organization logo

Plain Text Wikipedia 2020-11

De-marked up Wikipedia for offline use

Explore at:
6 scholarly articles cite this dataset (View in Google Scholar)
zip(8291577009 bytes)Available download formats
Dataset updated
Nov 27, 2020
Authors
DavidShapiro
License

Attribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically

Description

Context

Wikipedia dumps contain a tremendous amount of markup. WikiMedia Text is a hybrid of markdown and HTML, making it very difficult to use. Wikipedia, however, is an extremely valuable dataset. I wanted a more usable format.

Content

This dataset includes ~40MB JSON files, each of which contains a collection of Wikipedia articles. Each article element in the JSON contains only 3 keys: an ID number, the title of the article, and the text of the article. Each article has been "flattened" to occupy a single plain text string. This makes it easier for humans to read, as opposed to the markup version. It also makes it easier for NLP tasks. You will have much less cleanup to do.

Each file looks like this:

[
 {
 "id": "17279752",
 "text": "Hawthorne Road was a cricket and football ground in Bootle in England...",
 "title": "Hawthorne Road"
 }
]

Acknowledgements

Absolutely all thanks goes to Wikipedia! Everyone who helped build and fund Wikipedia, and everyone who has contributed their time and expertise to the content of Wikipedia

Inspiration

I was inspired to create this dataset because I needed an offline encyclopedia for an Artificial General Intelligence project I'm working on.

Search
Clear search
Close search
Google apps
Main menu