Facebook
TwitterWikidata is a free, CC0-licensed knowledge base operated by the Wikimedia Foundation. All public APIs (Wikibase REST, MediaWiki Action, SPARQL Query Service, Linked Data interface, Recent Changes event stream) are free of charge with no paid tiers. Heavy or commercial users are expected to follow the User-Agent and maxlag policies, mirror data via the database dumps, or contract with Wikimedia Enterprise for commercial-grade access.
Facebook
TwitterAttribution-ShareAlike 3.0 (CC BY-SA 3.0)https://creativecommons.org/licenses/by-sa/3.0/
License information was derived automatically
Regularly published dataset of PageRank scores for Wikidata entities. The underlying link graph is formed by a union of all links accross all Wikipedia language editions. Computation is performed by Andreas Thalhammer with 'danker' available at https://github.com/athalhammer/danker .
Facebook
TwitterAttribution-ShareAlike 3.0 (CC BY-SA 3.0)https://creativecommons.org/licenses/by-sa/3.0/
License information was derived automatically
Wikipedia, the free encyclopedia, and Wikidata, the free knowledge base, are crowd-sourced projects supported by the Wikimedia Foundation. Wikipedia is nearly 20 years old and recently added its six millionth article in English. Wikidata, its younger machine-readable sister project, was created in 2012 but has been growing rapidly and currently contains more than 75 million items.
These projects contribute to the Wikimedia Foundation's mission of empowering people to develop and disseminate educational content under a free license. They are also heavily utilized by computer science research groups, especially those interested in natural language processing (NLP). The Wikimedia Foundation periodically releases snapshots of the raw data backing these projects, but these are in a variety of formats and were not designed for use in NLP research. In the Kensho R&D group, we spend a lot of time downloading, parsing, and experimenting with this raw data. The Kensho Derived Wikimedia Dataset (KDWD) is a condensed subset of the raw Wikimedia data in a form that we find helpful for NLP work. The KDWD has a CC BY-SA 3.0 license, so feel free to use it in your work too.
https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4301984%2F972e4157b97efe8c2c5ea17c983b1504%2Fkdwd_header_logos_2.jpg?generation=1580510520532141&alt=media" alt="">
This particular release consists of two main components - a link annotated corpus of English Wikipedia pages and a compact sample of the Wikidata knowledge base. We version the KDWD using the raw Wikimedia snapshot dates. The version string for this dataset is kdwd_enwiki_20191201_wikidata_20191202 indicating that this KDWD was built from the English Wikipedia snapshot from 2019 December 1 and the Wikidata snapshot from 2019 December 2. Below we describe these components in more detail.
Dive right in by checking out some of our example notebooks:
page.csv (page metadata and Wikipedia-to-Wikidata mapping)link_annotated_text.jsonl (plaintext of Wikipedia pages with link offsets)item.csv (item labels and descriptions in English)item_aliases.csv (item aliases in English)property.csv (property labels and descriptions in English)property_aliases.csv (property aliases in English)statements.csv (truthy qpq statements)The KDWD is three connected layers of data. The base layer is a plain text English Wikipedia corpus, the middle layer annotates the corpus by indicating which text spans are links, and the top layer connects the link text spans to items in Wikidata. Below we'll describe these layers in more detail.
https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4301984%2F19663d43bade0e92f578255f6e0d9dcd%2Fkensho_wiki_triple_layer.svg?generation=1580347573004185&alt=media" alt="">
The first part of the KDWD is derived from Wikipedia. In order to create a corpus of mostly natural text, we restrict our English Wikipedia page sample to those that:
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
A copy of a dump which was available from WikiMedia: https://dumps.wikimedia.org/wikidatawiki/entities/
Facebook
Twitterhttps://choosealicense.com/licenses/cc0-1.0/https://choosealicense.com/licenses/cc0-1.0/
Wikidata Multilingual Label Maps 2025
Comprehensive multilingual label and description maps extracted from the 2025 Wikidata dump.This dataset contains labels and descriptions for Wikidata entities (Q-items and P-properties) across 613 languages.
Dataset Overview
📊 Total Records: 725,274,530 label/description pairs 🆔 Unique Entities: 117,229,348 (Q-items and P-properties) 🌍 Languages: 613 unique language codes 📝 With Descriptions: 339,691,043 pairs (46.8%… See the full description on the dataset page: https://huggingface.co/datasets/yashkumaratri/wikidata-label-maps-2025-all-languages.
Facebook
TwitterThe free knowledge base anyone can edit https://wikidata.org
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
Wikidata is an open-source knowledge base that serves as a repository for structured data used in various Wikimedia projects like Wikipedia and Wikivoyage. Much like other knowledge graphs, Wikidata organizes information into triples, which consist of a subject item, a property, and an object. We extract a specific set of triples from Wikidata based on certain criteria and finally obtain the carefully curated dataset. Currently, we are presenting 5 geographic datasets - Argentina, Australia, India, South Africa and Russia.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This dataset contains information about commercial organizations (companies) and their relations with other commercial organizations, persons, products, locations, groups and industries. The dataset has the form of a graph. It has been produced by the SmartDataLake project (https://smartdatalake.eu), using data collected from Wikidata (https://www.wikidata.org).
Facebook
Twitterhttps://choosealicense.com/licenses/cc0-1.0/https://choosealicense.com/licenses/cc0-1.0/
Wikidata parallel descriptions en-ja
Parallel corpus for machine translation generated from wikidata dump (2024-05-06). Currently we processed only English/Japanese pair. The jsonl file is ready-to-train by Hugging Face transformers trainer for translation tasks.
Dataset Details
https://www.wikidata.org/wiki/Wikidata:Database_download
Dataset Creation
As Wikidata description field does not represent exact direct translation, filtering is required for… See the full description on the dataset page: https://huggingface.co/datasets/Mitsua/wikidata-parallel-descriptions-en-ja.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset was created by JonathanFraine
Released under CC0: Public Domain
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
A wide array of active, passive, and potential end users has benefited from the gallery, library, archive, and museum (GLAM) communities tradition of acquiring, preserving, and providing access to a growing number of digital or digitized curated collections. For the custodians of the rare and unique material types there is greater demand for better tools to share data between systems, manage workflows, track metrics, assess user needs, while increasing visibility of these collections simultaneously while managing potential community concerns regarding privacy, sensitive materials, or sacred places. As the archives, special collections, and rare books library communities investigate opportunities to structure collection data as linked data, there are still questions and tensions balancing increasing access and discovery with ethical creation of linked data. One approach which is gaining momentum in the GLAM community is creating, consuming, and integrating Wikidata entities to enhance, augment, and link to digital collections and other web based applications. Wikidata is a multilingual crowdsourced structured data knowledge-base built upon Wikibase and MediaWiki technologies (Vrandečić, & Krötzsch, 2014). As cultural heritage professionals develop projects in Wikidata, pilot projects can reveal liminal spaces where ethical deliberation is both nuanced and necessary. In this chapter we will draw upon specific local use cases that illustrate these ethical questions and connect them to practical strategies that may be employed. Community feedback (from communities represented as well as the Wikidata community and the Linked Data communities) and a commitment to iterative local practice are two specific components that will be discussed. In addition, ethical principles from other big data projects and the movement for more inclusive archival and cultural heritage practices can provide translatable frameworks and accountability for this work. The chapter will conclude by connecting the current landscape to future challenges and opportunities with the goal of igniting interest and broader engagement as we work together to enact, empower, and reflect upon responsive approaches to ethical challenges.
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
RDF dump of wikidata produced with wdumper.
entity count: 0, statement count: 0, triple count: 0
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
Wikidata Descriptions Dataset
wikidata_descriptions pairs English Wikipedia article titles (wiki_title) and their Wikidata IDs (qid) with the English "description" available in Wikidata.The corpus contains 26 205 entities. Wikidata descriptions are short, one-line summaries that concisely state what an entity is.They can be used as lightweight contextual information in entity linking, search, question answering, knowledge-graph completion and many other NLP / IR tasks.… See the full description on the dataset page: https://huggingface.co/datasets/masaki-sakata/wikidata_descriptions.
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
RDF dump of wikidata produced with wdumps.
<p>
<br>
<a href="https://tools.wmflabs.org/wdumps/dump/669">View on wdumper</a>
</p>
<p>
<b>entity count</b>: 0, <b>statement count</b>: 0, <b>triple count</b>: 0
</p>
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
A Genism LDA Topic Model of English Wikipedia biographical articles with list of all 1.8M articles, and some associated Wikidata information The model has 150 Topics. This model was developed in the process of isolating a set of visual arts biographical articles, as described in "Clowns in the Visual Artists: Topic Modeling Wikipedia and Wikidata" in the Spring 2022 issue of Art Documentation - https://doi.org/10.1086/719999 Because names, nationalities, and birthdays are so prominent in biographies, the stopwords list removed 170,000 names, surnames, city names, place names, countries, days, months and other time related words (https://github.com/mandiberg/Names-Surnames-and-Countries-for-Stopwords). We also directly removed each article subject’s given and surname, which were almost always the most frequently occurring words in any given article. Otherwise, the model just produced topics based on nationality, and common names and surnames. Files: all_enwiki_bios_from_wikidata.csv The list of all Wikidata items for humans with an enwiki page (e.g biographical article) was extracted from Wikidata JSON dump; list includes gender, occupation, and nationality. This was joined with the converted plaintext from an English Wikipedia dump. This data was downloaded in March 2021. Wikipedia Biographies LDA Topic Model human readable summary.csv A human readable file with the 150 topics ranked by count of articles per topic from the 1.8M corpus. The most popular topics have categorical descriptions of the occupations of each cluster. Some are marked as not an occupation cluster. BoW_corpus.mm* model_lda_full_Sep2_150Tv2* These six files comprise the topic model. The code to load them is present in the python files. dict_full_Aug-28-2021 processed_docs_full_Aug-28-2021.txt processed_docs_1000_Aug-18-2021.txt These are the dictionary and processed corpuses required to build and implement the model using this code. The corpus with the first 1000 items is meant to be used for testing, as the full one is quite large and takes a long time to complete. topic-model-wikipedia-sept2021.zip The code and settings used for creating and implementing this model are included in this zip and are also available here: https://github.com/mandiberg/topic-model-wikipedia All-Wikipedia-Biographies-with-topic1.csv All-Wikipedia-Biographies-with-topic1and2.csv These are the list of 1.8M biographies matched to topics. The "topic1" file just includes the first topic, this is a slightly larger list. The "topic1and2" file is slightly smaller because about 2% articles do not match to a second topic. Analysis-for-Clowns-Visual-Arts.zip These are the raw data and final data produced for the "Clowns in the Visual Artists." Please see the article for context.
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
Structured factual statements about a named entity from Wikidata — the collaboratively-curated knowledge base behind Wikipedia. Resolves entity to its Wikidata item, then returns the headline facts: description, instance/class, key properties (for a country: capital, population, area, head of government, currency, ISO codes; for a person: birth/death, occupation, citizenship), cross-reference identifiers and the Wikidata QID. entity = a name to look up (Australia, Albert Einstein, Sydney Opera House, Bitcoin); default Australia. Reference/knowledge lookup, not live telemetry — values are as-current-as Wikidata's community edits. Source: Wikidata (wikidata.org), MediaWiki Action API — CC0 1.0 public domain: commercial use + redistribution permitted, no key, no attribution required.
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
This dataset was created by Jamesley Joseph
Released under MIT
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
The dataset corresponds to a preprocessed dump of wikidata, where all identifiers were mapped to a contiguous alphabet
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
The Wikimedia Foundation's open knowledge graph. Over 100 million entities with unique identifiers (Q-IDs). The most verifiable source for semantic identity.
Facebook
TwitterFiles in this dataset have been produced during Performance and Accuracy experiments of Wikidata subsetting practical tools: WDumper, KGTK, WDSub, WDF.
Facebook
TwitterWikidata is a free, CC0-licensed knowledge base operated by the Wikimedia Foundation. All public APIs (Wikibase REST, MediaWiki Action, SPARQL Query Service, Linked Data interface, Recent Changes event stream) are free of charge with no paid tiers. Heavy or commercial users are expected to follow the User-Agent and maxlag policies, mirror data via the database dumps, or contract with Wikimedia Enterprise for commercial-grade access.