Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
Wikidata is a collaboratively edited knowledge base operated by the Wikimedia Foundation. It is intended to provide a common source of certain types of data which can be used by Wikimedia projects such as Wikipedia. Wikidata functions as a document-oriented database, centred on individual items. Items represent topics, for which basic information is stored that identifies each topic.
Facebook
TwitterAttribution-NonCommercial 4.0 (CC BY-NC 4.0)https://creativecommons.org/licenses/by-nc/4.0/
License information was derived automatically
Wikidata Thematic Subgraph Selection
These datasets have been designed to train and evaluate algorithms to select thematic subgraphs of interest in a large knowledge graph from seed entities of interest. Specifically, we consider Wikidata. Given a set of seed QIDs of interest, a graph expansion is performed following P31, P279, and (-)P279 edges. Traversed classes that thematically deviates from seed QIDs of interest should be pruned. Datasets thus consist of classes reached from seed QIDs that are labeled as "to prune" or "to keep".
Available datasets
| Dataset | # Seed QIDs | # Labeled decisions | # Prune decisions | Min prune depth | Max prune depth | # Keep decisions | Min keep depth | Max keep depth | # Reached nodes up | # Reached nodes down |
|---|---|---|---|---|---|---|---|---|---|---|
| dataset1 | 455 | 5233 | 3464 | 1 | 4 | 1769 | 1 | 4 | 1507 | 2593609 |
| dataset2 | 105 | 982 | 388 | 1 | 2 | 594 | 1 | 3 | 1159 | 1247385 |
Each dataset folder contains
datasetX.csv: a CSV file containing one seed QID per line (not the complete URL, just the QID). This CSV file has no header.datasetX_labels.csv: a CSV file containing one seed QID per line and its label (not the complete URL, just the QID)datasetX_gold_decisions.csv: a CSV file with seed QIDs, reached QIDs, and the labeled decision (1: keep, 0: prune)datasetX_Y_folds.pkl: folds to train and test models based on the labeled decisionsdataset1-2 consists of using dataset1 for training and dataset2 for testing.
License
Datasets are available under the CC BY-NC license.
Facebook
Twitterhttps://choosealicense.com/licenses/cc0-1.0/https://choosealicense.com/licenses/cc0-1.0/
Wikidata Entity Embeddings 0.2
Dataset Summary
Wikidata Entity Embeddings is a dataset of embedding vectors for Wikidata entities. Each vector represents a Wikidata item (Q...) or property (P...) based on textual information extracted from Wikidata. The dataset is part of the Wikidata Embedding Project, an initiative led by Wikimedia Deutschland in collaboration with Jina AI and IBM DataStax. The project provides a publicly accessible Wikidata Vector Database to⊠See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/Wikidata_Vectors_0.2.
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
Dataset SummaryThe Triple-to-Text Alignment dataset aligns Knowledge Graph (KG) triples from Wikidata with diverse, real-world textual sources extracted from the web. Unlike previous datasets that rely primarily on Wikipedia text, this dataset provides a broader range of writing styles, tones, and structures by leveraging Wikidata references from various sources such as news articles, government reports, and scientific literature. Large language models (LLMs) were used to extract and validate text spans corresponding to KG triples, ensuring high-quality alignments. The dataset can be used for training and evaluating relation extraction (RE) and knowledge graph construction systems.Data FieldsEach row in the dataset consists of the following fields:subject (str): The subject entity of the knowledge graph triple.rel (str): The relation that connects the subject and object.object (str): The object entity of the knowledge graph triple.text (str): A natural language sentence that entails the given triple.validation (str): LLM-based validation results, including:Fluent Sentence(s): TRUE/FALSESubject mentioned in Text: TRUE/FALSERelation mentioned in Text: TRUE/FALSEObject mentioned in Text: TRUE/FALSEFact Entailed By Text: TRUE/FALSEFinal Answer: TRUE/FALSEreference_url (str): URL of the web source from which the text was extracted.subj_qid (str): Wikidata QID for the subject entity.rel_id (str): Wikidata Property ID for the relation.obj_qid (str): Wikidata QID for the object entity.Dataset CreationThe dataset was created through the following process:1. Triple-Reference Sampling and ExtractionAll relations from Wikidata were extracted using SPARQL queries.A sample of KG triples with associated reference URLs was collected for each relation.2. Domain Analysis and Web ScrapingURLs were grouped by domain, and sampled pages were analyzed to determine their primary language.English-language web pages were scraped and processed to extract plaintext content.3. LLM-Based Text Span Selection and ValidationLLMs were used to identify text spans from web content that correspond to KG triples.A Chain-of-Thought (CoT) prompting method was applied to validate whether the extracted text entailed the triple.The validation process included checking for fluency, subject mention, relation mention, object mention, and final entailment.4. Final Dataset Statistics12.5K Wikidata relations were analyzed, leading to 3.3M triple-reference pairs.After filtering for English content, 458K triple-web content pairs were processed with LLMs.80.5K validated triple-text alignments were included in the final dataset.
Facebook
TwitterWikidata is a free, CC0-licensed knowledge base operated by the Wikimedia Foundation. All public APIs (Wikibase REST, MediaWiki Action, SPARQL Query Service, Linked Data interface, Recent Changes event stream) are free of charge with no paid tiers. Heavy or commercial users are expected to follow the User-Agent and maxlag policies, mirror data via the database dumps, or contract with Wikimedia Enterprise for commercial-grade access.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
WikiSnap25 is a collection of analysisâready tables that integrate content, metadata, and readership information for English Wikipedia articles and Wikidata entities as of midâ2025. The goal is to lower the engineering barrier for empirical Wikimedia research by abstracting away the work of parsing XML/JSON dumps and aggregating pageview data into compact, wellâdocumented Parquet datasets.
WikiSnap25 combines several official Wikimedia sources: the June 1, 2025 English Wikipedia pagesâarticlesâmultistream XML dump for article text and metadata, the June 2, 2025 full Wikidata JSON dump for entityâlevel information, a decade of Wikimedia REST API monthly pageview data from July 2015 through June 2025, and additional dumps (e.g., stub metaâhistory and redirect mappings) for revision context and canonicalization. A modular processing pipeline (XML parsing, wikitext processing, mwparserfromhell, and aggregation scripts) produces a consistent articleâlevel snapshot anchored at midâ2025, enriched with longâterm readership and knowledgeâgraph features.
The following datasets are included:
WikiSnap_2025_WP_Articles.parquet â Articleâlevel data for 7,011,415 English Wikipedia articles, including normalized titles, total pageviews (2015â2025), first two sentences of article text, outlink metrics (e.g., num_articles_linked_out), and mappings to Wikidata QIDs. This table supports studies of article popularity, connectivity, and entity types.
WikiSnap_2025_WP_Edges.parquet â 314,337,621 intraâWikipedia hyperlinks between articles, with from/to article IDs and titles, endpoint pageviews, and a duplicateâlink flag for graph cleaning. This edge list is suitable for largeâscale network analyses of the article link graph.
WikiSnap_2025_WP_Network_Metrics.parquet â Network metrics for articles with intra-links to other Wikipedia articles (7,008,375 rows), including centrality scores, PageRank, trafficâweighted PageRank, HITS hub/authority, degrees, and Leiden community assignments computed from a deduplicated article link graph.
WikiSnap_2025_WP_Article_Metrics.parquet â Additional articleâlevel metrics (7,011,415 rows) such as creation timestamp, edit lifespan in days, total revisions, number of unique editors, and a Giniâstyle editor inequality index across redirect clusters.
WikiSnap_2025_WD_Entities.parquet â Wikidata entity information for 116,183,072 items, including labels, instanceâof labels, temporal attributes (begin/end year), country of citizenship labels, and sitelinks to English Wikipedia article titles, with disambiguation pages filtered and truthy claims prioritized.
WikiSnap_2025_WD_Metrics.parquet â Wikidata metrics for the same 116M entities, including counts of claims, in/out links, references, qualifiers, notable properties, articleâquality badges, and external identifiers, enabling notability and completeness assessments.
WikiSnap_2025_WD_Humans.parquet â A subset for human entities (P31=Q5) with 12,371,744 rows, containing birth and death years, occupation labels, gender, citizenship, and notability metrics, tailored to demographic and biographical analysis.
These datasets are intended to support a wide range of tasks, including correlating network centrality with pageview patterns, assessing Wikidata notability via property richness and sitelinks, and exploring editorial activity and inequality, without requiring researchers to build their own largeâscale Wikimedia processing pipelines from scratch.
Limitations. The current release focuses on English Wikipedia mainânamespace articles as of June 1, 2025; Wikidata content as of early June 2025; and articleâlevel pageviews aggregated over July 2015âJune 2025 without finer temporal granularity. Redirects, stubs, and full revision histories are not included in these tables.
All datasets and the full processing code are released to facilitate reproducible Wikimedia research and to enable crossâstudy comparability on a shared, 2025âanchored snapshot.
License and citation. All WikiSnap25 datasets and associated processing code archived in this record are dedicated to the public domain under the Creative Commons Zero v1.0 Universal (CC0 1.0) public domain dedication. Users are free to reuse, modify, and redistribute the materials without restriction.
Facebook
TwitterAttribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically
Wikidata Title and Description
Wikidata entity titles and short descriptions for all 324 Wikipedia language editions, extracted from the Wikimedia Analytics wmf.wikidata_entity table. Each row links a Wikidata QID to the title and description as they appear in a specific language edition of Wikipedia.
Dataset Details
Field Value
Source wmf.wikidata_entity (Wikidata + Wikipedia sitelinks)
Snapshot 2026-03-30
Languages 324
Total titles 88,292,409
Titles⊠See the full description on the dataset page: https://huggingface.co/datasets/wikimedia/wikidata-title-desc.
Facebook
TwitterAttribution-ShareAlike 3.0 (CC BY-SA 3.0)https://creativecommons.org/licenses/by-sa/3.0/
License information was derived automatically
Regularly published dataset of PageRank scores for Wikidata entities. The underlying link graph is formed by a union of all links accross all Wikipedia language editions. Computation is performed by Andreas Thalhammer with 'danker' available at https://github.com/athalhammer/danker .
Facebook
TwitterAttribution-NonCommercial 4.0 (CC BY-NC 4.0)https://creativecommons.org/licenses/by-nc/4.0/
License information was derived automatically
Category-based imports from Wikidata, the structured data version of Wikipedia.
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
Wikidata QRank computed by Andreas Thalhammer based on a modification of Wikidata QRank (source https://github.com/athalhammer/wikidata-qrank ; originally developed by Sascha Brawer https://github.com/brawer/wikidata-qrank ). Computation uses sitelinks extracted via danker from Wikipedia projects (no wikibooks, no commons, etc).
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
A free, collaborative, multilingual, secondary knowledge base operated by the Wikimedia Foundation. Wikidata collects structured data to support Wikipedia, Wikimedia Commons, and other Wikimedia projects, as well as third-party applications worldwide. Each data item is identified by a Q-number and described through property-value statements (properties use P-numbers), with optional qualifiers and source references. The data is published under the Creative Commons CC0 public domain dedication, enabling unrestricted reuse. Wikidata exposes its data through a SPARQL endpoint, REST APIs, and linked data serializations.
Facebook
TwitterWikidata APIs are governed by Wikimedia's shared platform limits rather than per-account quotas. Anonymous and authenticated traffic is best-effort; abusive or unidentified clients are throttled or blocked. The SPARQL Query Service enforces a hard query-execution timeout. The MediaWiki Action API uses the maxlag parameter as a courtesy throttle when replication is behind. The User-Agent policy is mandatory - requests with a generic / missing User-Agent header may be denied without further notice.
Facebook
TwitterAttribution-ShareAlike 3.0 (CC BY-SA 3.0)https://creativecommons.org/licenses/by-sa/3.0/
License information was derived automatically
Wikipedia, the free encyclopedia, and Wikidata, the free knowledge base, are crowd-sourced projects supported by the Wikimedia Foundation. Wikipedia is nearly 20 years old and recently added its six millionth article in English. Wikidata, its younger machine-readable sister project, was created in 2012 but has been growing rapidly and currently contains more than 75 million items.
These projects contribute to the Wikimedia Foundation's mission of empowering people to develop and disseminate educational content under a free license. They are also heavily utilized by computer science research groups, especially those interested in natural language processing (NLP). The Wikimedia Foundation periodically releases snapshots of the raw data backing these projects, but these are in a variety of formats and were not designed for use in NLP research. In the Kensho R&D group, we spend a lot of time downloading, parsing, and experimenting with this raw data. The Kensho Derived Wikimedia Dataset (KDWD) is a condensed subset of the raw Wikimedia data in a form that we find helpful for NLP work. The KDWD has a CC BY-SA 3.0 license, so feel free to use it in your work too.
https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4301984%2F972e4157b97efe8c2c5ea17c983b1504%2Fkdwd_header_logos_2.jpg?generation=1580510520532141&alt=media" alt="">
This particular release consists of two main components - a link annotated corpus of English Wikipedia pages and a compact sample of the Wikidata knowledge base. We version the KDWD using the raw Wikimedia snapshot dates. The version string for this dataset is kdwd_enwiki_20191201_wikidata_20191202 indicating that this KDWD was built from the English Wikipedia snapshot from 2019 December 1 and the Wikidata snapshot from 2019 December 2. Below we describe these components in more detail.
Dive right in by checking out some of our example notebooks:
page.csv (page metadata and Wikipedia-to-Wikidata mapping)link_annotated_text.jsonl (plaintext of Wikipedia pages with link offsets)item.csv (item labels and descriptions in English)item_aliases.csv (item aliases in English)property.csv (property labels and descriptions in English)property_aliases.csv (property aliases in English)statements.csv (truthy qpq statements)The KDWD is three connected layers of data. The base layer is a plain text English Wikipedia corpus, the middle layer annotates the corpus by indicating which text spans are links, and the top layer connects the link text spans to items in Wikidata. Below we'll describe these layers in more detail.
https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4301984%2F19663d43bade0e92f578255f6e0d9dcd%2Fkensho_wiki_triple_layer.svg?generation=1580347573004185&alt=media" alt="">
The first part of the KDWD is derived from Wikipedia. In order to create a corpus of mostly natural text, we restrict our English Wikipedia page sample to those that:
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
A canonical corpus of 100,000 structured Wikidata entities prepared for reproducible English-Vietnamese semantic-search and embedding experiments. Each row preserves its Wikidata QID and contains normalized English/Vietnamese labels, descriptions, aliases, direct P31 identifiers, and deterministic document text.
This release contains source text and metadata; it does not contain precomputed embedding vectors. It is the canonical input corpus for the Qwen3-Embedding-8B Kaggle T4 x2 demo.
| Property | Value |
|---|---|
| Canonical rows | 100,000 |
| Languages | English (en), Vietnamese (vi) |
| Canonical format | Parquet with Zstandard compression |
| Source | Wikidata structured data |
| License | CC0 1.0 |
| Canonical SHA-256 | 1f69b4b64a0aecf41505aef3cf6fa5c5ce19e4a733d198765ee6f431beb711e9 |
The canonical Parquet schema includes QID, EN/VI labels, EN/VI descriptions, EN/VI aliases, direct P31 QIDs, source/license fields, deterministic EN/VI retrieval documents, and quality flags. The release also includes the raw QID discovery snapshot, structured entity snapshot, runtime manifest, provenance metadata, data-quality report, license note, and SHA-256 checksums.
QIDs were discovered from twelve weighted direct-P31 seed classes through the official Wikidata Query Service. Structured entity fields were fetched sequentially through the official Wikibase Action API in batches of at most 50, with a descriptive User-Agent, maxlag=5, request pacing, retry/backoff, and resumable state.
The canonical transform uses strict UTF-8 decoding, NFC normalization, control/replacement-character rejection, whitespace normalization, alias deduplication, deterministic document construction, and stable sorting. Of 107,744 scanned structured records, 100,000 were accepted and 7,744 were rejected for a missing usable label.
Missing descriptions remain null rather than being fabricated.
Only structured Wikidata fields are included. Wikipedia article bodies, news, Common Crawl, and arbitrary web pages are excluded. Direct P31 seed sampling is a diversity policy, not a statistically representative sample of all Wikidata. Labels and descriptions reflect the acquisition snapshot and may contain normal Wikidata omissions or later become outdated.
Live Wikidata changes over time. Reproducibility for this release is defined by the published raw JSONL snapshots and SHA256SUMS, followed by the deterministic canonical transform documented in the manifest and provenance files.
Run sha256sum -c SHA256SUMS after downloading. The runtime manifest separately pins the canonical Parquet filename, row count, source/license declaration, and SHA-256 digest.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
KeySearchWiki is a dataset for evaluating keyword search systems over Wikidata.
The dataset was automatically generated by leveraging Wikidata and Wikipedia set categories (e.g., Category:American television directors) as data sources for both relevant entities and queries.
Relevant entities are gathered by carefully navigating the Wikipedia set categories hierarchy in all available languages. Furthermore, those categories are refined and combined to derive more complex queries.
Detailed information about KeySearchWiki and its generation can be found on the Github page.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
FinDailyKG Wikidata Entity Enrichment KG is an incremental (daily/weekly) enrichment layer that converts news-derived entity âsurfacesâ into stable Wikidata QIDs and produces graph-ready Parquet tables for downstream analytics and GNN training.
Whatâs inside
Per-day node tables (examples): entity_nodes_day=*.parquet, org_nodes_day=*.parquet, industry_nodes_day=*.parquet
Per-day edge tables (examples): entity_mentions_day=*.parquet, org_industry_edges_day=*.parquet
Mapping tables: entity_qid_map_day=*.parquet (surface â QID), plus diagnostics/rejection tables for quality control
Caches (for fast incremental runs): surface_link_cache*.parquet, qid_meta_cache.parquet, class_parent_cache.parquet, wp_*_cache.parquet
Config + provenance artifacts: entitykg_config.json, run_metadata_day=*.json, MERGE_REPORT.json
How itâs produced (high level)
Start from news/entity surfaces (from FinDailyKG Dataset1, GDELT-based).
Link surfaces to Wikidata QIDs (with caching + retry/diagnostics).
Emit graph tables (nodes/edges) and diagnostics per day.
Designed to be re-run incrementally by reading the prior published version as a resume/cache seed.
Intended uses
Dynamic knowledge graph experiments (time-sliced KGs)
Entity linking / entity resolution analysis
Feature generation for financial NLP
Graph ML / GNN pipelines over daily snapshots
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset provides entity mappings between Freebase and Wikidata, enabling seamless integration between two large-scale knowledge graphs. It is based on the Wikidata data dump from October 28, 2013, and was originally published by Google under the CC0 (Public Domain) license.
The mappings are carefully filtered to ensure high reliability:
This strict filtering results in high-confidence entity alignments, making the dataset useful for research and real-world applications in knowledge graph systems.
Facebook
TwitterAttribution-NonCommercial 4.0 (CC BY-NC 4.0)https://creativecommons.org/licenses/by-nc/4.0/
License information was derived automatically
Persons of interest profiles from Wikidata, the structured data version of Wikipedia.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This is a dump from Wikidata from 2018-12-17 in JSON. This one is not avavailable anymore from Wikidata. It was downloaded originally from https://dumps.wikimedia.org/other/wikidata/20181217.json.gz and recompressed to fit on Zenodo.
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
Wikidata SPARQL endpoint joined to Wikipedia article URLs. Tags companies that have a notable encyclopedic presence (typically mid-market and above). Useful trust signal for buyer-side filtering.
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
Wikidata is a collaboratively edited knowledge base operated by the Wikimedia Foundation. It is intended to provide a common source of certain types of data which can be used by Wikimedia projects such as Wikipedia. Wikidata functions as a document-oriented database, centred on individual items. Items represent topics, for which basic information is stored that identifies each topic.