100+ datasets found
  1. b

    Wikidata

    • bioregistry.io
    Updated Nov 13, 2021
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    (2021). Wikidata [Dataset]. http://identifiers.org/biolink:WIKIDATA
    Explore at:
    Dataset updated
    Nov 13, 2021
    License

    CC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
    License information was derived automatically

    Description

    Wikidata is a collaboratively edited knowledge base operated by the Wikimedia Foundation. It is intended to provide a common source of certain types of data which can be used by Wikimedia projects such as Wikipedia. Wikidata functions as a document-oriented database, centred on individual items. Items represent topics, for which basic information is stored that identifies each topic.

  2. Wikidata Thematic Subgraph Selection

    • zenodo.org
    • data.niaid.nih.gov
    zip
    Updated May 24, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Lucas Jarnac; Lucas Jarnac; Miguel Couceiro; Miguel Couceiro; Pierre Monnin; Pierre Monnin (2024). Wikidata Thematic Subgraph Selection [Dataset]. http://doi.org/10.5281/zenodo.8091584
    Explore at:
    zipAvailable download formats
    Dataset updated
    May 24, 2024
    Dataset provided by
    Zenodohttp://zenodo.org/
    Authors
    Lucas Jarnac; Lucas Jarnac; Miguel Couceiro; Miguel Couceiro; Pierre Monnin; Pierre Monnin
    License

    Attribution-NonCommercial 4.0 (CC BY-NC 4.0)https://creativecommons.org/licenses/by-nc/4.0/
    License information was derived automatically

    Description

    Wikidata Thematic Subgraph Selection

    These datasets have been designed to train and evaluate algorithms to select thematic subgraphs of interest in a large knowledge graph from seed entities of interest. Specifically, we consider Wikidata. Given a set of seed QIDs of interest, a graph expansion is performed following P31, P279, and (-)P279 edges. Traversed classes that thematically deviates from seed QIDs of interest should be pruned. Datasets thus consist of classes reached from seed QIDs that are labeled as "to prune" or "to keep".

    Available datasets

    Dataset# Seed QIDs# Labeled decisions# Prune decisionsMin prune depthMax prune depth# Keep decisionsMin keep depthMax keep depth# Reached nodes up# Reached nodes down
    dataset1455523334641417691415072593609
    dataset2105982388125941311591247385

    Each dataset folder contains

    • datasetX.csv: a CSV file containing one seed QID per line (not the complete URL, just the QID). This CSV file has no header.
    • datasetX_labels.csv: a CSV file containing one seed QID per line and its label (not the complete URL, just the QID)
    • datasetX_gold_decisions.csv: a CSV file with seed QIDs, reached QIDs, and the labeled decision (1: keep, 0: prune)
    • datasetX_Y_folds.pkl: folds to train and test models based on the labeled decisions

    dataset1-2 consists of using dataset1 for training and dataset2 for testing.

    License

    Datasets are available under the CC BY-NC license.

  3. h

    Wikidata_Vectors_0.2

    • huggingface.co
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Philippe Saadé, Wikidata_Vectors_0.2 [Dataset]. https://huggingface.co/datasets/philippesaade/Wikidata_Vectors_0.2
    Explore at:
    Authors
    Philippe Saadé
    License

    https://choosealicense.com/licenses/cc0-1.0/https://choosealicense.com/licenses/cc0-1.0/

    Description

    Wikidata Entity Embeddings 0.2

      Dataset Summary
    

    Wikidata Entity Embeddings is a dataset of embedding vectors for Wikidata entities. Each vector represents a Wikidata item (Q...) or property (P...) based on textual information extracted from Wikidata. The dataset is part of the Wikidata Embedding Project, an initiative led by Wikimedia Deutschland in collaboration with Jina AI and IBM DataStax. The project provides a publicly accessible Wikidata Vector Database to
 See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/Wikidata_Vectors_0.2.

  4. Wikidata Reference

    • figshare.com
    gz
    Updated Mar 17, 2025
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Sven Hertling; Nandana Mihindukulasooriya (2025). Wikidata Reference [Dataset]. http://doi.org/10.6084/m9.figshare.28602170.v2
    Explore at:
    gzAvailable download formats
    Dataset updated
    Mar 17, 2025
    Dataset provided by
    figshare
    Figsharehttp://figshare.com/
    Authors
    Sven Hertling; Nandana Mihindukulasooriya
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    Dataset SummaryThe Triple-to-Text Alignment dataset aligns Knowledge Graph (KG) triples from Wikidata with diverse, real-world textual sources extracted from the web. Unlike previous datasets that rely primarily on Wikipedia text, this dataset provides a broader range of writing styles, tones, and structures by leveraging Wikidata references from various sources such as news articles, government reports, and scientific literature. Large language models (LLMs) were used to extract and validate text spans corresponding to KG triples, ensuring high-quality alignments. The dataset can be used for training and evaluating relation extraction (RE) and knowledge graph construction systems.Data FieldsEach row in the dataset consists of the following fields:subject (str): The subject entity of the knowledge graph triple.rel (str): The relation that connects the subject and object.object (str): The object entity of the knowledge graph triple.text (str): A natural language sentence that entails the given triple.validation (str): LLM-based validation results, including:Fluent Sentence(s): TRUE/FALSESubject mentioned in Text: TRUE/FALSERelation mentioned in Text: TRUE/FALSEObject mentioned in Text: TRUE/FALSEFact Entailed By Text: TRUE/FALSEFinal Answer: TRUE/FALSEreference_url (str): URL of the web source from which the text was extracted.subj_qid (str): Wikidata QID for the subject entity.rel_id (str): Wikidata Property ID for the relation.obj_qid (str): Wikidata QID for the object entity.Dataset CreationThe dataset was created through the following process:1. Triple-Reference Sampling and ExtractionAll relations from Wikidata were extracted using SPARQL queries.A sample of KG triples with associated reference URLs was collected for each relation.2. Domain Analysis and Web ScrapingURLs were grouped by domain, and sampled pages were analyzed to determine their primary language.English-language web pages were scraped and processed to extract plaintext content.3. LLM-Based Text Span Selection and ValidationLLMs were used to identify text spans from web content that correspond to KG triples.A Chain-of-Thought (CoT) prompting method was applied to validate whether the extracted text entailed the triple.The validation process included checking for fluency, subject mention, relation mention, object mention, and final entailment.4. Final Dataset Statistics12.5K Wikidata relations were analyzed, leading to 3.3M triple-reference pairs.After filtering for English content, 458K triple-web content pairs were processed with LLMs.80.5K validated triple-text alignments were included in the final dataset.

  5. Wikidata Plans Pricing

    • apis.io
    Updated May 7, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Wikidata (2026). Wikidata Plans Pricing [Dataset]. https://apis.io/plans/wikidata/wikidata-plans-pricing/
    Explore at:
    Dataset updated
    May 7, 2026
    Dataset authored and provided by
    Wikidatahttp://wikidata.org/
    Description

    Wikidata is a free, CC0-licensed knowledge base operated by the Wikimedia Foundation. All public APIs (Wikibase REST, MediaWiki Action, SPARQL Query Service, Linked Data interface, Recent Changes event stream) are free of charge with no paid tiers. Heavy or commercial users are expected to follow the User-Agent and maxlag policies, mirror data via the database dumps, or contract with Wikimedia Enterprise for commercial-grade access.

  6. Wikipedia Articles/Links and Wikidata Entities

    • kaggle.com
    zip
    Updated Mar 26, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Vizalyst (2026). Wikipedia Articles/Links and Wikidata Entities [Dataset]. https://www.kaggle.com/datasets/mrpaulalbert/wikipedia-articleslinks-and-wikidata-entities
    Explore at:
    zip(13104756484 bytes)Available download formats
    Dataset updated
    Mar 26, 2026
    Authors
    Vizalyst
    License

    https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

    Description

    WikiSnap25 is a collection of analysis‑ready tables that integrate content, metadata, and readership information for English Wikipedia articles and Wikidata entities as of mid‑2025. The goal is to lower the engineering barrier for empirical Wikimedia research by abstracting away the work of parsing XML/JSON dumps and aggregating pageview data into compact, well‑documented Parquet datasets.

    WikiSnap25 combines several official Wikimedia sources: the June 1, 2025 English Wikipedia pages‑articles‑multistream XML dump for article text and metadata, the June 2, 2025 full Wikidata JSON dump for entity‑level information, a decade of Wikimedia REST API monthly pageview data from July 2015 through June 2025, and additional dumps (e.g., stub meta‑history and redirect mappings) for revision context and canonicalization. A modular processing pipeline (XML parsing, wikitext processing, mwparserfromhell, and aggregation scripts) produces a consistent article‑level snapshot anchored at mid‑2025, enriched with long‑term readership and knowledge‑graph features.

    The following datasets are included:

    WikiSnap_2025_WP_Articles.parquet – Article‑level data for 7,011,415 English Wikipedia articles, including normalized titles, total pageviews (2015–2025), first two sentences of article text, outlink metrics (e.g., num_articles_linked_out), and mappings to Wikidata QIDs. This table supports studies of article popularity, connectivity, and entity types.

    WikiSnap_2025_WP_Edges.parquet – 314,337,621 intra‑Wikipedia hyperlinks between articles, with from/to article IDs and titles, endpoint pageviews, and a duplicate‑link flag for graph cleaning. This edge list is suitable for large‑scale network analyses of the article link graph.

    WikiSnap_2025_WP_Network_Metrics.parquet – Network metrics for articles with intra-links to other Wikipedia articles (7,008,375 rows), including centrality scores, PageRank, traffic‑weighted PageRank, HITS hub/authority, degrees, and Leiden community assignments computed from a deduplicated article link graph.

    WikiSnap_2025_WP_Article_Metrics.parquet – Additional article‑level metrics (7,011,415 rows) such as creation timestamp, edit lifespan in days, total revisions, number of unique editors, and a Gini‑style editor inequality index across redirect clusters.

    WikiSnap_2025_WD_Entities.parquet – Wikidata entity information for 116,183,072 items, including labels, instance‑of labels, temporal attributes (begin/end year), country of citizenship labels, and sitelinks to English Wikipedia article titles, with disambiguation pages filtered and truthy claims prioritized.

    WikiSnap_2025_WD_Metrics.parquet – Wikidata metrics for the same 116M entities, including counts of claims, in/out links, references, qualifiers, notable properties, article‑quality badges, and external identifiers, enabling notability and completeness assessments.

    WikiSnap_2025_WD_Humans.parquet – A subset for human entities (P31=Q5) with 12,371,744 rows, containing birth and death years, occupation labels, gender, citizenship, and notability metrics, tailored to demographic and biographical analysis.

    These datasets are intended to support a wide range of tasks, including correlating network centrality with pageview patterns, assessing Wikidata notability via property richness and sitelinks, and exploring editorial activity and inequality, without requiring researchers to build their own large‑scale Wikimedia processing pipelines from scratch.

    Limitations. The current release focuses on English Wikipedia main‑namespace articles as of June 1, 2025; Wikidata content as of early June 2025; and article‑level pageviews aggregated over July 2015–June 2025 without finer temporal granularity. Redirects, stubs, and full revision histories are not included in these tables.

    All datasets and the full processing code are released to facilitate reproducible Wikimedia research and to enable cross‑study comparability on a shared, 2025‑anchored snapshot.

    License and citation. All WikiSnap25 datasets and associated processing code archived in this record are dedicated to the public domain under the Creative Commons Zero v1.0 Universal (CC0 1.0) public domain dedication. Users are free to reuse, modify, and redistribute the materials without restriction.

  7. wikidata-title-desc

    • huggingface.co
    Updated Jun 4, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Wikimedia (2026). wikidata-title-desc [Dataset]. https://huggingface.co/datasets/wikimedia/wikidata-title-desc
    Explore at:
    Dataset updated
    Jun 4, 2026
    Dataset provided by
    Wikimedia Foundationhttp://www.wikimedia.org/
    Authors
    Wikimedia
    License

    Attribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
    License information was derived automatically

    Description

    Wikidata Title and Description

    Wikidata entity titles and short descriptions for all 324 Wikipedia language editions, extracted from the Wikimedia Analytics wmf.wikidata_entity table. Each row links a Wikidata QID to the title and description as they appear in a specific language edition of Wikipedia.

      Dataset Details
    

    Field Value

    Source wmf.wikidata_entity (Wikidata + Wikipedia sitelinks)

    Snapshot 2026-03-30

    Languages 324

    Total titles 88,292,409

    Titles
 See the full description on the dataset page: https://huggingface.co/datasets/wikimedia/wikidata-title-desc.

  8. a

    Wikidata PageRank

    • danker.s3.amazonaws.com
    application/vnd.hdt +3
    Updated Sep 10, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Andreas Thalhammer (2026). Wikidata PageRank [Dataset]. https://danker.s3.amazonaws.com/index.html
    Explore at:
    tsv, ttl, application/vnd.hdt, ntAvailable download formats
    Dataset updated
    Sep 10, 2026
    Authors
    Andreas Thalhammer
    License

    Attribution-ShareAlike 3.0 (CC BY-SA 3.0)https://creativecommons.org/licenses/by-sa/3.0/
    License information was derived automatically

    Description

    Regularly published dataset of PageRank scores for Wikidata entities. The underlying link graph is formed by a union of all links accross all Wikipedia language editions. Computation is performed by Andreas Thalhammer with 'danker' available at https://github.com/athalhammer/danker .

  9. Wikidata Persons in Relevant Categories

    • opensanctions.org
    Updated Sep 16, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Wikidata (2026). Wikidata Persons in Relevant Categories [Dataset]. https://www.opensanctions.org/datasets/wd_categories/
    Explore at:
    Dataset updated
    Sep 16, 2026
    Dataset authored and provided by
    Wikidatahttp://wikidata.org/
    License

    Attribution-NonCommercial 4.0 (CC BY-NC 4.0)https://creativecommons.org/licenses/by-nc/4.0/
    License information was derived automatically

    Description

    Category-based imports from Wikidata, the structured data version of Wikipedia.

  10. a

    Wikidata QRank [athaMod]

    • danker.s3.amazonaws.com
    csv
    Updated Sep 10, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Andreas Thalhammer (2026). Wikidata QRank [athaMod] [Dataset]. https://danker.s3.amazonaws.com/index.html
    Explore at:
    csvAvailable download formats
    Dataset updated
    Sep 10, 2026
    Authors
    Andreas Thalhammer
    License

    CC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
    License information was derived automatically

    Description

    Wikidata QRank computed by Andreas Thalhammer based on a modification of Wikidata QRank (source https://github.com/athalhammer/wikidata-qrank ; originally developed by Sascha Brawer https://github.com/brawer/wikidata-qrank ). Computation uses sitelinks extracted via danker from Wikipedia projects (no wikibooks, no commons, etc).

  11. Wikidata

    • msi.dublincore.org
    Updated 2012
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Wikimedia Foundation (2012). Wikidata [Dataset]. https://msi.dublincore.org/standards/wikidata
    Explore at:
    Dataset updated
    2012
    Dataset authored and provided by
    Wikimedia Foundationhttp://www.wikimedia.org/
    License

    CC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
    License information was derived automatically

    Description

    A free, collaborative, multilingual, secondary knowledge base operated by the Wikimedia Foundation. Wikidata collects structured data to support Wikipedia, Wikimedia Commons, and other Wikimedia projects, as well as third-party applications worldwide. Each data item is identified by a Q-number and described through property-value statements (properties use P-numbers), with optional qualifiers and source references. The data is published under the Creative Commons CC0 public domain dedication, enabling unrestricted reuse. Wikidata exposes its data through a SPARQL endpoint, REST APIs, and linked data serializations.

  12. Wikidata Rate Limits

    • apis.io
    Updated Jun 30, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Wikidata (2026). Wikidata Rate Limits [Dataset]. https://apis.io/rate-limits/wikidata/wikidata-rate-limits/
    Explore at:
    Dataset updated
    Jun 30, 2026
    Dataset authored and provided by
    Wikidatahttp://wikidata.org/
    Description

    Wikidata APIs are governed by Wikimedia's shared platform limits rather than per-account quotas. Anonymous and authenticated traffic is best-effort; abusive or unidentified clients are throttled or blocked. The SPARQL Query Service enforces a hard query-execution timeout. The MediaWiki Action API uses the maxlag parameter as a courtesy throttle when replication is behind. The User-Agent policy is mandatory - requests with a generic / missing User-Agent header may be denied without further notice.

  13. Kensho Derived Wikimedia Dataset

    • kaggle.com
    zip
    Updated Jan 24, 2020
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Kensho R&D (2020). Kensho Derived Wikimedia Dataset [Dataset]. https://www.kaggle.com/datasets/kenshoresearch/kensho-derived-wikimedia-data/code
    Explore at:
    zip(8760044227 bytes)Available download formats
    Dataset updated
    Jan 24, 2020
    Authors
    Kensho R&D
    License

    Attribution-ShareAlike 3.0 (CC BY-SA 3.0)https://creativecommons.org/licenses/by-sa/3.0/
    License information was derived automatically

    Description

    Kensho Derived Wikimedia Dataset

    Wikipedia, the free encyclopedia, and Wikidata, the free knowledge base, are crowd-sourced projects supported by the Wikimedia Foundation. Wikipedia is nearly 20 years old and recently added its six millionth article in English. Wikidata, its younger machine-readable sister project, was created in 2012 but has been growing rapidly and currently contains more than 75 million items.

    These projects contribute to the Wikimedia Foundation's mission of empowering people to develop and disseminate educational content under a free license. They are also heavily utilized by computer science research groups, especially those interested in natural language processing (NLP). The Wikimedia Foundation periodically releases snapshots of the raw data backing these projects, but these are in a variety of formats and were not designed for use in NLP research. In the Kensho R&D group, we spend a lot of time downloading, parsing, and experimenting with this raw data. The Kensho Derived Wikimedia Dataset (KDWD) is a condensed subset of the raw Wikimedia data in a form that we find helpful for NLP work. The KDWD has a CC BY-SA 3.0 license, so feel free to use it in your work too.

    https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4301984%2F972e4157b97efe8c2c5ea17c983b1504%2Fkdwd_header_logos_2.jpg?generation=1580510520532141&alt=media" alt="">

    This particular release consists of two main components - a link annotated corpus of English Wikipedia pages and a compact sample of the Wikidata knowledge base. We version the KDWD using the raw Wikimedia snapshot dates. The version string for this dataset is kdwd_enwiki_20191201_wikidata_20191202 indicating that this KDWD was built from the English Wikipedia snapshot from 2019 December 1 and the Wikidata snapshot from 2019 December 2. Below we describe these components in more detail.

    Example Notebooks

    Dive right in by checking out some of our example notebooks:

    Updates / Changelog

    • initial release 2020-01-31

    File Summary

    • Wikipedia
      • page.csv (page metadata and Wikipedia-to-Wikidata mapping)
      • link_annotated_text.jsonl (plaintext of Wikipedia pages with link offsets)
    • Wikidata
      • item.csv (item labels and descriptions in English)
      • item_aliases.csv (item aliases in English)
      • property.csv (property labels and descriptions in English)
      • property_aliases.csv (property aliases in English)
      • statements.csv (truthy qpq statements)

    Three Layers of Data

    The KDWD is three connected layers of data. The base layer is a plain text English Wikipedia corpus, the middle layer annotates the corpus by indicating which text spans are links, and the top layer connects the link text spans to items in Wikidata. Below we'll describe these layers in more detail.

    https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4301984%2F19663d43bade0e92f578255f6e0d9dcd%2Fkensho_wiki_triple_layer.svg?generation=1580347573004185&alt=media" alt="">

    Wikipedia Sample

    The first part of the KDWD is derived from Wikipedia. In order to create a corpus of mostly natural text, we restrict our English Wikipedia page sample to those that:

  14. Wikidata EN-VI Semantic Search 100K

    • kaggle.com
    zip
    Updated Aug 27, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Dang Khoa (2026). Wikidata EN-VI Semantic Search 100K [Dataset]. https://www.kaggle.com/datasets/dangkhoa2016/wikidata-en-vi-semantic-search-100k
    Explore at:
    zip(17412704 bytes)Available download formats
    Dataset updated
    Aug 27, 2026
    Authors
    Dang Khoa
    License

    https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

    Description

    Overview

    A canonical corpus of 100,000 structured Wikidata entities prepared for reproducible English-Vietnamese semantic-search and embedding experiments. Each row preserves its Wikidata QID and contains normalized English/Vietnamese labels, descriptions, aliases, direct P31 identifiers, and deterministic document text.

    This release contains source text and metadata; it does not contain precomputed embedding vectors. It is the canonical input corpus for the Qwen3-Embedding-8B Kaggle T4 x2 demo.

    Dataset summary

    PropertyValue
    Canonical rows100,000
    LanguagesEnglish (en), Vietnamese (vi)
    Canonical formatParquet with Zstandard compression
    SourceWikidata structured data
    LicenseCC0 1.0
    Canonical SHA-2561f69b4b64a0aecf41505aef3cf6fa5c5ce19e4a733d198765ee6f431beb711e9

    Contents

    The canonical Parquet schema includes QID, EN/VI labels, EN/VI descriptions, EN/VI aliases, direct P31 QIDs, source/license fields, deterministic EN/VI retrieval documents, and quality flags. The release also includes the raw QID discovery snapshot, structured entity snapshot, runtime manifest, provenance metadata, data-quality report, license note, and SHA-256 checksums.

    Collection and processing

    QIDs were discovered from twelve weighted direct-P31 seed classes through the official Wikidata Query Service. Structured entity fields were fetched sequentially through the official Wikibase Action API in batches of at most 50, with a descriptive User-Agent, maxlag=5, request pacing, retry/backoff, and resumable state.

    The canonical transform uses strict UTF-8 decoding, NFC normalization, control/replacement-character rejection, whitespace normalization, alias deduplication, deterministic document construction, and stable sorting. Of 107,744 scanned structured records, 100,000 were accepted and 7,744 were rejected for a missing usable label.

    Language coverage

    • Vietnamese label coverage: 100%
    • English label coverage: 99.967%
    • Vietnamese document coverage: 100%
    • English document coverage: 99.967%
    • Vietnamese description coverage: 90.961%
    • English description coverage: 90.939%

    Missing descriptions remain null rather than being fabricated.

    Intended use

    • English-Vietnamese semantic-search demonstrations
    • multilingual embedding and retrieval evaluation
    • FAISS/Qdrant/vector-database indexing experiments
    • entity retrieval, reranking, and query-instruction studies
    • reproducible Kaggle notebooks and integration tests

    Scope and limitations

    Only structured Wikidata fields are included. Wikipedia article bodies, news, Common Crawl, and arbitrary web pages are excluded. Direct P31 seed sampling is a diversity policy, not a statistically representative sample of all Wikidata. Labels and descriptions reflect the acquisition snapshot and may contain normal Wikidata omissions or later become outdated.

    Live Wikidata changes over time. Reproducibility for this release is defined by the published raw JSONL snapshots and SHA256SUMS, followed by the deterministic canonical transform documented in the manifest and provenance files.

    Integrity

    Run sha256sum -c SHA256SUMS after downloading. The runtime manifest separately pins the canonical Parquet filename, row count, source/license declaration, and SHA-256 digest.

  15. KeySearchWiki

    • zenodo.org
    • data.niaid.nih.gov
    zip
    Updated Feb 14, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Leila Feddoul; Leila Feddoul; Frank Löffler; Frank Löffler; Sirko Schindler; Sirko Schindler (2022). KeySearchWiki [Dataset]. http://doi.org/10.5281/zenodo.4955200
    Explore at:
    zipAvailable download formats
    Dataset updated
    Feb 14, 2022
    Dataset provided by
    Zenodohttp://zenodo.org/
    Authors
    Leila Feddoul; Leila Feddoul; Frank Löffler; Frank Löffler; Sirko Schindler; Sirko Schindler
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    KeySearchWiki is a dataset for evaluating keyword search systems over Wikidata.

    The dataset was automatically generated by leveraging Wikidata and Wikipedia set categories (e.g., Category:American television directors) as data sources for both relevant entities and queries.
    Relevant entities are gathered by carefully navigating the Wikipedia set categories hierarchy in all available languages. Furthermore, those categories are refined and combined to derive more complex queries.

    Detailed information about KeySearchWiki and its generation can be found on the Github page.

  16. finDailyKG wikidata entity enrichment KG

    • kaggle.com
    zip
    Updated Aug 5, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    The citation is currently not available for this dataset.
    Explore at:
    zip(1332901334 bytes)Available download formats
    Dataset updated
    Aug 5, 2026
    Authors
    Wilmer E. Henao
    License

    https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

    Description

    FinDailyKG Wikidata Entity Enrichment KG is an incremental (daily/weekly) enrichment layer that converts news-derived entity “surfaces” into stable Wikidata QIDs and produces graph-ready Parquet tables for downstream analytics and GNN training.

    What’s inside

    Per-day node tables (examples): entity_nodes_day=*.parquet, org_nodes_day=*.parquet, industry_nodes_day=*.parquet

    Per-day edge tables (examples): entity_mentions_day=*.parquet, org_industry_edges_day=*.parquet

    Mapping tables: entity_qid_map_day=*.parquet (surface → QID), plus diagnostics/rejection tables for quality control

    Caches (for fast incremental runs): surface_link_cache*.parquet, qid_meta_cache.parquet, class_parent_cache.parquet, wp_*_cache.parquet

    Config + provenance artifacts: entitykg_config.json, run_metadata_day=*.json, MERGE_REPORT.json

    How it’s produced (high level)

    Start from news/entity surfaces (from FinDailyKG Dataset1, GDELT-based).

    Link surfaces to Wikidata QIDs (with caching + retry/diagnostics).

    Emit graph tables (nodes/edges) and diagnostics per day.

    Designed to be re-run incrementally by reading the prior published version as a resume/cache seed.

    Intended uses

    Dynamic knowledge graph experiments (time-sliced KGs)

    Entity linking / entity resolution analysis

    Feature generation for financial NLP

    Graph ML / GNN pipelines over daily snapshots

  17. Freebase/Wikidata Mappings

    • kaggle.com
    zip
    Updated Mar 13, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    The citation is currently not available for this dataset.
    Explore at:
    zip(21894706 bytes)Available download formats
    Dataset updated
    Mar 13, 2026
    Authors
    Dhruv Bansal
    License

    https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

    Description

    This dataset provides entity mappings between Freebase and Wikidata, enabling seamless integration between two large-scale knowledge graphs. It is based on the Wikidata data dump from October 28, 2013, and was originally published by Google under the CC0 (Public Domain) license.

    The mappings are carefully filtered to ensure high reliability:

    • Each mapping includes at least two shared Wikipedia links
    • There are no conflicting Wikipedia links

    This strict filtering results in high-confidence entity alignments, making the dataset useful for research and real-world applications in knowledge graph systems.

  18. Wikidata Entities of Interest

    • opensanctions.org
    csv
    Updated Aug 9, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Wikidata (2026). Wikidata Entities of Interest [Dataset]. https://www.opensanctions.org/datasets/wd_curated/
    Explore at:
    csvAvailable download formats
    Dataset updated
    Aug 9, 2026
    Dataset authored and provided by
    Wikidatahttp://wikidata.org/
    License

    Attribution-NonCommercial 4.0 (CC BY-NC 4.0)https://creativecommons.org/licenses/by-nc/4.0/
    License information was derived automatically

    Description

    Persons of interest profiles from Wikidata, the structured data version of Wikipedia.

  19. Wikidata dump from 2018-12-17 in JSON

    • zenodo.org
    • data.niaid.nih.gov
    gz
    Updated Jan 15, 2021
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Jakub KlĂ­mek; Jakub KlĂ­mek; Petr Ć koda; Petr Ć koda (2021). Wikidata dump from 2018-12-17 in JSON [Dataset]. http://doi.org/10.5281/zenodo.4436356
    Explore at:
    gzAvailable download formats
    Dataset updated
    Jan 15, 2021
    Dataset provided by
    Zenodohttp://zenodo.org/
    Authors
    Jakub KlĂ­mek; Jakub KlĂ­mek; Petr Ć koda; Petr Ć koda
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    This is a dump from Wikidata from 2018-12-17 in JSON. This one is not avavailable anymore from Wikidata. It was downloaded originally from https://dumps.wikimedia.org/other/wikidata/20181217.json.gz and recompressed to fit on Zenodo.

  20. s

    Wikidata + Wikipedia

    • smbs.com
    Updated Jun 20, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    The citation is currently not available for this dataset.
    Explore at:
    Dataset updated
    Jun 20, 2026
    Dataset provided by
    Wikimedia Foundation
    License

    CC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
    License information was derived automatically

    Description

    Wikidata SPARQL endpoint joined to Wikipedia article URLs. Tags companies that have a notable encyclopedic presence (typically mid-market and above). Useful trust signal for buyer-side filtering.

Share
FacebookFacebook
TwitterTwitter
Email
Click to copy link
Link copied
Close
Cite
(2021). Wikidata [Dataset]. http://identifiers.org/biolink:WIKIDATA

Wikidata

Explore at:
Dataset updated
Nov 13, 2021
License

CC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically

Description

Wikidata is a collaboratively edited knowledge base operated by the Wikimedia Foundation. It is intended to provide a common source of certain types of data which can be used by Wikimedia projects such as Wikipedia. Wikidata functions as a document-oriented database, centred on individual items. Items represent topics, for which basic information is stored that identifies each topic.

Search
Clear search
Close search
Google apps
Main menu