12 datasets found
  1. Wiki STEM Corpus

    • kaggle.com
    zip
    Updated Apr 12, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Raja Biswas (2024). Wiki STEM Corpus [Dataset]. https://www.kaggle.com/datasets/conjuring92/wiki-stem-corpus
    Explore at:
    zip(891581618 bytes)Available download formats
    Dataset updated
    Apr 12, 2024
    Authors
    Raja Biswas
    License

    Attribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
    License information was derived automatically

    Description

    We created a STEM (Science, Technology, Engineering and Mathematics) corpus by filtering wikipedia articles based on their category metadata. During extraction of wiki page contents, we mitigated the frequent rendering issues (number, equations & symbols) prevalent in existing wiki datasets.

    For filtering, we first defined a set of seed wikipedia categories related to STEM topics such as Category:Concepts in physics, Category:Physical quantities, etc. For each category, recursively collect the member pages and subcategories up to a certain depth. We next extracted the page contents of the collected wiki URLs using Wikipedia-API (400k+ pages).

    Chunking: We first split the full text from each article based on different sections. The longer sections were further broken down into smaller chunks containing approximately 300 tokens (deberta-v3 tokenizer).

    This dataset can be embedded and used for RAG over STEM wiki.

    References: - Wiki STEM url collection: https://www.kaggle.com/code/conjuring92/d01-wiki-urls/notebook - Extraction of page content: https://www.kaggle.com/code/conjuring92/s04-stem-wiki-fetch - Chunking: https://www.kaggle.com/code/conjuring92/d504-chunking/notebook

  2. h

    1.5-Million-English-STEM-Test-Questions-Data-Sample

    • huggingface.co
    Updated Aug 11, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Nexdata-AI (2026). 1.5-Million-English-STEM-Test-Questions-Data-Sample [Dataset]. https://huggingface.co/datasets/Nexdata-AI/1.5-Million-English-STEM-Test-Questions-Data-Sample
    Explore at:
    Dataset updated
    Aug 11, 2026
    Authors
    Nexdata-AI
    Description

    Description

    This dataset contains 1.5 million English science and engineering test questions, including mathematics, physics, chemistry, biology, and other STEM subjects at the university level. Each questions contain title, answer, parse, type, subject, grade. The dataset can be used for large model subject knowledge enhancement tasks. For more details, please refer to the link: https://www.nexdata.ai/datasets/llm/1881?source=Huggingface

      Content
    

    Science subjects
 See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/1.5-Million-English-STEM-Test-Questions-Data-Sample.

  3. F

    STEM-NER-60k

    • data.uni-hannover.de
    zip
    Updated May 24, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    TIB (2022). STEM-NER-60k [Dataset]. https://data.uni-hannover.de/dataset/stem-ner-60k
    Explore at:
    zipAvailable download formats
    Dataset updated
    May 24, 2022
    Dataset authored and provided by
    TIB
    License

    Attribution-ShareAlike 3.0 (CC BY-SA 3.0)https://creativecommons.org/licenses/by-sa/3.0/
    License information was derived automatically

    Description

    A Large-scale Dataset of STEM Science as PROCESS, METHOD, MATERIAL, and DATA Named Entities

    This repository hosts data as a follow-up study to the following publications

    D'Souza, J., Hoppe, A., Brack, A., Jaradeh, M., Auer, S., & Ewerth, R. (2020). The STEM-ECR Dataset: Grounding Scientific Entity References in STEM Scholarly Content to Authoritative Encyclopedic and Lexicographic Sources. In Proceedings of The 12th Language Resources and Evaluation Conference (pp. 2192–2203). European Language Resources Association.

    Brack, A., D’Souza, J., Hoppe, A., Auer, S., Ewerth, R. (2020). Domain-Independent Extraction of Scientific Concepts from Research Articles. In: , et al. Advances in Information Retrieval. ECIR 2020. Lecture Notes in Computer Science, vol 12035. Springer, Cham. https://doi.org/10.1007/978-3-030-45439-5_17

    Supporting dataset link https://data.uni-hannover.de/dataset/stem-ecr-v1-0

    Description

    Roughly 60,000 titles and abstracts of scholarly articles with the CC-BY redistributable license were downloaded from Elsevier. The articles spanned 10 STEM domains which were the most prolific on Elsevier viz., Agriculture, Astronomy, Biology, Chemistry, Computer Science, Earth Science, Engineering, Material Science, and Mathematics. The STEM NER system reported in the publication above was applied on these articles. An automatically extracted dataset of 4 typed entities, viz., Process, Method, Material, and Data was created.

    What this repository contains?

    Aggregated lists of Process, Method, Material, and Data entities with respective occurrence counts extracted from 59,984 scholarly publications organized per the 10 STEM domains considered.

    Additionally, the list of Elsevier CC-BY articles used in this study are provided in the raw-data directory of the repository.

    Useful Links

  4. r

    STEM-NER-60k

    • resodate.org
    Updated May 24, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Jennifer D'Souza (2022). STEM-NER-60k [Dataset]. https://resodate.org/resources/aHR0cHM6Ly9kYXRhLnVuaS1oYW5ub3Zlci5kZS9kYXRhc2V0L3N0ZW0tbmVyLTYwaw==
    Explore at:
    Dataset updated
    May 24, 2022
    Dataset provided by
    Technische Informationsbibliothek (TIB)
    Forschungsdaten UniversitÀt Hannover
    Authors
    Jennifer D'Souza
    Description

    A Large-scale Dataset of STEM Science as PROCESS, METHOD, MATERIAL, and DATA Named Entities

    This repository hosts data as a follow-up study to the following publications

    D'Souza, J., Hoppe, A., Brack, A., Jaradeh, M., Auer, S., & Ewerth, R. (2020). The STEM-ECR Dataset: Grounding Scientific Entity References in STEM Scholarly Content to Authoritative Encyclopedic and Lexicographic Sources. In Proceedings of The 12th Language Resources and Evaluation Conference (pp. 2192–2203). European Language Resources Association.

    Brack, A., D’Souza, J., Hoppe, A., Auer, S., Ewerth, R. (2020). Domain-Independent Extraction of Scientific Concepts from Research Articles. In: , et al. Advances in Information Retrieval. ECIR 2020. Lecture Notes in Computer Science, vol 12035. Springer, Cham. https://doi.org/10.1007/978-3-030-45439-5_17

    Supporting dataset link https://data.uni-hannover.de/dataset/stem-ecr-v1-0

    Description

    Roughly 60,000 titles and abstracts of scholarly articles with the CC-BY redistributable license were downloaded from Elsevier. The articles spanned 10 STEM domains which were the most prolific on Elsevier viz., Agriculture, Astronomy, Biology, Chemistry, Computer Science, Earth Science, Engineering, Material Science, and Mathematics. The STEM NER system reported in the publication above was applied on these articles. An automatically extracted dataset of 4 typed entities, viz., Process, Method, Material, and Data was created.

    What this repository contains?

    Aggregated lists of Process, Method, Material, and Data entities with respective occurrence counts extracted from 59,984 scholarly publications organized per the 10 STEM domains considered.

    Additionally, the list of Elsevier CC-BY articles used in this study are provided in the raw-data directory of the repository.

    Useful Links

  5. t

    Jennifer D'Souza (2022). Dataset: STEM-NER-60k....

    • service.tib.eu
    Updated Apr 26, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    (2022). Jennifer D'Souza (2022). Dataset: STEM-NER-60k. https://doi.org/10.25835/heyid7l7 [Dataset]. https://service.tib.eu/ldmservice/dataset/luh-stem-ner-60k
    Explore at:
    Dataset updated
    Apr 26, 2022
    License

    Attribution-ShareAlike 3.0 (CC BY-SA 3.0)https://creativecommons.org/licenses/by-sa/3.0/
    License information was derived automatically

    Description

    A Large-scale Dataset of STEM Science as PROCESS, METHOD, MATERIAL, and DATA Named Entities This repository hosts data as a follow-up study to the following publications D'Souza, J., Hoppe, A., Brack, A., Jaradeh, M., Auer, S., & Ewerth, R. (2020). The STEM-ECR Dataset: Grounding Scientific Entity References in STEM Scholarly Content to Authoritative Encyclopedic and Lexicographic Sources. In Proceedings of The 12th Language Resources and Evaluation Conference (pp. 2192–2203). European Language Resources Association. Brack, A., D’Souza, J., Hoppe, A., Auer, S., Ewerth, R. (2020). Domain-Independent Extraction of Scientific Concepts from Research Articles. In: , et al. Advances in Information Retrieval. ECIR 2020. Lecture Notes in Computer Science, vol 12035. Springer, Cham. https://doi.org/10.1007/978-3-030-45439-5_17 Supporting dataset link https://data.uni-hannover.de/dataset/stem-ecr-v1-0 Description Roughly 60,000 titles and abstracts of scholarly articles with the CC-BY redistributable license were downloaded from Elsevier. The articles spanned 10 STEM domains which were the most prolific on Elsevier viz., Agriculture, Astronomy, Biology, Chemistry, Computer Science, Earth Science, Engineering, Material Science, and Mathematics. The STEM NER system reported in the publication above was applied on these articles. An automatically extracted dataset of 4 typed entities, viz., Process, Method, Material, and Data was created. What this repository contains?

  6. Scientists and Engineers Statistical Data System (SESTAT), United States,...

    • datalumos.org
    Updated Jul 16, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    National Center for Science and Engineering Statistics/National Science Foundation (2025). Scientists and Engineers Statistical Data System (SESTAT), United States, 1993-1999, 2001, 2003, 2006, 2008, 2010, 2013 [Dataset]. http://doi.org/10.3886/E236241V1
    Explore at:
    Dataset updated
    Jul 16, 2025
    Dataset provided by
    National Science Foundationhttp://www.nsf.gov/
    National Center for Science and Engineering Statisticshttp://ncses.nsf.gov/
    Authors
    National Center for Science and Engineering Statistics/National Science Foundation
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Area covered
    United States
    Description

    SummarySESTAT is the Scientists and Engineers Statistical Data System. SESTAT was established in 1993 and comprised of three workforce surveys from the National Center for Science and Engineering Statistics within the U.S. National Science Foundation. This integrated data system is a unique source of longitudinal information on the education and employment of the college-educated U.S. science, technology, engineering, and mathematics (STEM) workforce. It was integrated through 2013.The Scientists and Engineers Statistical Data System (SESTAT) comprises three demographic surveys of scientists and engineers sponsored by sponsored by the National Center for Science and Engineering Statistics (NCSES) within the U.S. National Science Foundation (NSF): the National Survey of College Graduates (NSCG), the National Survey of Recent College Graduates (NSRCG), and the Survey of Doctorate Recipients (SDR). The three component surveys used similar questionnaires, survey reference dates, data collection periods, and data-processing procedures to facilitate integration for SESTAT. The three surveys were designed to provide maximum coverage of the target population—namely, scientists and engineers—with special emphasis given to relatively rare populations (e.g., doctorate recipients, recent graduates, and minorities). Overall, SESTAT provides a comprehensive picture of the number and characteristics of individuals in the United States with a bachelor's or higher-level degree and their employment, with a focus on those having science and engineering (S&E) degrees or working in S&E occupations. In the 2000s, this definition was expanded to include S&E-related degrees and occupations.Background. Since 1993, SESTAT has provided a unique source of information on the education and employment of the college-educated U.S. S&E workforce by integrating three surveys: NSCG, NSRCG, and SDR. The establishment of SESTAT design was based on recommendations from a 1989 CNSTAT panel study report, Surveying the Nation's Scientists and Engineers—A Data System for the 1990s. This CNSTAT recommendation encouraged NSF to target the population of college graduates trained in S&E fields and those with employment in S&E occupations, to conduct the postcensal survey of college graduates, to conduct a survey for new graduating bachelor's or master's degree recipients, and to continue to support the ongoing SDR. The 2010 survey cycle introduced the first redesign of SESTAT during its nearly 20 years of existence.

  7. f

    Supplementary file 2_STEM approach using soccer: improving academic...

    • figshare.com
    docx
    Updated Feb 24, 2025
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Miguel Ángel Queiruga-Dios; José Benito Våzquez Dorrío; María Consuelo Såiz-Manzanares; Emilia López-Iñesta; María Diez-Ojeda (2025). Supplementary file 2_STEM approach using soccer: improving academic performance in Physics and Mathematics in a real-world context.docx [Dataset]. http://doi.org/10.3389/fpsyg.2025.1503397.s002
    Explore at:
    docxAvailable download formats
    Dataset updated
    Feb 24, 2025
    Dataset provided by
    Frontiers
    Authors
    Miguel Ángel Queiruga-Dios; José Benito Våzquez Dorrío; María Consuelo Såiz-Manzanares; Emilia López-Iñesta; María Diez-Ojeda
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Area covered
    World
    Description

    This proposal adds original approaches to the currently scarce body of practical evidence on the application of STEM innovations in the curriculum. A teaching-learning program was designed in a real-world context such as the game of soccer with a STEM (Science, Technology, Engineering and Mathematics) approach through a cooperative problem-solving methodology. The objectives of the research focus on analyzing the effect of the use of this STEM unit on the academic performance of students, taking into account the gender variable; and their appreciation of the activities and methodology used, as well as the challenges encountered and their solutions. The intervention was implemented in the 4th year of Compulsory Secondary Education in a school in Spain with 36 students (24 girls and 12 boys). Academic performance was analyzed taking into account the gender variable, for which a quasi-experimental design was applied before and after with a control group. The appreciation and interest of the experimental group regarding the methodology used as well as the difficulties that arose were studied. As a result, there is an improvement in the academic performance, which is more evident in girls. The methodology has been valued positively and the greatest difficulties refer to the distribution of roles and understanding and carrying out the activities, however, these difficulties were resolved with the help of classmates and the teacher.

  8. d

    How syllabi relate to outcomes in higher education: An evaluation of syllabi...

    • search.dataone.org
    Updated Jul 29, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Maryam Eslami; Brian Sato; Kameryn Denaro (2025). How syllabi relate to outcomes in higher education: An evaluation of syllabi learner-centeredness and grade inequities in STEM [Dataset]. http://doi.org/10.7280/D1NH6N
    Explore at:
    Dataset updated
    Jul 29, 2025
    Dataset provided by
    Dryad Digital Repository
    Authors
    Maryam Eslami; Brian Sato; Kameryn Denaro
    Time period covered
    Jan 1, 2022
    Description

    Fostering equity in undergraduate science, technology, engineering, and mathematics (STEM) programs can be accomplished by incorporating learner-centered pedagogies, resulting in the closing of opportunity gaps (defined in this research as the difference in grades earned by minoritized and non-minoritized students). We assessed STEM courses that exhibit small and large opportunity gaps at a minority-serving, research-intensive university, and evaluated the degree to which their syllabi are learner-centered, according to a previously validated rubric. We specifically chose syllabi as they are often the first interaction a student has with a course and can serve to establish expectations for course policies and practices. We found that STEM courses with more learner-centered syllabi had smaller opportunity gaps. The syllabus rubric factor that most correlated with smaller opportunity gaps was Power and Control, which reflects the Student's Role, Outside Resources, and Syllabus Focus. This..., This dataset is composed of rubric scores for 50 course syllabi of STEM classes in a research-intensive university with a large population of minoritized students as well as some institutional data (here defined as African-American, Latinx, Pacific Islander, and American Indian). We wanted to examine the relationship between racial grade gaps (here labeled as opportunity gaps) and the degree of learner-centeredness of the syllabi since course syllabi are good representations of classroom pedagogy according to the previous literature. The 50 syllabi were evaluated with a previously validated and peer-reviewed rubric designed by Cullen and Harris in 2009 and published in Assessment and Evaluation in Higher Education journal. The rubric measures the degree of learner-centeredness of syllabi. It has 13 items categorized under 3 factors plus the number of pages of the syllabi. We have modified the rubric to be on a 5-point scale (0-4). Zero represents the lowest degree of learner-centerednes..., , # How syllabi relate to outcomes in higher education: An evaluation of syllabi learner-centeredness and grade inequities in STEM

    https://doi.org/10.7280/D1NH6N

    DATA-SPECIFIC INFORMATION

    Eslami_2022_How_Syllabi_Relate_to_Outcomes_in_Higher_Education.csv

    Number of variables: 26

    Number of cases/rows: 50

    Variable names with descriptions and/or their values in parenthesis:

    small_opportunity_gap; delta_GP (average grade point difference between minoritized and non-minoritized students in a STEM course); additional_item_Length_of_Syllabus (number of pages for each syllabus);

    For the definition of rubric items, refer to the rubric designed by Cullen & Harris in 2009 published in Assessment and Evaluation in Higher Education journal; Rubric items are scored on a 5-point scale (0, 1, 2, 3, and 4): rubric_item_Accessibility_of_Teacher; rubric_item_Learning_Rationale; rubric_item_Collaboration; rubric_item_Teachers_Role; rubric_item_Stud...

  9. f

    Table 1_Integrated STEM for sustainability in school and early teacher...

    • figshare.com
    xlsx
    Updated Oct 16, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Amandyk Kopbossyn; Shakhislam Laiskhanov; BĂŒlent Aksoy; Aigul Tokbergenova; Mukhit Nametkulov; Assel Kozybakova (2025). Table 1_Integrated STEM for sustainability in school and early teacher education: a systematic review (2019–2025).xlsx [Dataset]. http://doi.org/10.3389/feduc.2025.1697058.s001
    Explore at:
    xlsxAvailable download formats
    Dataset updated
    Oct 16, 2025
    Dataset provided by
    Frontiers
    Authors
    Amandyk Kopbossyn; Shakhislam Laiskhanov; BĂŒlent Aksoy; Aigul Tokbergenova; Mukhit Nametkulov; Assel Kozybakova
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    This systematic review synthesizes research on school-focused initiatives that integrate science, technology, engineering, and mathematics (STEM) with sustainability goals, published between 2019 and 2025. Searches of Scopus, Web of Science, and SpringerLink, along with reference checks, identified 49 studies. We coded approaches, topics, technology use, outcomes, and implementation features. Of these 19 studies, 42 empirical interventions were mapped by topic and subject, while seven conceptual or non-anchored pieces were excluded from topic counts but were used for informed interpretation. Publications accelerated after 2020 and clustered in North America and Southeast/East Asia. Climate dominated the topic distributions, followed by water and circularity; biodiversity and energy were at moderate levels, while smaller clusters addressed disaster, built environment, and justice/policy. Technology integration was most prevalent in water and circularity units, moderate in disaster and built environment, and comparatively limited in climate; energy and justice/policy showed minimal technology integration. Outcome synthesis indicated broad gains from project-based and inquiry-oriented designs and from context/place-based approaches; socio-scientific argumentation most consistently advanced agency and values; modeling and engineering design excelled on skills and, with coherence supports, also improved concepts. A synthesized framework addresses key implementation challenges—curriculum fit, teacher capacity, cognitive load, assessment alignment, and equity logistics. The review offers design-ready guidance for selecting approaches that match desired learning and participation outcomes.

  10. Data from: arXiv Dataset

    • kaggle.com
    zip
    Updated Oct 7, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Cornell University (2023). arXiv Dataset [Dataset]. https://www.kaggle.com/dsv/6634444
    Explore at:
    zip(1289748687 bytes)Available download formats
    Dataset updated
    Oct 7, 2023
    Dataset authored and provided by
    Cornell University
    License

    https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

    Description

    About ArXiv

    For nearly 30 years, ArXiv has served the public and research communities by providing open access to scholarly articles, from the vast branches of physics to the many subdisciplines of computer science to everything in between, including math, statistics, electrical engineering, quantitative biology, and economics. This rich corpus of information offers significant, but sometimes overwhelming depth.

    In these times of unique global challenges, efficient extraction of insights from data is essential. To help make the arXiv more accessible, we present a free, open pipeline on Kaggle to the machine-readable arXiv dataset: a repository of 1.7 million articles, with relevant features such as article titles, authors, categories, abstracts, full text PDFs, and more.

    Our hope is to empower new use cases that can lead to the exploration of richer machine learning techniques that combine multi-modal features towards applications like trend analysis, paper recommender engines, category prediction, co-citation networks, knowledge graph construction and semantic search interfaces.

    The dataset is freely available via Google Cloud Storage buckets (more info here). Stay tuned for weekly updates to the dataset!

    ArXiv is a collaboratively funded, community-supported resource founded by Paul Ginsparg in 1991 and maintained and operated by Cornell University.

    The release of this dataset was featured further in a Kaggle blog post here.

    https://storage.googleapis.com/kaggle-public-downloads/arXiv.JPG" alt="">

    See here for more information.

    ArXiv On Kaggle

    Metadata

    This dataset is a mirror of the original ArXiv data. Because the full dataset is rather large (1.1TB and growing), this dataset provides only a metadata file in the json format. This file contains an entry for each paper, containing: - id: ArXiv ID (can be used to access the paper, see below) - submitter: Who submitted the paper - authors: Authors of the paper - title: Title of the paper - comments: Additional info, such as number of pages and figures - journal-ref: Information about the journal the paper was published in - doi: https://www.doi.org - abstract: The abstract of the paper - categories: Categories / tags in the ArXiv system - versions: A version history

    You can access each paper directly on ArXiv using these links: - https://arxiv.org/abs/{id}: Page for this paper including its abstract and further links - https://arxiv.org/pdf/{id}: Direct link to download the PDF

    Bulk access

    The full set of PDFs is available for free in the GCS bucket gs://arxiv-dataset or through Google API (json documentation and xml documentation).

    You can use for example gsutil to download the data to your local machine. ```

    List files:

    gsutil cp gs://arxiv-dataset/arxiv/

    Download pdfs from March 2020:

    gsutil cp gs://arxiv-dataset/arxiv/arxiv/pdf/2003/ ./a_local_directory/

    Download all the source files

    gsutil cp -r gs://arxiv-dataset/arxiv/ ./a_local_directory/ ```

    Update Frequency

    We're automatically updating the metadata as well as the GCS bucket on a weekly basis.

    License

    Creative Commons CC0 1.0 Universal Public Domain Dedication applies to the metadata in this dataset. See https://arxiv.org/help/license for further details and licensing on individual papers.

    Acknowledgements

    The original data is maintained by ArXiv, huge thanks to the team for building and maintaining this dataset.

    We're using https://github.com/mattbierbaum/arxiv-public-datasets to pull the original data, thanks to Matt Bierbaum for providing this tool.

  11. Employed persons with tertiary education in STEM fields by occupation...

    • autario.com
    csv, json
    Updated Aug 4, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Eurostat (2026). Employed persons with tertiary education in STEM fields by occupation (Eurostat) [Dataset]. https://autario.com/data/employed-persons-with-tertiary-education-in-stem-fields-by-occupation-eurostat
    Explore at:
    csv, jsonAvailable download formats
    Dataset updated
    Aug 4, 2026
    Dataset provided by
    Eurostathttp://ec.europa.eu/eurostat
    autario
    Authors
    Eurostat
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Area covered
    Worldwide
    Variables measured
    geo, freq, unit, isco08, dataflow, obs_flag, obs_value, conf_status, last_update, time_period, and 5 more
    Description

    This dataset tracks employed persons with tertiary education in science, technology, engineering, and mathematics fields, broken down by occupation across the European Union and member states. Understanding the distribution of STEM-educated talent across different occupations reveals how advanced economies allocate their highest-skilled human capital and identifies sectoral demand for specialized expertise.

    The dataset covers 28 European entities with 1,568 data records spanning from 2021 to 2024. All 28 entities, including Austria, Belgium, Bulgaria, Cyprus, and the Czech Republic, reported data through 2024. This comprehensive geographic and temporal scope enables both cross-country comparisons of STEM employment patterns and analysis of how occupational distribution shifted across the four-year period.

    Researchers and labor economists use this data to understand workforce composition, identify skills gaps, and forecast demand for STEM talent by occupation type. Policymakers reference occupational breakdowns to guide education and immigration strategies, while business analysts track employment trends to anticipate hiring needs in specialized technical roles. The data reveals whether STEM graduates concentrate in engineering and computing roles or distribute across broader professional occupations.

    This publicly available Eurostat resource supports evidence-based workforce planning, academic research on skills markets, and strategic business intelligence for companies competing for technical talent across Europe.

  12. f

    Data_Sheet_1_On the Design and Validation of Assessing Tools for Measuring...

    • frontiersin.figshare.com
    pdf
    Updated Jun 2, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    María Pilar Herce-Palomares; Carmen Botella-Mascarell; Esther de Ves; Emilia López-Iñesta; Anabel Forte; Xaro Benavent; Silvia Rueda (2023). Data_Sheet_1_On the Design and Validation of Assessing Tools for Measuring the Impact of Programs Promoting STEM Vocations.pdf [Dataset]. http://doi.org/10.3389/fpsyg.2022.937058.s001
    Explore at:
    pdfAvailable download formats
    Dataset updated
    Jun 2, 2023
    Dataset provided by
    Frontiers
    Authors
    María Pilar Herce-Palomares; Carmen Botella-Mascarell; Esther de Ves; Emilia López-Iñesta; Anabel Forte; Xaro Benavent; Silvia Rueda
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    This paper presents the design and validation process of a set of instruments to evaluate the impact of an informal learning initiative to promote Science, Technology, Engineering, and Mathematics (STEM) vocations in students, their families (parents), and teachers. The proposed set of instruments, beyond assessing the satisfaction of the public involved, allow collecting data to evaluate the impact in terms of changes in the consideration of the role of women in STEM areas and STEM vocations. The procedure followed to develop the set of instruments consisted of two phases. In the first phase, a preliminary version (v1) of the questionnaires was designed based on the objectives of the Girls4STEM initiative, an inclusive project promoting STEM vocations between 6 and 18 years old boys and girls. Five specific questionnaires were designed, one for the families (post activity), two for the students (pre and post activity) and two for the teachers (pre and post avitivity). A refined version (v2) of each questionnaire was obtained with evidence of content validity after undergoing an expert judgment process. The second phase was the refinement of the (v2) instruments, to ascertain the evidence of reliability and validity so that a final version (v3) was derived. In the paper, a high-quality set of good practices focused on promoting diversity and gender equality in the STEM sector are presented from a Higher Education Institution perspective, the University of Valencia. The main contribution of this work is the achievement of a set of instruments, rigorously designed for the evaluation of the implementation and effectiveness of a STEM promoting program, with sufficient validity evidence. Moreover, the proposed instruments can be a reference for the evaluation of other projects aimed at diversifying the STEM sector.

  13. Not seeing a result you expected?
    Learn how you can add new datasets to our index.

Share
FacebookFacebook
TwitterTwitter
Email
Click to copy link
Link copied
Close
Cite
Raja Biswas (2024). Wiki STEM Corpus [Dataset]. https://www.kaggle.com/datasets/conjuring92/wiki-stem-corpus
Organization logo

Wiki STEM Corpus

A subset of wikipedia focusing on STEM articles

Explore at:
zip(891581618 bytes)Available download formats
Dataset updated
Apr 12, 2024
Authors
Raja Biswas
License

Attribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically

Description

We created a STEM (Science, Technology, Engineering and Mathematics) corpus by filtering wikipedia articles based on their category metadata. During extraction of wiki page contents, we mitigated the frequent rendering issues (number, equations & symbols) prevalent in existing wiki datasets.

For filtering, we first defined a set of seed wikipedia categories related to STEM topics such as Category:Concepts in physics, Category:Physical quantities, etc. For each category, recursively collect the member pages and subcategories up to a certain depth. We next extracted the page contents of the collected wiki URLs using Wikipedia-API (400k+ pages).

Chunking: We first split the full text from each article based on different sections. The longer sections were further broken down into smaller chunks containing approximately 300 tokens (deberta-v3 tokenizer).

This dataset can be embedded and used for RAG over STEM wiki.

References: - Wiki STEM url collection: https://www.kaggle.com/code/conjuring92/d01-wiki-urls/notebook - Extraction of page content: https://www.kaggle.com/code/conjuring92/s04-stem-wiki-fetch - Chunking: https://www.kaggle.com/code/conjuring92/d504-chunking/notebook

Search
Clear search
Close search
Google apps
Main menu