Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
STEM Dataset
📃 [Paper] • 💻 [Github] • 🤗 [Dataset] • 🏆 [Leaderboard] • 📽 [Slides] • 📋 [Poster]
This dataset is proposed in the ICLR 2024 paper: Measuring Vision-Language STEM Skills of Neural Models. We introduce a new challenge to test the STEM skills of neural models. The problems in the real world often require solutions, combining knowledge from STEM (science, technology, engineering, and math). Unlike existing datasets, our dataset requires the understanding of… See the full description on the dataset page: https://huggingface.co/datasets/stemdataset/STEM.
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
Black scientists are major contributors to the advancement of Science, Technology, Engineering, and Mathematics (STEM). Yet, most of us know very little about these accomplishments. Here, we provide the first volume of the Atlas of Black Scholarship (A.B.S.) for inclusion to help science educators in the Life Sciences and Chemistry integrate the work of Black scientists into their curricula.
Facebook
TwitterAttribution-ShareAlike 3.0 (CC BY-SA 3.0)https://creativecommons.org/licenses/by-sa/3.0/
License information was derived automatically
The STEM ECR v1.0 dataset has been developed to provide a benchmark for the evaluation of scientific entity extraction, classification, and resolution tasks in a domain-independent fashion. It comprises annotations for scientific entities in scientific Abstracts drawn from 10 disciplines in Science, Technology, Engineering, and Medicine. The annotated entities are further grounded to Wikipedia and Wiktionary, respectively.
The dataset is organized in the following folders:
The annotation guidelines that supported the creation of this corpus can be found here.
D'Souza, J., Hoppe, A., Brack, A., Jaradeh, M., Auer, S., & Ewerth, R. (2020). The STEM-ECR Dataset: Grounding Scientific Entity References in STEM Scholarly Content to Authoritative Encyclopedic and Lexicographic Sources. In Proceedings of The 12th Language Resources and Evaluation Conference (pp. 2192–2203). European Language Resources Association.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
About ArXiv For nearly 30 years, ArXiv has served the public and research communities by providing open access to scholarly articles, from the vast branches of physics to the many subdisciplines of computer science to everything in between, including math, statistics, electrical engineering, quantitative biology, and economics. This rich corpus of information offers significant, but sometimes overwhelming depth.
In these times of unique global challenges, efficient extraction of insights from data is essential. To help make the arXiv more accessible, we present a free, open pipeline on Kaggle to the machine-readable arXiv dataset: a repository of 1.7 million articles, with relevant features such as article titles, authors, categories, abstracts, full text PDFs, and more.
Our hope is to empower new use cases that can lead to the exploration of richer machine learning techniques that combine multi-modal features towards applications like trend analysis, paper recommender engines, category prediction, co-citation networks, knowledge graph construction and semantic search interfaces.
ArXiv is a collaboratively funded, community-supported resource founded by **Paul Ginsparg **in 1991 and maintained and operated by Cornell University.
ArXiv On Kaggle Metadata This dataset is a mirror of the original ArXiv data. Because the full dataset is rather large (1.1TB and growing), this dataset provides only a metadata file in the json format. This file contains an entry for each paper, containing:
id: ArXiv ID (can be used to access the paper, see below) submitter: Who submitted the paper authors: Authors of the paper title: Title of the paper comments: Additional info, such as number of pages and figures journal-ref: Information about the journal the paper was published in doi: https://www.doi.org abstract: The abstract of the paper categories: Categories / tags in the ArXiv system versions: A version history
Facebook
TwitterAttribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically
We created a STEM (Science, Technology, Engineering and Mathematics) corpus by filtering wikipedia articles based on their category metadata. During extraction of wiki page contents, we mitigated the frequent rendering issues (number, equations & symbols) prevalent in existing wiki datasets.
For filtering, we first defined a set of seed wikipedia categories related to STEM topics such as Category:Concepts in physics, Category:Physical quantities, etc. For each category, recursively collect the member pages and subcategories up to a certain depth. We next extracted the page contents of the collected wiki URLs using Wikipedia-API (400k+ pages).
Chunking: We first split the full text from each article based on different sections. The longer sections were further broken down into smaller chunks containing approximately 300 tokens (deberta-v3 tokenizer).
This dataset can be embedded and used for RAG over STEM wiki.
References: - Wiki STEM url collection: https://www.kaggle.com/code/conjuring92/d01-wiki-urls/notebook - Extraction of page content: https://www.kaggle.com/code/conjuring92/s04-stem-wiki-fetch - Chunking: https://www.kaggle.com/code/conjuring92/d504-chunking/notebook
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This systematic review synthesizes research on school-focused initiatives that integrate science, technology, engineering, and mathematics (STEM) with sustainability goals, published between 2019 and 2025. Searches of Scopus, Web of Science, and SpringerLink, along with reference checks, identified 49 studies. We coded approaches, topics, technology use, outcomes, and implementation features. Of these 19 studies, 42 empirical interventions were mapped by topic and subject, while seven conceptual or non-anchored pieces were excluded from topic counts but were used for informed interpretation. Publications accelerated after 2020 and clustered in North America and Southeast/East Asia. Climate dominated the topic distributions, followed by water and circularity; biodiversity and energy were at moderate levels, while smaller clusters addressed disaster, built environment, and justice/policy. Technology integration was most prevalent in water and circularity units, moderate in disaster and built environment, and comparatively limited in climate; energy and justice/policy showed minimal technology integration. Outcome synthesis indicated broad gains from project-based and inquiry-oriented designs and from context/place-based approaches; socio-scientific argumentation most consistently advanced agency and values; modeling and engineering design excelled on skills and, with coherence supports, also improved concepts. A synthesized framework addresses key implementation challenges—curriculum fit, teacher capacity, cognitive load, assessment alignment, and equity logistics. The review offers design-ready guidance for selecting approaches that match desired learning and participation outcomes.
Facebook
TwitterFostering equity in undergraduate science, technology, engineering, and mathematics (STEM) programs can be accomplished by incorporating learner-centered pedagogies, resulting in the closing of opportunity gaps (defined in this research as the difference in grades earned by minoritized and non-minoritized students). We assessed STEM courses that exhibit small and large opportunity gaps at a minority-serving, research-intensive university, and evaluated the degree to which their syllabi are learner-centered, according to a previously validated rubric. We specifically chose syllabi as they are often the first interaction a student has with a course and can serve to establish expectations for course policies and practices. We found that STEM courses with more learner-centered syllabi had smaller opportunity gaps. The syllabus rubric factor that most correlated with smaller opportunity gaps was Power and Control, which reflects the Student's Role, Outside Resources, and Syllabus Focus. This..., This dataset is composed of rubric scores for 50 course syllabi of STEM classes in a research-intensive university with a large population of minoritized students as well as some institutional data (here defined as African-American, Latinx, Pacific Islander, and American Indian). We wanted to examine the relationship between racial grade gaps (here labeled as opportunity gaps) and the degree of learner-centeredness of the syllabi since course syllabi are good representations of classroom pedagogy according to the previous literature. The 50 syllabi were evaluated with a previously validated and peer-reviewed rubric designed by Cullen and Harris in 2009 and published in Assessment and Evaluation in Higher Education journal. The rubric measures the degree of learner-centeredness of syllabi. It has 13 items categorized under 3 factors plus the number of pages of the syllabi. We have modified the rubric to be on a 5-point scale (0-4). Zero represents the lowest degree of learner-centerednes..., , # How syllabi relate to outcomes in higher education: An evaluation of syllabi learner-centeredness and grade inequities in STEM
https://doi.org/10.7280/D1NH6N
Number of variables: 26
Number of cases/rows: 50
Variable names with descriptions and/or their values in parenthesis:
small_opportunity_gap; delta_GP (average grade point difference between minoritized and non-minoritized students in a STEM course); additional_item_Length_of_Syllabus (number of pages for each syllabus);
For the definition of rubric items, refer to the rubric designed by Cullen & Harris in 2009 published in Assessment and Evaluation in Higher Education journal; Rubric items are scored on a 5-point scale (0, 1, 2, 3, and 4): rubric_item_Accessibility_of_Teacher; rubric_item_Learning_Rationale; rubric_item_Collaboration; rubric_item_Teachers_Role; rubric_item_Stud...
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
For nearly 30 years, ArXiv has served the public and research communities by providing open access to scholarly articles, from the vast branches of physics to the many subdisciplines of computer science to everything in between, including math, statistics, electrical engineering, quantitative biology, and economics. This rich corpus of information offers significant, but sometimes overwhelming depth.
In these times of unique global challenges, efficient extraction of insights from data is essential. To help make the arXiv more accessible, we present a free, open pipeline on Kaggle to the machine-readable arXiv dataset: a repository of 1.7 million articles, with relevant features such as article titles, authors, categories, abstracts, full text PDFs, and more.
Our hope is to empower new use cases that can lead to the exploration of richer machine learning techniques that combine multi-modal features towards applications like trend analysis, paper recommender engines, category prediction, co-citation networks, knowledge graph construction and semantic search interfaces.
The dataset is freely available via Google Cloud Storage buckets (more info here). Stay tuned for weekly updates to the dataset!
ArXiv is a collaboratively funded, community-supported resource founded by Paul Ginsparg in 1991 and maintained and operated by Cornell University.
The release of this dataset was featured further in a Kaggle blog post here.
https://storage.googleapis.com/kaggle-public-downloads/arXiv.JPG" alt="">
See here for more information.
This dataset is a mirror of the original ArXiv data. Because the full dataset is rather large (1.1TB and growing), this dataset provides only a metadata file in the json format. This file contains an entry for each paper, containing:
- id: ArXiv ID (can be used to access the paper, see below)
- submitter: Who submitted the paper
- authors: Authors of the paper
- title: Title of the paper
- comments: Additional info, such as number of pages and figures
- journal-ref: Information about the journal the paper was published in
- doi: https://www.doi.org
- abstract: The abstract of the paper
- categories: Categories / tags in the ArXiv system
- versions: A version history
You can access each paper directly on ArXiv using these links:
- https://arxiv.org/abs/{id}: Page for this paper including its abstract and further links
- https://arxiv.org/pdf/{id}: Direct link to download the PDF
The full set of PDFs is available for free in the GCS bucket gs://arxiv-dataset or through Google API (json documentation and xml documentation).
You can use for example gsutil to download the data to your local machine. ```
gsutil cp gs://arxiv-dataset/arxiv/
gsutil cp gs://arxiv-dataset/arxiv/arxiv/pdf/2003/ ./a_local_directory/
gsutil cp -r gs://arxiv-dataset/arxiv/ ./a_local_directory/ ```
We're automatically updating the metadata as well as the GCS bucket on a weekly basis.
Creative Commons CC0 1.0 Universal Public Domain Dedication applies to the metadata in this dataset. See https://arxiv.org/help/license for further details and licensing on individual papers.
The original data is maintained by ArXiv, huge thanks to the team for building and maintaining this dataset.
We're using https://github.com/mattbierbaum/arxiv-public-datasets to pull the original data, thanks to Matt Bierbaum for providing this tool.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This paper presents the design and validation process of a set of instruments to evaluate the impact of an informal learning initiative to promote Science, Technology, Engineering, and Mathematics (STEM) vocations in students, their families (parents), and teachers. The proposed set of instruments, beyond assessing the satisfaction of the public involved, allow collecting data to evaluate the impact in terms of changes in the consideration of the role of women in STEM areas and STEM vocations. The procedure followed to develop the set of instruments consisted of two phases. In the first phase, a preliminary version (v1) of the questionnaires was designed based on the objectives of the Girls4STEM initiative, an inclusive project promoting STEM vocations between 6 and 18 years old boys and girls. Five specific questionnaires were designed, one for the families (post activity), two for the students (pre and post activity) and two for the teachers (pre and post avitivity). A refined version (v2) of each questionnaire was obtained with evidence of content validity after undergoing an expert judgment process. The second phase was the refinement of the (v2) instruments, to ascertain the evidence of reliability and validity so that a final version (v3) was derived. In the paper, a high-quality set of good practices focused on promoting diversity and gender equality in the STEM sector are presented from a Higher Education Institution perspective, the University of Valencia. The main contribution of this work is the achievement of a set of instruments, rigorously designed for the evaluation of the implementation and effectiveness of a STEM promoting program, with sufficient validity evidence. Moreover, the proposed instruments can be a reference for the evaluation of other projects aimed at diversifying the STEM sector.
Facebook
TwitterAttribution-NonCommercial-ShareAlike 4.0 (CC BY-NC-SA 4.0)https://creativecommons.org/licenses/by-nc-sa/4.0/
License information was derived automatically
https://i.imgur.com/fbinH6B.jpg" alt="visa">
This program consists of 3 visa programs : H-1B, H-1B1 and E-3
For more information about individual columns, refer the column metadata. A detailed description of the underlying raw datasets is available in official site.
The Office of Foreign Labor Certification (OFLC) is in charge of compiling programme statistics, which includes information on H-1B, H-1B1 and E-3 visas. Quarterly, the disclosure data is updated and made available online.
The raw data available contains private data of some kind. As a result, all of the private data has been masked or dropped, and various transformations have been applied to make the data more available for speedy investigation.
For data before 2017, please visit H-1B Visa Petitions 2011-2016
Keywords : Time Series Analysis, Survey Analysis, Statistical Analysis, Exploratory Data Analysis, Human Rights, Tabular Data, History, Data Cleaning, Income, Data Visualization, Data Analytics, Business, United States, Education, Travel
Not seeing a result you expected?
Learn how you can add new datasets to our index.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically