Facebook
Twitterfarsi-asr/ganjoor-dataset dataset hosted on Hugging Face and contributed by the HF Datasets community
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
The application of Artificial Intelligence (AI) has been evident in the agricultural sector recently. The main goal of AI in agriculture is to improve crop yield, control crop pests/diseases, and reduce cost. The agricultural sector in developing countries faces severe in the form of disease and pest infestation, the knowledge gap between farmers and technology, and a lack of storage facilities, among others. To help address some of these challenges, this work presents crop pests/disease datasets sourced from local farms in Ghana. The dataset is presented in two folds; the raw images which consists of 24,881 images ( 6,549-Cashew, 7,508-Cassava, 5,389-Maize, and 5,435-Tomato) and augmented images which is further split into train and test set consists of 102,976 images (25,811-Cashew, 26,330-Cassava, 23,657-Maize, and 27,178-Tomato), categorized into 22 classes. All images are de-identified, validated by expert plant virologists, and freely available for use by the research community.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
Copernicus Dataset
A curated, multi-domain dataset for tokenizer training and small language-model pretraining, covering natural language, source code, and mathematics. The dataset was created as the training corpus for the Copernicus Tokenizer, with the goal of providing broad token and pattern coverage rather than relying on a single text domain.
Dataset Overview
Property Details
Total size ~12.3 GB
Domains NLP, Code, Mathematics
Format Parquet… See the full description on the dataset page: https://huggingface.co/datasets/Nj-1111/Copernicus-Dataset.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
Addv2_v1 is a dataset for object detection tasks - it contains Can Foam Plastic Plasticbottle U YgPQ annotations for 1,461 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
TwitterIBM Employee Salary Dataset: is a cleaned, analyzed dataset now you need to visualize the data and predict accordingly how much IBM pays its employees per month, per year or per day etc. I am a beginner at data analysis but have worked sincerely on this data.
Facebook
TwitterAttribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically
Multi-file Spotify charts dataset built from the scraper pipeline. Includes daily songs, daily artists, weekly albums, normalized entity tables, Spotify monthly listener history, artwork, and external links.
Files currently shipped:
- charts_songs_daily.csv.gz
- charts_artists_daily.csv.gz
- charts_albums_weekly.csv.gz
- artist_listeners_daily.csv
- songs.csv
- artists.csv
- albums.csv
- artwork.csv
- links.csv
Updated by the automated pipeline.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This dataset contains 10,000 samples designed for drone navigation and obstacle avoidance research. It includes RGB images (320x320), Depth maps (320x320), and corresponding Commands (vx, vy, vz, yaw_rate). The data was collected in AirSim, a realistic drone simulator by Microsoft, using a drone controlled by a script implementing potential fields for navigation and obstacle avoidance.
This dataset is ideal for researchers and developers working on autonomous drone navigation, computer vision, or robotics projects involving RGB and Depth data.
rgb/: Directory containing 10,000 RGB images (e.g., 000000.png, ..., 009999.png)depth/: Directory containing 10,000 Depth maps as NumPy arrays (e.g., 000000.npy, ..., 009999.npy)commands/: Directory containing 10,000 Commands as NumPy arrays (e.g., 000000.npy, ..., 009999.npy), each file with 4 values: vx, vy, vz, yaw_rateThis dataset is suitable for: - Developing models for autonomous drone navigation - Research in obstacle avoidance and path planning - Computer vision tasks involving RGB and Depth data - Robotics and simulation-based studies
Example use case: Use the RGB and Depth data to develop algorithms for real-time obstacle avoidance in drones.
This dataset is licensed under CC BY 4.0. You are free to use, modify, and distribute it as long as you provide attribution to the author and acknowledge the source of the data: - Attribution: "Dataset DroneFlight_Obs_AvoidanceAirSimRGBDepth10k_320x320 by https://www.kaggle.com/lukpellant, data generated using AirSim (MIT License)." - AirSim License: The data was collected in AirSim, which is licensed under the MIT License (https://github.com/microsoft/AirSim).
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
This dataset was created by Lê Đức Tùng Dương
Released under Apache 2.0
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
This dataset contains the digitized treatments in Plazi based on the original journal article Ng, Peter K. L., Bouchet, Philippe (2015): Actaea grimaldii, a new species of reef crab from Papua New Guinea (Crustacea, Brachyura, Xanthidae). European Journal of Taxonomy 140: 1-18, DOI: 10.5852/ejt.2015.140
Abstract. A new species of xanthid crab, Actaea grimaldii, is described from the coral reefs of Papua New Guinea. This species has a distinctive red and white coloration and is closest to Actaea spinosissima Borradaile, 1902, from the Indian Ocean. However, the new species can be distinguished by the arrangement of spines on the carapace, chelipeds and ambulatory legs, and the structure of the male gonopods. Actaea grimaldii sp. nov. has also been confused with A. polyacantha (Heller, 1861), but differs markedly in the carapace armature.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset contains detailed information on 2000 learners enrolled in English learning programs. It includes learner profiles, study habits, language skill levels, and cognitive abilities. Key attributes cover initial English proficiency, daily study time, study frequency, mastery in vocabulary, grammar, listening, speaking, reading, and writing, as well as cognitive skills such as memorization, understanding, application, analysis, evaluation, and creation.
The dataset also includes a target column representing the recommended learning path for each learner, helping to plan individualized learning sequences based on their current abilities and progress. It is suitable for research and development in adaptive, learner-centered English education systems.
Columns Description:
LearnerID – A unique identifier assigned to each learner for tracking purposes.
InitialEnglishLevel – The learner’s starting English proficiency level (Beginner, Elementary, Intermediate, Upper-Intermediate, Advanced).
DailyStudyTime – Number of hours the learner spends studying English per day.
StudyFrequency – Number of days per week the learner engages in English study.
VocabularyMastery – Learner’s proficiency in English vocabulary, scored from 0 to 100.
GrammarMastery – Learner’s proficiency in English grammar, scored from 0 to 100.
ListeningComprehension – Learner’s ability to understand spoken English, scored from 0 to 100.
OralExpression – Learner’s ability to speak English effectively, scored from 0 to 100.
ReadingAbility – Learner’s ability to comprehend written English, scored from 0 to 100.
WritingAbility – Learner’s ability to write in English clearly and correctly, scored from 0 to 100.
Memorization – Cognitive skill level in remembering information, scored from 0 (low) to 5 (high).
Understanding – Cognitive skill level in comprehending concepts, scored from 0 to 5.
Application – Cognitive skill level in applying knowledge to tasks, scored from 0 to 5.
Analysis – Cognitive skill level in analyzing and interpreting information, scored from 0 to 5.
Evaluation – Cognitive skill level in evaluating or judging information, scored from 0 to 5.
Creation – Cognitive skill level in creating or producing new ideas or content, scored from 0 to 5.
LearningPathCategory – Recommended learning path based on the learner’s current skills and mastery, indicating the next focus level (Beginner_Path, Elementary_Path, Intermediate_Path, UpperIntermediate_Path, Advanced_Path).
Facebook
TwitterThis dataset was created by Ahmed Hamada
Released under Data files © Original Authors
It contains the following files:
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
VLM-3R Training Data
Training QA data for VLM-3R: vsibench_train/ (VSI-Bench-style tasks) and vstibench_train/ (VSTI-Bench tasks over ScanNet train split).
Erratum (2026-07-13): corrected camera-position ground truth
A bug in the QA generation pipeline (reported by Jacob Yeung, CMU) extracted the camera center from camera-to-world poses using -R.T @ t instead of pose[:3, 3]. Answers in five vstibench_train files depended on the camera's world position and have… See the full description on the dataset page: https://huggingface.co/datasets/Journey9ni/VLM-3R-DATA.
Facebook
TwitterThis dataset was created by Aryan Khandal
Facebook
TwitterVegAnn Dataset
Vegetation Annotation of a large multi-crop RGB Dataset acquired under diverse conditions for image semantic segmentation
Keypoints ⏳
VegAnn contains 3775 images Images are 512*512 pixels Corresponding binary masks is 0 for soil + crop residues (background) 255 for Vegetation (foreground) The dataset includes images of 26+ crop species, which are not evenly represented VegAnn was compiled using a variety of outdoor images captured with… See the full description on the dataset page: https://huggingface.co/datasets/simonMadec/VegAnn.
Facebook
Twitterstarkosae/test-dataset dataset hosted on Hugging Face and contributed by the HF Datasets community
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
DamageR is a dataset for object detection tasks - it contains DamageR annotations for 480 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
Tugas Akhir Variasi 500 is a dataset for object detection tasks - it contains Potato Disease DRNC annotations for 493 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
TwitterSave Mart Supermarkets location dataset — United States subset. Verified addresses, coordinates, and phones. Licensed via CREHQ Data Store.
Facebook
TwitterThis is the edited sequence set for the data in the cowbird dataset. Due to file size limitations, these are only the unique sequences. For the unedited dataset, please contact the authors of the manuscript: Hird et al. 2014. Sampling locality is more detectable than taxonomy or ecology in the brood-parasitic Brown-Headed Cowbird. PeerJ. DOI:10.7717/peerj.321. Available: https://peerj.com/articles/321/
Facebook
TwitterThe EMITL2BCO2ENH Version 1 data product was decommissioned on March 26, 2026. Users are encouraged to use the EMITL2BCO2ENH Version 2 data product. The Earth Surface Mineral Dust Source Investigation (EMIT) instrument measures surface mineralogy, targeting the Earth’s arid dust source regions. EMIT is installed on the International Space Station (ISS) and uses imaging spectroscopy to take measurements of the sunlit regions of interest between 52° N latitude and 52° S latitude. An interactive map showing the regions being investigated, current and forecasted data coverage, and additional data resources can be found on the VSWIR Imaging Spectroscopy Interface for Open Science (VISIONS) EMIT Open Data Portal. In addition to its primary objective described above, EMIT has demonstrated the capacity to characterize carbon dioxide (CO2) and methane (CH4) point-source emissions by measuring gas absorption features in the short-wave infrared bands. The EMIT Level 2B Greenhouse Gas (GHG) series of products can be used to identify and quantify point source emissions. The EMIT Level 2B Carbon Dioxide Enhancement Data (EMITL2BCO2ENH) Version 1 data product is a total vertical column enhancement estimate of CO2 in parts per million meter (ppm m) based on an adaptive matched filter approach. EMITL2BCO2ENH provides per-pixel CO2 enhancement data used to identify CO2 plume complexes. The initial release of the EMITL2BCO2ENH data product will only include granules where CO2 plume complexes have been identified. Each granule contains one Cloud Optimized GeoTIFF (COG) file at a spatial resolution of 60 meters (m): Carbon Dioxide Enhancement (EMIT_L2B_CO2ENH). The EMITL2BCO2ENH COG file contains methane enhancement data based primarily on EMITL1BRAD radiance values. Each granule is approximately 75 kilometers (km) by 75 km, nominal at the equator, with some granules near the end of an orbit segment reaching 150 km in length. Known Issues Data acquisition gap: From September 13, 2022, through January 6, 2023, a power issue outside of EMIT caused a pause in operations. Due to this shutdown, no data were acquired during that timeframe.
Facebook
Twitterfarsi-asr/ganjoor-dataset dataset hosted on Hugging Face and contributed by the HF Datasets community