Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
Copernicus Dataset
A curated, multi-domain dataset for tokenizer training and small language-model pretraining, covering natural language, source code, and mathematics. The dataset was created as the training corpus for the Copernicus Tokenizer, with the goal of providing broad token and pattern coverage rather than relying on a single text domain.
Dataset Overview
Property Details
Total size ~12.3 GB
Domains NLP, Code, Mathematics
Format Parquet… See the full description on the dataset page: https://huggingface.co/datasets/Nj-1111/Copernicus-Dataset.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
The application of Artificial Intelligence (AI) has been evident in the agricultural sector recently. The main goal of AI in agriculture is to improve crop yield, control crop pests/diseases, and reduce cost. The agricultural sector in developing countries faces severe in the form of disease and pest infestation, the knowledge gap between farmers and technology, and a lack of storage facilities, among others. To help address some of these challenges, this work presents crop pests/disease datasets sourced from local farms in Ghana. The dataset is presented in two folds; the raw images which consists of 24,881 images ( 6,549-Cashew, 7,508-Cassava, 5,389-Maize, and 5,435-Tomato) and augmented images which is further split into train and test set consists of 102,976 images (25,811-Cashew, 26,330-Cassava, 23,657-Maize, and 27,178-Tomato), categorized into 22 classes. All images are de-identified, validated by expert plant virologists, and freely available for use by the research community.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
G1_football is a dataset for object detection tasks - it contains G1 Football annotations for 303 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
Twitterstarkosae/test-dataset dataset hosted on Hugging Face and contributed by the HF Datasets community
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
Tugas Akhir Variasi 500 is a dataset for object detection tasks - it contains Potato Disease DRNC annotations for 493 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
Twitterhttps://baristalife.co/pages/caffeine-datahttps://baristalife.co/pages/caffeine-data
Verified caffeine content for 154 drinks across coffee, tea, energy drinks, soda, ready to drink, chocolate, dessert, and decaf, with serving size, caffeine per ounce, and a cited source on every row. Updated quarterly.
Facebook
TwitterAttribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically
Multi-file Spotify charts dataset built from the scraper pipeline. Includes daily songs, daily artists, weekly albums, normalized entity tables, Spotify monthly listener history, artwork, and external links.
Files currently shipped:
- charts_songs_daily.csv.gz
- charts_artists_daily.csv.gz
- charts_albums_weekly.csv.gz
- artist_listeners_daily.csv
- songs.csv
- artists.csv
- albums.csv
- artwork.csv
- links.csv
Updated by the automated pipeline.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Test dataset for a manuscript: "Parvalbumin-expressing basal forebrain neurons mediate learning from negative experience" from Hegedüs et al (https://doi.org/10.1101/2023.03.31.535018).
This dataset contains a folder structure with tetrode recording, behavioral and fiber photometry data to test functions in https://github.com/hangyabalazs/PV_Pavlovian_analysis GitHub repository
Facebook
TwitterThis database, compiled by Matthews and Fung (1987), provides information on the distribution and environmental characteristics of natural wetlands. The database was developed to evaluate the role of wetlands in the annual emission of methane from terrestrial sources. The original data consists of five global 1-degree latitude by 1-degree longitude arrays. This subset, for the study area of the Large Scale Biosphere-Atmosphere Experiment in Amazonia (LBA) in South America, retains all five arrays at the 1-degree resolution but only for the area of interest (i.e., longitude 85 deg to 30 deg W, latitude 25 deg S to 10 deg N). The arrays are (1) wetland data source, (2) wetland type, (3) fractional inundation, (4) vegetation type, and (5) soil type. The data subsets are in both ASCII GRID and binary image file formats.The data base is the result of the integration of three independent digital sources: (1) vegetation classified according to the United Nations Educational Scientific and Cultural Organization (UNESCO) system (Matthews, 1983), (2) soil properties from the Food and Agriculture Organization (FAO) soil maps (Zobler, 1986), and (3) fractional inundation in each 1-degree cell compiled from a global map survey of Operational Navigation Charts (ONC). With vegetation, soil, and inundation characteristics of each wetland site identified, the data base has been used for a coherent and systematic estimate of methane emissions from wetlands and for an analysis of the causes for uncertainties in the emission estimate.The complete global data base is available from NASA/GISS [http://www.giss.nasa.gov] and NCAR data set ds765.5 [http://www.ncar.ucar.edu]; the global vegetation types data are available from ORNL DAAC [http://www.daac.ornl.gov].
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
VLM-3R Training Data
Training QA data for VLM-3R: vsibench_train/ (VSI-Bench-style tasks) and vstibench_train/ (VSTI-Bench tasks over ScanNet train split).
Erratum (2026-07-13): corrected camera-position ground truth
A bug in the QA generation pipeline (reported by Jacob Yeung, CMU) extracted the camera center from camera-to-world poses using -R.T @ t instead of pose[:3, 3]. Answers in five vstibench_train files depended on the camera's world position and have… See the full description on the dataset page: https://huggingface.co/datasets/Journey9ni/VLM-3R-DATA.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
DamageR is a dataset for object detection tasks - it contains DamageR annotations for 480 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
This dataset contains the digitized treatments in Plazi based on the original journal article Sawyer, Roy T. (2019): Observations on the terrestrial leech Haemopis septagon Sawyer & Shelley, 1976 (Annelida: Hirudinea) from the Outer Banks, North Carolina, USA, with a revision of the species. Zootaxa 4658 (2): 275-296, DOI: 10.11646/zootaxa.4658.2.4
Abstract
The terrestrial leech Haemopis septagon Sawyer & Shelley, 1976, is indigenous to the Great Dismal Swamp and environs of northeastern North Carolina and southeastern Virginia, USA. Ever since its discovery in 1895 at Lake Drummond in the Dismal Swamp, this elusive species has been recognized as taxonomically aberrant. For example, it is the only jawed leech in the United States with seven annuli between gonopores, and the only one with sixteen complete (5-annulate) segments, both highly conserved characters in the Hirudinidae.
The discovery of two populations of H. septagon in the Albemarle Peninsula in the Outer Banks region of North Carolina afforded an opportunity to investigate the taxonomy and biology of this inadequately characterized species. Its description in this study is the first comprehensive account of the external and internal anatomy of this species since its incomplete original description in 1976. This study is also an opportunity to correct errors in the incomplete original description, and to elucidate morphological and developmental variability of taxonomic significance. Evidence is presented for the first time of a possible aquatic or semi-aquatic form of H. septagon.
These Albemarle individuals were compared to the holotype from Durham County, NC, specimens from southeastern Virginia and a terrestrial leech recently reported from southern New Jersey. All of these fall within the variability demonstrated in this study for the Albemarle populations, and are accordingly recognized as the same species, H. septagon. Consequentially, Haemopis ottorum Wirchansky & Shain, 2010, is recognized as a junior synonym of Haemopis septagon.
Facebook
TwitterThis is the edited sequence set for the data in the cowbird dataset. Due to file size limitations, these are only the unique sequences. For the unedited dataset, please contact the authors of the manuscript: Hird et al. 2014. Sampling locality is more detectable than taxonomy or ecology in the brood-parasitic Brown-Headed Cowbird. PeerJ. DOI:10.7717/peerj.321. Available: https://peerj.com/articles/321/
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
## Overview
Smoke & Fire Dec is a dataset for object detection tasks - it contains Smoke Fire CTmX annotations for 5,190 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [MIT license](https://creativecommons.org/licenses/MIT).
Facebook
TwitterThe development plan (BPL) contains the legally binding stipulations for the urban development order. In principle, the development plan must be developed from the land use plan. The available data are the development plan ‘Bei der Schule, 1st Amendment’ of the municipality of Bretzfeld from XPlanung 5.0. Description: Development Plan At School, 1st Amendment.
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
Dataset Card for Synthetic Inline Holographical Images v3 (224px Highly Diverse)
This dataset provides synthetic image triplets representing inline holographical imaging in a simulated environment. This version (v3) uses a native 224x224 resolution optimized for modern Vision Transformers (ViT, Swin) and contains 25,000 samples across 8 noise configurations. Each data sample consists of:
An object-domain field (ground truth), Its corresponding forward-propagated hologram (the… See the full description on the dataset page: https://huggingface.co/datasets/gokhankocmarli/inline-digital-holography-v3.
Facebook
Twitter100 real-world household objects. Tasks: Cross-sensory retrieval, Contact localization, Material classification, Reconstruction, Manipulation benchmarks. The official ObjectFolder site provides dataset and benchmark download routes.
Facebook
TwitterFirst Choice Financial location dataset — United States subset. Verified addresses, coordinates, and phones. Licensed via CREHQ Data Store.
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
EngDesign Dataset
A dataset of 101 structured engineering design prompts across multiple domains.
Facebook
TwitterOpen Government Licence 3.0http://www.nationalarchives.gov.uk/doc/open-government-licence/version/3/
License information was derived automatically
Young people's open spaces (play areas) in York. For further information please visit City of York Council's website. *Please note that the data published within this dataset is a live API link to CYC's GIS server. Any changes made to the master copy of the data will be immediately reflected in the resources of this dataset.The date shown in the "Last Updated" field of each GIS resource reflects when the data was first published.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
Copernicus Dataset
A curated, multi-domain dataset for tokenizer training and small language-model pretraining, covering natural language, source code, and mathematics. The dataset was created as the training corpus for the Copernicus Tokenizer, with the goal of providing broad token and pattern coverage rather than relying on a single text domain.
Dataset Overview
Property Details
Total size ~12.3 GB
Domains NLP, Code, Mathematics
Format Parquet… See the full description on the dataset page: https://huggingface.co/datasets/Nj-1111/Copernicus-Dataset.