Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
VLM-3R Training Data
Training QA data for VLM-3R: vsibench_train/ (VSI-Bench-style tasks) and vstibench_train/ (VSTI-Bench tasks over ScanNet train split).
Erratum (2026-07-13): corrected camera-position ground truth
A bug in the QA generation pipeline (reported by Jacob Yeung, CMU) extracted the camera center from camera-to-world poses using -R.T @ t instead of pose[:3, 3]. Answers in five vstibench_train files depended on the camera's world position and have… See the full description on the dataset page: https://huggingface.co/datasets/Journey9ni/VLM-3R-DATA.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
The "Plant Pathogen Dataset" is a comprehensive collection of labeled images depicting various types of pathogens affecting plant species. This dataset is curated to facilitate research and development in the field of plant pathology, enabling the development of machine learning models for automated disease diagnosis and monitoring.
Image Categories: The dataset contains images representing different types of plant diseases, including bacterial infections, fungal diseases, pest infestations, and viral infections.
The images in this dataset were sourced from various sources, including research institutions, agricultural organizations, and open-access repositories. Care was taken to ensure high-quality images with accurate disease annotations.
Disease Diagnosis: The dataset can be used to train machine learning models for automated diagnosis of plant diseases based on image analysis. Disease Monitoring: By continuously monitoring plant health using machine learning models trained on this dataset, farmers and agricultural professionals can detect diseases early and implement timely interventions.
We would like to acknowledge the contributions of the research community, agricultural experts, and dataset contributors who have made this dataset possible. Their efforts in collecting, labeling, and sharing plant disease images are invaluable to advancing research in plant pathology and agricultural technology.
Facebook
TwitterThis dataset was created by Singh Prince Rinku
Released under Other (specified in description)
Facebook
TwitterThis repo consists of the datasets used for the TaCo paper. There are four datasets:
Multilingual Alpaca-52K GPT-4 dataset Multilingual Dolly-15K GPT-4 dataset TaCo dataset Multilingual Vicuna Benchmark dataset
We translated the first three datasets using Google Cloud Translation. The TaCo dataset is created by using the TaCo approach as described in our paper, combining the Alpaca-52K and Dolly-15K datasets. If you would like to create the TaCo dataset for a specific language, you can… See the full description on the dataset page: https://huggingface.co/datasets/saillab/taco-datasets.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset was created by SeshuRaju 🧘♂️
Released under CC0: Public Domain
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
492_ShoeBrandClassification is a dataset for object detection tasks - it contains Shoes annotations for 693 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
Twitterhttps://geolocet.com/pages/terms-of-usehttps://geolocet.com/pages/terms-of-use
Complete dataset containing all 17734 supermarkets in Belgium, including branded and independent locations. Includes geocoded addresses, latitude and longitude coordinates, contact details, opening hours, and administrative areas in CSV format for retail analysis, market research, logistics, and geospatial applications. Last updated: 26 May 2026.
Facebook
TwitterThe First National Bank of Elmer location dataset — United States subset. Verified addresses, coordinates, phones, and operating hours. Licensed via CREHQ Data Store.
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
This dataset contains the digitized treatments in Plazi based on the original journal article Heffern, Daniel, Santos-Silva, Antonio (2025): New species of Elaphidiini (Coleoptera, Cerambycidae, Cerambycinae) from Mexico and Central America, and new records in Cerambycidae and Disteniidae. Zootaxa 5569 (2): 231-252, DOI: 10.11646/zootaxa.5569.2.2, URL: https://doi.org/10.11646/zootaxa.5569.2.2
Abstract
Five new species are described in Elaphidiini: Aneflomorpha spinifera sp. nov., from Mexico (Jalisco); Aneflomorpha makra sp. nov., from Mexico (Oaxaca); Aneflomorpha guatemalana sp. nov., from Guatemala; Aneflomorpha gracilenta sp. nov., from Mexico (Oaxaca); and Aneflus macraei sp. nov., from Mexico (Oaxaca). Additionally, new records are provided for: Susuacanga marcelae Botero, 2015 (Cerambycidae, Cerambycinae, Eburiini), with the female illustrated for the first time; Ameriphoderes cribricollis (Bates, 1892) (Cerambycidae, Cerambycinae, Rhinotragini); Spinestoloides hefferni Santos-Silva, Wappes & Galileo, 2018 (Cerambycidae, Lamiinae, Desmiphorini), with the male illustrated for the first time; and Novantinoe solisi Santos-Silva & Hovore, 2007 (Disteniidae, Disteniinae).
Facebook
TwitterHome Depot location dataset — Mexico subset. Verified addresses, coordinates, phones, and operating hours. Licensed via CREHQ Data Store.
Facebook
TwitterThe Allen Brain Observatory – Visual Coding is a large-scale, standardized survey of physiological activity across the mouse visual cortex, hippocampus, and thalamus. It includes datasets collected with both two-photon imaging and Neuropixels probes, two complementary techniques for measuring the activity of neurons in vivo. The two-photon imaging dataset features visually evoked calcium responses from GCaMP6-expressing neurons in a range of cortical layers, visual areas, and Cre lines. The Neuropixels dataset features spiking activity from distributed cortical and subcortical brain regions, collected under analogous conditions to the two-photon imaging experiments. We hope that experimentalists and modelers will use these comprehensive, open datasets as a testbed for theories of visual information processing.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
CWHR species range datasets represent the maximum current geographic extent of each species within California. Ranges were originally delineated at a scale of 1:5,000,000 by species-level experts more than 30 years ago and have gradually been revised at a scale of 1:1,000,000. Species occurrence data are used in defining species ranges, but range polygons may extend beyond the limits of extant occurrence data for a particular species. When drawing range boundaries, CDFW seeks to err on the side of commission rather than omission. This means that CDFW may include areas within a range based on expert knowledge or other available information, despite an absence of confirmed occurrences, which may be due to a lack of survey effort. The degree to which a range polygon is extended beyond occurrence data will vary among species, depending upon each species’ vagility, dispersal patterns, and other ecological and life history factors. The boundary line of a range polygon is drawn with consideration of these factors and is aligned with standardized boundaries including watersheds (NHD), ecoregions (USDA), or other ecologically meaningful delineations such as elevation contour lines. While CWHR ranges are meant to represent the current range, once an area has been designated as part of a species’ range in CWHR, it will remain part of the range even if there have been no documented occurrences within recent decades. An area is not removed from the range polygon unless experts indicate that it has not been occupied for a number of years after repeated surveys or is deemed no longer suitable and unlikely to be recolonized. It is important to note that range polygons typically contain areas in which a species is not expected to be found due to the patchy configuration of suitable habitat within a species’ range. In this regard, range polygons are coarse generalizations of where a species may be found. This data is available for download from the CDFW website: https://www.wildlife.ca.gov/Data/CWHR. The following data sources were collated for the purposes of range mapping and species habitat modeling by RADMAP. Each focal taxon’s location data was extracted (when applicable) from the following list of sources. BIOS datasets are bracketed with their “ds” numbers and can be located on CDFW’s BIOS viewer: https://wildlife.ca.gov/Data/BIOS. California Natural Diversity Database, Terrestrial Species Monitoring [ds2826], North American Bat Monitoring Data Portal, VertNet, Breeding Bird Survey, Wildlife Insights, eBird, iNaturalist, other available CDFW or partner data.
Facebook
TwitterGDPa1: Antibody developability dataset
Contains the assay data for 242 antibodies across 10 assays as described in our latest preprint, PROPHET-Ab: A high-throughput platform for biophysical antibody developability assessment to enable AI/ML model training.
Example usage
Using pandas: import pandas as pd
huggingface-cli login to access this datasetdf = pd.read_csv("hf://datasets/ginkgo-datapoints/GDPa1/GDPa1_v1.2_20250814.csv")
Using Hugging… See the full description on the dataset page: https://huggingface.co/datasets/ginkgo-datapoints/GDPa1.
Facebook
Twitterfarsi-asr/ganjoor-dataset dataset hosted on Hugging Face and contributed by the HF Datasets community
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
This dataset contains the digitized treatments in Plazi based on the original journal article Pei, Rui, Li, Quanguo, Meng, Qingjin, Norell, Mark A., Gao, Ke-Qin (2017): New Specimens Of Anchiornis Huxleyi (Theropoda: Paraves) From The Late Jurassic Of Northeastern China. Bulletin of the American Museum of Natural History 2017 (411): 1-66, DOI: 10.1206/0003-0090-411.1.1, URL: http://www.bioone.org/doi/10.1206/0003-0090-411.1.1
ABSTRACT
Four new specimens of Anchiornis huxleyi (PKUP V 1068, BMNHC PH 804, BMNHC PH 822, and BMNHC PH 823) were recently recovered from the Late Jurassic fossil beds of the Tiaojishan Formation in northeastern China. These new specimens are almost completely preserved with cranial and postcranial skeletons. Morphological features of Anchiornis huxleyi have implications for paravian character evolution and provide insights into the relationships of major paravian lineages. Anchiornis huxleyi shares derived features with avialans, such as a straight nasal process of the premaxilla and the absence of an external mandibular fenestra in lateral view. However, Anchiornis huxleyi lacks several derived deinonychosaurian features, including a laterally exposed splenial and a specialized raptorial pedal digit II. Morphological comparisons strongly suggest Anchiornis is more closely related to avialans than to deinonychosaurians or troodontids. Anchiornis huxleyi exhibits many conservative paravian features, and closely resembles Archaeopteryx and other Jurassic paravians from Jianchang County, such as Xiaotingia and Eosinopteryx. The other Jianchang paravian, Aurornis xui, is likely a junior synonym of Anchiornis huxleyi.
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
This dataset contains the digitized treatments in Plazi based on the original journal article Sinev, Sergey Yu., Korb, Stanislav K. (2022): A preliminary list of the Pyraloid moths (Lepidoptera: Pyraloidea) of Kyrgyzstan. Zootaxa 5138 (2): 101-136, DOI: 10.11646/zootaxa.5138.2.1
Abstract
A preliminary list of the pyraloid moths of Kyrgyzstan is provided. In total, 264 species are mentioned for this country, among them 117 listed by literature data and 146 species recorded in current paper for the first time. Another 97 species are known from the neighboring territories and most probably are also present in Kyrgyzstan.
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
## Overview
Only Basketball is a dataset for object detection tasks - it contains Basketball Q9SX annotations for 1,334 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [Public Domain license](https://creativecommons.org/licenses/Public Domain).
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
Fruit And Vegetable Finder is a dataset for object detection tasks - it contains Fruits Vegetables annotations for 230 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
TwitterDC3_AircraftRemoteSensing_DIAL_DC8_Data are remotely sensed data collected by the Differential Absorption Lidar (DIAL) onboard the DC-8 aircraft during the Deep Convective Clouds and Chemistry (DC3) field campaign. Data collection for this product is complete. The Deep Convective Clouds and Chemistry (DC3) field campaign sought to understand the dynamical, physical, and lightning processes of deep, mid-latitude continental convective clouds and to define the impact of these clouds on upper tropospheric composition and chemistry. DC3 was conducted from May to June 2012 with a base location of Salina, Kansas. Observations were conducted in northeastern Colorado, west Texas to central Oklahoma, and northern Alabama in order to provide a wide geographic sample of storm types and boundary layer compositions, as well as to sample convection. DC3 had two primary science objectives. The first was to investigate storm dynamics and physics, lightning and its production of nitrogen oxides, cloud hydrometeor effects on wet deposition of species, surface emission variability, and chemistry in anvil clouds. Observations related to this objective focused on the early stages of active convection. The second objective was to investigate changes in upper tropospheric chemistry and composition after active convection. Observations related to this objective focused on the 12-48 hours following convection. This objective also served to explore seasonal change of upper tropospheric chemistry. In addition to using the NSF/NCAR Gulfstream-V (GV) aircraft, the NASA DC-8 was used during DC3 to provide in-situ measurements of the convective storm inflow and remotely-sensed measurements used for flight planning and column characterization. DC3 utilized ground-based radar networks spread across its observation area to measure the physical and kinematic characteristics of storms. Additional sampling strategies relied on lightning mapping arrays, radiosondes, and precipitation collection. Lastly, DC3 used data collected from various satellite instruments to achieve its goals, focusing on measurements from CALIOP onboard CALIPSO and CPL onboard CloudSat. In addition to providing an extensive set of data related to deep, mid-latitude continental convective clouds and analyzing their impacts on upper tropospheric composition and chemistry, DC3 improved models used to predict convective transport. DC3 improved knowledge of convection and chemistry, and provided information necessary to understanding the processes relating to ozone in the upper troposphere.
Facebook
Twitterhttp://opendatacommons.org/licenses/dbcl/1.0/http://opendatacommons.org/licenses/dbcl/1.0/
The Hops Health Image Dataset is a meticulously curated collection, originally assembled by Kaggle user scruggzilla, featuring high-definition images that depict various health conditions of hop plants. These conditions range from diseases like Downy and Powdery mildew to nutrient deficiencies and pest attacks, contrasted with images of healthy hops.
My transition from a brewmaster to a data scientist has inspired the repurposing of this dataset, aiming to harness advanced data techniques to ensure the vitality of ingredients crucial to brewing. This dataset serves as a confluence of agricultural insight and data science, intended to push forward the frontiers of research in plant pathology, brewing science, and agricultural technology.
File Structure and Contents: The dataset, structured in intuitive subdirectories, simplifies access to specific categories of hop plant health. Each category is represented in its dedicated folder, housing numerous JPEG images that capture the nuances of each health condition. The categories include:
Disease-Downy: This folder features images of hop plants grappling with Downy mildew, a fungal disease that poses a significant threat to yield and quality. Disease-Powdery: Here, you'll find images detailing the impact of Powdery mildew, another fungal adversary known to the hop community. Healthy: A collection of hop plants in their robust form, these images represent the desired standard for cultivation and further brewing processes. Nutrient: This section highlights the subtle yet impactful signs of nutrient deficiencies in hop plants, conditions that can alter plant health and compromise produce quality. Pests: Unveiling the effects of pest infestations, the images in this folder show hop plants in distress, emphasizing the need for effective pest control measures.
Usage: This dataset is a valuable asset for stakeholders across various fields. Agricultural researchers can delve into disease patterns, educators can illustrate plant health concepts, and machine learning enthusiasts can develop models for automated disease detection. Most personally, it allows brewing professionals to leverage data-driven insights for quality control and innovation in brewing practices.
Acknowledgments: I extend my deepest gratitude to scruggzilla for their foundational work in creating this dataset. Their contribution has set the stage for further exploration and innovation in combining the realms of brewing and data science.
Personal Note: Transitioning from brewmaster to data scientist, I chose this dataset to contribute to an industry I hold dear. I aspire to integrate modern data science solutions with traditional brewing wisdom, ensuring the health of hop plants, a critical ingredient in brewing. This endeavor aims to enhance the consistency and caliber of produce used, potentially transforming brewing methodologies.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
VLM-3R Training Data
Training QA data for VLM-3R: vsibench_train/ (VSI-Bench-style tasks) and vstibench_train/ (VSTI-Bench tasks over ScanNet train split).
Erratum (2026-07-13): corrected camera-position ground truth
A bug in the QA generation pipeline (reported by Jacob Yeung, CMU) extracted the camera center from camera-to-world poses using -R.T @ t instead of pose[:3, 3]. Answers in five vstibench_train files depended on the camera's world position and have… See the full description on the dataset page: https://huggingface.co/datasets/Journey9ni/VLM-3R-DATA.