Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
YOLO Version Test Dataset is a dataset for object detection tasks - it contains Objects annotations for 1,992 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
midooy/RDRF-dataset dataset hosted on Hugging Face and contributed by the HF Datasets community
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset was created by Farshad Tofighi
Released under CC0: Public Domain
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
Sweetness Watermelon is a dataset for classification tasks - it contains Watermelon annotations for 700 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
TwitterAttribution-NonCommercial-ShareAlike 4.0 (CC BY-NC-SA 4.0)https://creativecommons.org/licenses/by-nc-sa/4.0/
License information was derived automatically
LlamaLens: Specialized Multilingual LLM Dataset
This dataset supports the research presented in the paper LlamaLens: Specialized Multilingual LLM for Analyzing News and Social Media Content.
Overview
LlamaLens is a specialized multilingual LLM designed for analyzing news and social media content. It focuses on 18 NLP tasks, leveraging 52 datasets across Arabic, English, and Hindi. This repository contains the English-language portion of the data.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/LlamaLens-English.
Facebook
Twitterfarsi-asr/ganjoor-dataset dataset hosted on Hugging Face and contributed by the HF Datasets community
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This dataset provides real skeletal radiographs sampled from the FracAtlas dataset (Abedeen et al., Scientific Data 2023, doi: 10.1038/s41597-023-02932-1, CC BY 4.0), one of the largest publicly available collections of annotated bone X-rays.
The dataset includes images across multiple skeletal regions. It is designed to challenge models in identifying whether a fracture is present from medical imaging.
| Field | Value |
|---|---|
| Original Dataset | FracAtlas |
| Authors | Abedeen et al. |
| Publication | Scientific Data 10, 521 (2023) |
| License | CC BY 4.0 — commercial and academic use permitted |
| Task Type | Binary / Multi-class Bone Fracture classification |
| Modality | Conventional radiograph (X-ray) |
The raw downloaded artifact contains the following structure:
images/ folder:-
4,083 images in two subdirectories: Fractured/ (717 images) and Non_fractured/ (3,366 images)
Original resolution varies; resized to ≤384px PNG during processing
Annotations/ folder:- Contains detailed fracture bounding box and polygon annotations in multiple formats: COCO JSON/, PASCAL VOC/, VGG JSON/, and YOLO/. (Note: This challenge focuses on image-level multi-task classification and does not evaluate bounding box metrics).
Utilities/ folder:-
Contains helper notebooks (coco2yolo.ipynb, yolo2voc.ipynb) and metadata for standard dataset splits (Fracture Split/)
dataset.csv
| Column | Data Type | Description |
|:---|:---|:---|
| image_id | String | Original filename (e.g., IMG0000019.jpg) |
| hand, leg, hip, shoulder | Integer (0/1) | One-hot anatomical region encoding |
| mixed | Integer (0/1) | Mixed anatomical regions |
| hardware | Integer (0/1) | Orthopedic hardware visible |
| multiscan | Integer (0/1) | Multiple views in single image |
| fractured | Integer (0/1) | Binary fracture label |
| fracture_count | Integer | Number of visible fractures |
| frontal, lateral, oblique | Integer (0/1) | X-ray projection type |
The images are real, down-sampled radiographs.
The splits generated in the public dataset correspond strictly to standard diagnosis workflows, mapping visual radiograph properties to human-readable radiology impressions.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
VLM-3R Training Data
Training QA data for VLM-3R: vsibench_train/ (VSI-Bench-style tasks) and vstibench_train/ (VSTI-Bench tasks over ScanNet train split).
Erratum (2026-07-13): corrected camera-position ground truth
A bug in the QA generation pipeline (reported by Jacob Yeung, CMU) extracted the camera center from camera-to-world poses using -R.T @ t instead of pose[:3, 3]. Answers in five vstibench_train files depended on the camera's world position and have… See the full description on the dataset page: https://huggingface.co/datasets/Journey9ni/VLM-3R-DATA.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
Digital Caps'ule Memories is a dataset for object detection tasks - it contains Beer Bottle Caps annotations for 248 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
The VCF files contained all of cleaned SNPs and small InDels across the whole genome in lychee population.
Facebook
TwitterFood safety practices play a crucial role in the prevention of foodborne diseases, particularly in low- and middle-income countries like Bangladesh. This study assessed food safety practices among household food handlers in Patuakhali, Bangladesh, and identified associated factors influencing these practices. A cross-sectional study was conducted among 300 randomly selected households, using structured interviews and direct observations. The findings revealed that only 46% of participants demonstrated good food safety practices, with notable deficiencies in proper handwashing techniques (36.7%). Multiple logistic regression analysis identified that secondary education (AOR = 2.84; 95% CI: 1.44, 5.59), government employment (AOR = 5.74; 95% CI: 1.24, 26.53), monthly income between 15,000 and 30,000 BDT (AOR = 4.50; 95% CI: 2.17, 9.31), and participation in food safety training (AOR = 5.01; 95% CI: 1.95, 12.90) were significantly associated with good food safety practices. Conversely, living in rural areas (AOR = 0.30; 95% CI: 0.13–0.67) and, being aged 39–58 years (AOR = 0.36; 95% CI: 0.15–0.84) were associated with poor food safety practices. Addressing these factors, particularly socioeconomic disparities and offering targeted food safety education, could significantly improve public health outcomes and overall food safety practices.
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
This dataset contains the digitized treatments in Plazi based on the original journal article Skoracka, Anna, Shi, Aoxiang, Pacyna, Anna (2001): New eriophyoid mites (Acari: Eriophyoidea) associated with grasses from Mongolia. Zootaxa 9: 1-18, DOI: 10.5281/zenodo.4620032
Abstract
Two new species, Aculodes mongolicus sp. n. from Hordeum brevisubulatum (Trin.) Link, Eriophyes bromusi sp. n. from Bromus inermis Leyss, are described from Mongolia. Four new records of eriophyoid mites collected from grasses in Mongolia are also presented.
Key words: Aculodes mongolicus, Eriophyes bromusi, Eriophyoidea, grasses, Mongolia, taxonomy
Facebook
TwitterSundance State Bank location dataset — United States subset. Verified addresses and coordinates. Licensed via CREHQ Data Store.
Facebook
TwitterThis is the edited sequence set for the data in the cowbird dataset. Due to file size limitations, these are only the unique sequences. For the unedited dataset, please contact the authors of the manuscript: Hird et al. 2014. Sampling locality is more detectable than taxonomy or ecology in the brood-parasitic Brown-Headed Cowbird. PeerJ. DOI:10.7717/peerj.321. Available: https://peerj.com/articles/321/
Facebook
TwitterCAL_LID_L2_05kmALay-Standard-V5-00 is the Cloud-Aerosol Lidar and Infrared Pathfinder Satellite Observation (CALIPSO) Lidar Level 2 5 km Aerosol Layer data product. This data product was collected using the Cloud-Aerosol Lidar with Orthogonal Polarization (CALIOP) instrument. Within this aerosol layer product, generated at a horizontal resolution of 5 km, are two general classes of data: Column Properties (including position data and viewing geometry) and Layer Properties. The aerosol layer products consist of a sequence of column descriptors, each associated with a variable number of cloud layer descriptors. The column descriptors specify the temporal and geophysical location of the column of the atmosphere through which a given lidar pulse travels. Also included in the column descriptors are indicators of surface lighting conditions, information about the surface type, and the number of features (e.g., aerosol layers) identified within the column. For each feature within a column, a set of layer descriptors is reported. The layer descriptors provide information about the spatial and optical characteristics of a feature, such as base and top altitudes, integrated attenuated backscatter, and optical depth. CALIPSO was a partnership between NASA and the French Space Agency, CNES. CALIPSO was launched on April 28, 2006 to study the many roles played by clouds and aerosols in Earth’s climate and weather. It flew in the international A-Train constellation for coincident Earth observations from launch until September 13, 2018,when CALIPSO began lowering its orbit from 705 km to 688 km (428 miles) above the Earth to resume formation flying with CloudSat as part of the “C-Train”. The CALIPSO satellite carried three remote sensing instruments: the Cloud-Aerosol Lidar with Orthogonal Polarization(CALIOP), the Imaging Infrared Radiometer (IIR), and the Wide Field-of-View Camera (WFC). By mutual agreement between NASA and CNES, the CALIPSO science mission concluded on August1, 2023.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
http://data.europa.eu/eli/dec/2011/833/ojhttp://data.europa.eu/eli/dec/2011/833/oj
Eine Teilmenge der Tenders Electronic Daily (TED)-Daten für die Vergabe öffentlicher Aufträge in der Europäischen Union und darüber hinaus vom 1.1.2006 bis zum 31.12.2023 im Format „Comma Separated Value“ (CSV). Diese Daten umfassen die wichtigsten Felder aus den Standardformularen für die Auftragsbekanntmachung und die Vergabebekanntmachung, z. B. wer was von wem gekauft hat, für wie viel und welches Verfahren und welche Zuschlagskriterien verwendet wurden. Im Allgemeinen bestehen die Daten aus Angeboten oberhalb der Beschaffungsschwellen. Die Veröffentlichung von Angeboten unterhalb des Schwellenwerts in TED gilt jedoch als bewährtes Verfahren, sodass auch eine nicht zu vernachlässigende Anzahl von Angeboten unterhalb des Schwellenwerts vorliegt.Bitte beachten Sie die nachstehenden Unterlagen für wichtige Informationen zu den Daten und ihrer Verwendung, einschließlich einer Versionshistorie des Exports.Die Europäische Kommission ist an den Ergebnissen der Forschung zum öffentlichen Beschaffungswesen interessiert, die aus der Weiterverwendung dieser Daten resultieren. Wir freuen uns daher über Links zu Papieren, Berichten oder Bewerbungen unter GROW-G4@ec.europa.eu. TED mit breiterer Abdeckung ist auch im XML-Format unter https://data.europa.eu/euodp/en/data/dataset/ted-1.eForms verfügbarAm 14. November 2022 änderte sich das Format der in TED veröffentlichten Bekanntmachungen: Das Amt für Veröffentlichungen zeigt sowohl die aktuellen Standardformulare als auch die eForms an und stellt sie zur Weiterverwendung zur Verfügung. Wenn Sie TED-Daten wiederverwenden, müssen Ihre Systeme bereit sein, beide Arten von Mitteilungen zu verarbeiten. Zur Anpassung Ihrer Systeme finden Sie Ressourcen, Modelle und Schemata im eForms Software Development Kit auf GitHub (https://github.com/OP-TED/eForms-SDK/https://github.com/OP-TED/eForms-SDK/). Die Dokumentation ist auf der Website „Ted Developers Documentation“ (https://docs.ted.europa.eu/) verfügbar, einschließlich häufig gestellter Fragen zu eForms (https://docs.ted.europa.eu/home/FAQ/eforms.html).
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
Futbol Amateur is a dataset for object detection tasks - it contains Futbol Amateur annotations for 511 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
TwitterThis dataset was created by Kuleen
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
COCO-Counterfactuals is a high quality synthetic dataset for multimodal vision-language model evaluation and for training data augmentation. Each COCO-Counterfactuals example includes a pair of image-text pairs; one is a counterfactual variation of the other. The two captions are identical to each other except a noun subject. The two corresponding synthetic images differ only in terms of the altered subject in the two captions. In our accompanying paper, we showed that the COCO-Counterfactuals dataset is challenging for existing pre-trained multimodal models and significantly increase the difficulty of the zero-shot image-text retrieval and image-text matching tasks. Our experiments also demonstrate that augmenting training data with COCO-Counterfactuals improves OOD generalization on multiple downstream tasks.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Context
The dataset tabulates the Stanford population over the last 20 plus years. It lists the population for each year, along with the year on year change in population, as well as the change in percentage terms for each year. The dataset can be utilized to understand the population change of Stanford across the last two decades. For example, using this dataset, we can identify if the population is declining or increasing. If there is a change, when the population peaked, or if it is still growing and has not reached its peak. We can also compare the trend with the overall trend of United States population over the same period of time.
Key observations
In 2022, the population of Stanford was 592, a 0.67% decrease year-by-year from 2021. Previously, in 2021, Stanford population was 596, a decline of 0.17% compared to a population of 597 in 2020. Over the last 20 plus years, between 2000 and 2022, population of Stanford decreased by 108. In this period, the peak population was 709 in the year 2001. The numbers suggest that the population has already reached its peak and is showing a trend of decline. Source: U.S. Census Bureau Population Estimates Program (PEP).
When available, the data consists of estimates from the U.S. Census Bureau Population Estimates Program (PEP).
Data Coverage:
Variables / Data Columns
Good to know
Margin of Error
Data in the dataset are based on the estimates and are subject to sampling variability and thus a margin of error. Neilsberg Research recommends using caution when presening these estimates in your research.
Custom data
If you do need custom data for any of your research project, report or presentation, you can contact our research staff at research@neilsberg.com for a feasibility of a custom tabulation on a fee-for-service basis.
Neilsberg Research Team curates, analyze and publishes demographics and economic data from a variety of public and proprietary sources, each of which often includes multiple surveys and programs. The large majority of Neilsberg Research aggregated datasets and insights is made available for free download at https://www.neilsberg.com/research/.
This dataset is a part of the main dataset for Stanford Population by Year. You can refer the same here
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
YOLO Version Test Dataset is a dataset for object detection tasks - it contains Objects annotations for 1,992 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).