Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
VLM-3R Training Data
Training QA data for VLM-3R: vsibench_train/ (VSI-Bench-style tasks) and vstibench_train/ (VSTI-Bench tasks over ScanNet train split).
Erratum (2026-07-13): corrected camera-position ground truth
A bug in the QA generation pipeline (reported by Jacob Yeung, CMU) extracted the camera center from camera-to-world poses using -R.T @ t instead of pose[:3, 3]. Answers in five vstibench_train files depended on the camera's world position and have… See the full description on the dataset page: https://huggingface.co/datasets/Journey9ni/VLM-3R-DATA.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
VQA is a multimodal task wherein, given an image and a natural language question related to the image, the objective is to produce a natural language answer correctly as output.
It involves understanding the content of the image and correlating it with the context of the question asked. Because we need to compare the semantics of information present in both of the modalities — the image and natural language question related to it — VQA entails a wide range of sub-problems in both CV and NLP (such as object detection and recognition, scene classification, counting, and so on). Thus, it is considered an AI-complete task.
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
This dataset is designed for traffic surveillance anomaly detection, originally from the WSAL (Weakly-Supervised Anomaly Localization) repository. It consists of 500 short video clips totaling approximately 25 hours of footage. Each clip averages around 1,075 frames, and anomalies, when present, typically span around 80 frames.
Each video is labeled to indicate whether it contains an anomaly or not, enabling both supervised training and evaluation. You can use the labels to develop or compare different anomaly detection methods.
If you use this dataset for your research, please cite the following paper:
@article{wsal_tip21,
author = {Hui Lv and
Chuanwei Zhou and
Zhen Cui and
Chunyan Xu and
Yong Li and
Jian Yang},
title = {Localizing Anomalies from Weakly-Labeled Videos},
journal = {IEEE Transactions on Image Processing (TIP)},
year = {2021}
}
For more details about how the dataset was created and used, see the original WSAL GitHub repository.
Facebook
TwitterThis repo consists of the datasets used for the TaCo paper. There are four datasets:
Multilingual Alpaca-52K GPT-4 dataset Multilingual Dolly-15K GPT-4 dataset TaCo dataset Multilingual Vicuna Benchmark dataset
We translated the first three datasets using Google Cloud Translation. The TaCo dataset is created by using the TaCo approach as described in our paper, combining the Alpaca-52K and Dolly-15K datasets. If you would like to create the TaCo dataset for a specific language, you can… See the full description on the dataset page: https://huggingface.co/datasets/saillab/taco-datasets.
Facebook
TwitterThis dataset was created by Rithik Kotha
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
Dataset Ow is a dataset for object detection tasks - it contains Player annotations for 10,000 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Context
The dataset tabulates the population of Ronceverte by gender, including both male and female populations. This dataset can be utilized to understand the population distribution of Ronceverte across both sexes and to determine which sex constitutes the majority.
Key observations
There is a majority of female population, with 55.02% of total population being female. Source: U.S. Census Bureau American Community Survey (ACS) 2019-2023 5-Year Estimates.
When available, the data consists of estimates from the U.S. Census Bureau American Community Survey (ACS) 2019-2023 5-Year Estimates.
Scope of gender :
Please note that American Community Survey asks a question about the respondents current sex, but not about gender, sexual orientation, or sex at birth. The question is intended to capture data for biological sex, not gender. Respondents are supposed to respond with the answer as either of Male or Female. Our research and this dataset mirrors the data reported as Male and Female for gender distribution analysis. No further analysis is done on the data reported from the Census Bureau.
Variables / Data Columns
Good to know
Margin of Error
Data in the dataset are based on the estimates and are subject to sampling variability and thus a margin of error. Neilsberg Research recommends using caution when presening these estimates in your research.
Custom data
If you do need custom data for any of your research project, report or presentation, you can contact our research staff at research@neilsberg.com for a feasibility of a custom tabulation on a fee-for-service basis.
Neilsberg Research Team curates, analyze and publishes demographics and economic data from a variety of public and proprietary sources, each of which often includes multiple surveys and programs. The large majority of Neilsberg Research aggregated datasets and insights is made available for free download at https://www.neilsberg.com/research/.
This dataset is a part of the main dataset for Ronceverte Population by Race & Ethnicity. You can refer the same here
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
Humanoid_grip_4class is a dataset for object detection tasks - it contains Humanoid_grip_4class annotations for 15,639 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
Wet Floor is a dataset for object detection tasks - it contains Sign annotations for 218 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
G1_football is a dataset for object detection tasks - it contains G1 Football annotations for 303 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
Bitcoin, the pioneering cryptocurrency, has become a significant asset class in the global financial landscape. This dataset provides a comprehensive historical record of Bitcoin's price movements and trading volume.
Data Source: The dataset is sourced from reputable financial data providers and aggregators. It encompasses a wide range of historical Bitcoin price and volume data.
Content:
Date: The date of the recorded data point. Open: The opening price of Bitcoin on the given date. High: The highest price of Bitcoin reached during the day. Low: The lowest price of Bitcoin reached during the day. Close: The closing price of Bitcoin on the given date. Adj Close: The adjusted closing price, considering factors such as dividends and stock splits. Volume: The trading volume of Bitcoin on the given date. Time Period: The dataset spans from [start date] to [end date], providing a comprehensive historical perspective on Bitcoin's price and trading activity.
Frequency: The data is recorded at [daily/hourly/etc.] intervals.
Missing Values: Any missing values in the dataset have been appropriately handled.
Data Format: The dataset is provided in CSV format for easy integration and analysis.
Facebook
TwitterInformation on beneficiary, financial, quality, and cost‑and‑use measures for organizations participating in the Next Generation Accountable Care Organization (NGACO) Model.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
This dataset contains the digitized treatments in Plazi based on the original journal article Slager, David L., Klicka, John (2014): A new genus for the American Tree Sparrow (Aves: Passeriformes: Passerellidae). Zootaxa 3821 (3): 398-400, DOI: 10.11646/zootaxa.3821.3.9
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset was created by mdnurhossen
Released under CC0: Public Domain
Facebook
TwitterIntrusion Detection Systems (IDS) for IoT and edge environments require datasets with unambiguous labels. Existing corpora often mix benign and malicious traffic within the same capture window, producing ambiguous flow labels that distort model evaluation. This work introduces the TRUST Lab Dataset, a flow-based traffic corpus designed under a single-class session policy: each capture contains exclusively benign traffic or a single attack family, preventing temporal overlap and ensuring label integrity at the bi-flow level. The dataset was generated in an operational testbed reproducing enterprise-grade services (HTTP/S, DNS, email, SSH, SNMP, NTP, MySQL) and modern APIs (REST, GraphQL, SOAP), combined with 15 attack families spanning DDoS/DoS, port scanning, brute force, web/API exploits, DNS abuses, MITM, NIDS evasion, tunneling/exfiltration, C2/beaconing, TLS/SSL anomalies, buffer overflow, and slowloris. Traffic was processed with CICFlowMeter into 16 single-class CSV files totaling ~4.6million bi-flows with 80 features per flow. Comprehensive statistical analyses (distributions, correlation, ANOVA, PCA) confirm discriminative signal without payload inspection. A binary meta-classifier achieves ROC-AUC 0.9676 and recall 0.95, validating TRUST Lab's utility for lightweight edge-oriented IDS evaluation.
Facebook
TwitterAttribution-NonCommercial-ShareAlike 4.0 (CC BY-NC-SA 4.0)https://creativecommons.org/licenses/by-nc-sa/4.0/
License information was derived automatically
This dataset contains 195,000+ raw records of Escherichia coli clinical isolates and their antimicrobial susceptibility test results. The data was extracted from the Bacterial and Viral Bioinformatics Resource Center (BV-BRC), a public repository funded by NIAID.
Each entry captures how a specific E. coli genome responds to a given antibiotic, along with phenotypic interpretation, lab methods, measurement values (e.g., MIC), and supporting publication links.
Facebook
TwitterThe Allen Brain Observatory – Visual Coding is a large-scale, standardized survey of physiological activity across the mouse visual cortex, hippocampus, and thalamus. It includes datasets collected with both two-photon imaging and Neuropixels probes, two complementary techniques for measuring the activity of neurons in vivo. The two-photon imaging dataset features visually evoked calcium responses from GCaMP6-expressing neurons in a range of cortical layers, visual areas, and Cre lines. The Neuropixels dataset features spiking activity from distributed cortical and subcortical brain regions, collected under analogous conditions to the two-photon imaging experiments. We hope that experimentalists and modelers will use these comprehensive, open datasets as a testbed for theories of visual information processing.
Facebook
Twitter7-Eleven location dataset — Denmark subset. Verified addresses, coordinates, and phones. Licensed via CREHQ Data Store.
Facebook
TwitterThe focus of this work is the analysis of different degradation phenomena based on thermal overstress and electrical overstress accelerated aging systems and the use of accelerated aging techniques for prognostics algorithm development. Results on thermal overstress and electrical overstress experiments are presented. In addition, preliminary results toward the development of physics-based degradation models are presented focusing on the electrolyte evaporation failure mechanism. An empirical degradation model based on percentage capacitance loss under electrical overstress is presented and used in: (i) a Bayesian-based implementation of model-based prognostics using a discrete Kalman filter for health state estimation, and (ii) a dynamic system representation of the degradation model for forecasting and remaining useful life (RUL) estimation. A leave-one-out validation methodology is used to assess the validity of the methodology under the small sample size constrain. The results observed on the RUL estimation are consistent through the validation tests comparing relative accuracy and prediction error. It has been observed that the inaccuracy of the model to represent the change in degradation behavior observed at the end of the test data is consistent throughout the validation tests, indicating the need of a more detailed degradation model or the use of an algorithm that could estimate model parameters on-line. Based on the observed degradation process under different stress intensity with rest periods, the need for more sophisticated degradation models is further supported. The current degradation model does not represent the capacitance recovery over rest periods following an accelerated aging stress period.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
VLM-3R Training Data
Training QA data for VLM-3R: vsibench_train/ (VSI-Bench-style tasks) and vstibench_train/ (VSTI-Bench tasks over ScanNet train split).
Erratum (2026-07-13): corrected camera-position ground truth
A bug in the QA generation pipeline (reported by Jacob Yeung, CMU) extracted the camera center from camera-to-world poses using -R.T @ t instead of pose[:3, 3]. Answers in five vstibench_train files depended on the camera's world position and have… See the full description on the dataset page: https://huggingface.co/datasets/Journey9ni/VLM-3R-DATA.