Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
Video-As-Prompt: Unified Semantic Control for Video Generation
🔥 News
Oct 24, 2025: 📖 We release the first unified semantic video generation model, Video-As-Prompt (VAP)! Oct 24, 2025: 🤗 We release the VAP-Data, the largest semantic-controlled video generation datasets with more than $100K$ samples! Oct 24, 2025: 👋 We present the technical report of Video-As-Prompt, please check out the details and spark some discussion!… See the full description on the dataset page: https://huggingface.co/datasets/BianYx/VAP-Data.
Facebook
TwitterGDPa1: Antibody developability dataset
Contains the assay data for 242 antibodies across 10 assays as described in our latest preprint, PROPHET-Ab: A high-throughput platform for biophysical antibody developability assessment to enable AI/ML model training.
Example usage
Using pandas: import pandas as pd
huggingface-cli login to access this datasetdf = pd.read_csv("hf://datasets/ginkgo-datapoints/GDPa1/GDPa1_v1.2_20250814.csv")
Using Hugging… See the full description on the dataset page: https://huggingface.co/datasets/ginkgo-datapoints/GDPa1.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
VQA is a multimodal task wherein, given an image and a natural language question related to the image, the objective is to produce a natural language answer correctly as output.
It involves understanding the content of the image and correlating it with the context of the question asked. Because we need to compare the semantics of information present in both of the modalities — the image and natural language question related to it — VQA entails a wide range of sub-problems in both CV and NLP (such as object detection and recognition, scene classification, counting, and so on). Thus, it is considered an AI-complete task.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
YOLO Version Test Dataset is a dataset for object detection tasks - it contains Objects annotations for 1,992 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
Dataset Ow is a dataset for object detection tasks - it contains Player annotations for 10,000 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Context
The dataset tabulates the population of Onawa by gender, including both male and female populations. This dataset can be utilized to understand the population distribution of Onawa across both sexes and to determine which sex constitutes the majority.
Key observations
There is a majority of female population, with 53.95% of total population being female. Source: U.S. Census Bureau American Community Survey (ACS) 2019-2023 5-Year Estimates.
When available, the data consists of estimates from the U.S. Census Bureau American Community Survey (ACS) 2019-2023 5-Year Estimates.
Scope of gender :
Please note that American Community Survey asks a question about the respondents current sex, but not about gender, sexual orientation, or sex at birth. The question is intended to capture data for biological sex, not gender. Respondents are supposed to respond with the answer as either of Male or Female. Our research and this dataset mirrors the data reported as Male and Female for gender distribution analysis. No further analysis is done on the data reported from the Census Bureau.
Variables / Data Columns
Good to know
Margin of Error
Data in the dataset are based on the estimates and are subject to sampling variability and thus a margin of error. Neilsberg Research recommends using caution when presening these estimates in your research.
Custom data
If you do need custom data for any of your research project, report or presentation, you can contact our research staff at research@neilsberg.com for a feasibility of a custom tabulation on a fee-for-service basis.
Neilsberg Research Team curates, analyze and publishes demographics and economic data from a variety of public and proprietary sources, each of which often includes multiple surveys and programs. The large majority of Neilsberg Research aggregated datasets and insights is made available for free download at https://www.neilsberg.com/research/.
This dataset is a part of the main dataset for Onawa Population by Race & Ethnicity. You can refer the same here
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Context
The dataset presents the distribution of median household income among distinct age brackets of householders in Shepherd. Based on the latest 2019-2023 5-Year Estimates from the American Community Survey, it displays how income varies among householders of different ages in Shepherd. It showcases how household incomes typically rise as the head of the household gets older. The dataset can be utilized to gain insights into age-based household income trends and explore the variations in incomes across households.
Key observations: Insights from 2023
In terms of income distribution across age cohorts, in Shepherd, the median household income stands at $71,563 for householders within the 25 to 44 years age group, followed by $57,672 for the 65 years and over age group. Notably, householders within the 45 to 64 years age group, had the lowest median household income at $42,596.
When available, the data consists of estimates from the U.S. Census Bureau American Community Survey (ACS) 2019-2023 5-Year Estimates. All incomes have been adjusting for inflation and are presented in 2023-inflation-adjusted dollars.
Age groups classifications include:
Variables / Data Columns
Good to know
Margin of Error
Data in the dataset are based on the estimates and are subject to sampling variability and thus a margin of error. Neilsberg Research recommends using caution when presening these estimates in your research.
Custom data
If you do need custom data for any of your research project, report or presentation, you can contact our research staff at research@neilsberg.com for a feasibility of a custom tabulation on a fee-for-service basis.
Neilsberg Research Team curates, analyze and publishes demographics and economic data from a variety of public and proprietary sources, each of which often includes multiple surveys and programs. The large majority of Neilsberg Research aggregated datasets and insights is made available for free download at https://www.neilsberg.com/research/.
This dataset is a part of the main dataset for Shepherd median household income by age. You can refer the same here
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
Kahve is a dataset for object detection tasks - it contains Kahve annotations for 865 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
The application of Artificial Intelligence (AI) has been evident in the agricultural sector recently. The main goal of AI in agriculture is to improve crop yield, control crop pests/diseases, and reduce cost. The agricultural sector in developing countries faces severe in the form of disease and pest infestation, the knowledge gap between farmers and technology, and a lack of storage facilities, among others. To help address some of these challenges, this work presents crop pests/disease datasets sourced from local farms in Ghana. The dataset is presented in two folds; the raw images which consists of 24,881 images ( 6,549-Cashew, 7,508-Cassava, 5,389-Maize, and 5,435-Tomato) and augmented images which is further split into train and test set consists of 102,976 images (25,811-Cashew, 26,330-Cassava, 23,657-Maize, and 27,178-Tomato), categorized into 22 classes. All images are de-identified, validated by expert plant virologists, and freely available for use by the research community.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This dataset contains information on patent applications and grants at the 2-digit ISIC Rev.4/NACE Rev.2 sector level for a set of 24 European countries from 1985 to 2020. The data are assigned to sectors by applying the Lybbert & Zolas (2014) weights of sector-technology field correspondence to European Patent Office (EPO) patent data retrieved from OECD. The dataset provides patent data based on the applicant’s and inventor’s country of residence, on an annual basis, in two formats: i) patent applications and grants (flows), and ii) patent applications’ and grants’ stocks, estimated through the perpetual inventory method (PIM) using a 15% deprecation rate.
It was developed in the Laboratory of Industrial and Energy Economics of the National Technical University of Athens by Dr. D. Stamopoulos and Dr. P. Dimas under the supervision of Assistant Professor Aimilia Protogerou (principal investigator of GRinGVCs). The "Leveraging Global Value Chains for Innovation and Competitiveness: The Case of Greece" (GRinGVCs) project is carried out within the framework of the National Recovery and Resilience Plan Greece 2.0, funded by the European Union – NextGenerationEU (Implementation body: HFRI - Project Number: HFRI-016667).
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
Conti 2 is a dataset for object detection tasks - it contains Numbers annotations for 293 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
TwitterAction Recognition in video is known to be more challenging than image recognition problems. Unlike image recognition models which use 2D convolutional neural blocks, action classification models require additional dimensionality to capture the spatio-temporal information in video sequences. This intrinsically makes video action recognition models computationally intensive and significantly more data-hungry than image recognition counterparts. Unequivocally, existing video datasets such as Kinetics, AVA, Charades, Something-Something, HMDB51, and UFC101 have had tremendous impact on the recently evolving video recognition technologies. Artificial Intelligence models trained on these datasets have largely benefited applications such as behavior monitoring in elderly people, video summarization, and content-based retrieval. However, this growing concept of action recognition has yet to be explored in Intelligent Transportation System (ITS), particularly in vital applications such as incidents detection. This is partly due to the lack of availability of annotated dataset adequate for training models suitable for such direct ITS use cases. In this paper, the concept of video action recognition is explored to tackle the problem of highway incident detection and classification from live surveillance footage. First, a novel dataset - HWID12 (Highway Incidents Detection) dataset is introduced. The HWAD12 consists of 11 distinct highway incidents categories, and one additional category for negative samples representing normal traffic. The proposed dataset also includes 2780+ video segments of 3 to 8 seconds on average each, and 500k+ temporal frames. Next, the baseline for highway accident detection and classification is established with a state-of-the-art action recognition model trained on the proposed HWID12 dataset. Performance benchmarking for 12-class (normal traffic vs 11 accident categories), and 2-class (incident vs normal traffic) settings is performed. This benchmarking reveals a recognition accuracy of up to 88% and 98% for 12-class and 2-class recognition setting, respectively.
The Proposed Highway Incidents Detection Dataset (HWID12) is the first of its kind dataset aimed at fostering experimentation of video action recognition technologies to solve the practical problem of real-time highway incident detections which currently challenges intelligent transportation systems. The lack of such dataset has limited the expansion of the recent breakthroughs in video action classification for practical uses cases in intelligent transportation systems.. The proposed dataset contains more than 2780 video clips of length varying between 3 to 8 seconds. These video clips capture moments leading to, up until right after an incident occurred. The clips were manually segmented from accident compilations videos sourced from YouTube and other videos data platforms.
There is one main zip file available for download. The zip file contains 2780+ video clips.
1) 12 folders
2) each folder represents an incident category. One of the classes represent the negative sample class which simulates normal traffic.
Any publication using this database must reference to the following journal manuscript:
Note: if the link is broken, please use http instead of https.
In Chrome, use the steps recommended in the following website to view the webpage if it appears to be broken https://www.technipages.com/chrome-enabledisable-not-secure-warning
Other relevant datasets VCoR dataset: https://www.kaggle.com/landrykezebou/vcor-vehicle-color-recognition-dataset VRiV dataset: https://www.kaggle.com/landrykezebou/vriv-vehicle-recognition-in-videos-dataset
For any enquires regarding the HWID12 dataset, contact: landrykezebou@gmail.com
Facebook
Twitterhttps://baristalife.co/pages/caffeine-datahttps://baristalife.co/pages/caffeine-data
Verified caffeine content for 154 drinks across coffee, tea, energy drinks, soda, ready to drink, chocolate, dessert, and decaf, with serving size, caffeine per ounce, and a cited source on every row. Updated quarterly.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
So, a few days ago I bought this game by Reiner Knizia, called Lost Cities. Got totally hooked on the game. But the scoring of the game is a bit hard, so let's try to come up with some kind of a model that can identify the set and a small program that can calculate the score.
The following 11 classes are used: * 2-10 - numbers * w - a bet * set - a set of cards that together make up the score
Thanks to Erik Dekker for sending some images my way. If you have more images for me; please let me know: https://twitter.com/keestalkstech
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
REU Original Metadataset is a dataset for object detection tasks - it contains Transmission Lines UJFL annotations for 2,485 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset was created by SeshuRaju 🧘♂️
Released under CC0: Public Domain
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
Ayakkabı New is a dataset for object detection tasks - it contains Ayakkabi annotations for 3,594 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
TwitterAttribution 3.0 (CC BY 3.0)https://creativecommons.org/licenses/by/3.0/
License information was derived automatically
This is a working unpublished document based on the NZMS260 Map Series, and is a precursor to the publication of QMAP geological map 18 Wakatipu. Map, ink and pencil on paper, medium detail, poor condition. - Observation measure: Mainly interpretation. - Map size: 900 x 700 mm. Notes: This is a copy of the original and is cellotaped together. Keywords: LAKE WAKATIPU; GEOLOGIC MAPS; QMAP; ARROWTOWN; AERIAL PHOTOGRAPHY; PHOTOINTERPRETATION; LANDSLIDES; QUATERNARY
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
VLM-3R Training Data
Training QA data for VLM-3R: vsibench_train/ (VSI-Bench-style tasks) and vstibench_train/ (VSTI-Bench tasks over ScanNet train split).
Erratum (2026-07-13): corrected camera-position ground truth
A bug in the QA generation pipeline (reported by Jacob Yeung, CMU) extracted the camera center from camera-to-world poses using -R.T @ t instead of pose[:3, 3]. Answers in five vstibench_train files depended on the camera's world position and have… See the full description on the dataset page: https://huggingface.co/datasets/Journey9ni/VLM-3R-DATA.
Facebook
TwitterU.S. Government Workshttps://www.usa.gov/government-works
License information was derived automatically
This archive contains raw observations of the 2009-10-09 impact of the LCROSS spacecraft on the moon by the CLIO instrument on the MMT Observatory 6.5m telescope. The archive consists of uncalibrated FITS images of the event. This is one of several data sets of Earth-based observations of the impact.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
Video-As-Prompt: Unified Semantic Control for Video Generation
🔥 News
Oct 24, 2025: 📖 We release the first unified semantic video generation model, Video-As-Prompt (VAP)! Oct 24, 2025: 🤗 We release the VAP-Data, the largest semantic-controlled video generation datasets with more than $100K$ samples! Oct 24, 2025: 👋 We present the technical report of Video-As-Prompt, please check out the details and spark some discussion!… See the full description on the dataset page: https://huggingface.co/datasets/BianYx/VAP-Data.