Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
VLM-3R Training Data
Training QA data for VLM-3R: vsibench_train/ (VSI-Bench-style tasks) and vstibench_train/ (VSTI-Bench tasks over ScanNet train split).
Erratum (2026-07-13): corrected camera-position ground truth
A bug in the QA generation pipeline (reported by Jacob Yeung, CMU) extracted the camera center from camera-to-world poses using -R.T @ t instead of pose[:3, 3]. Answers in five vstibench_train files depended on the camera's world position and have… See the full description on the dataset page: https://huggingface.co/datasets/Journey9ni/VLM-3R-DATA.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
The "Plant Pathogen Dataset" is a comprehensive collection of labeled images depicting various types of pathogens affecting plant species. This dataset is curated to facilitate research and development in the field of plant pathology, enabling the development of machine learning models for automated disease diagnosis and monitoring.
Image Categories: The dataset contains images representing different types of plant diseases, including bacterial infections, fungal diseases, pest infestations, and viral infections.
The images in this dataset were sourced from various sources, including research institutions, agricultural organizations, and open-access repositories. Care was taken to ensure high-quality images with accurate disease annotations.
Disease Diagnosis: The dataset can be used to train machine learning models for automated diagnosis of plant diseases based on image analysis. Disease Monitoring: By continuously monitoring plant health using machine learning models trained on this dataset, farmers and agricultural professionals can detect diseases early and implement timely interventions.
We would like to acknowledge the contributions of the research community, agricultural experts, and dataset contributors who have made this dataset possible. Their efforts in collecting, labeling, and sharing plant disease images are invaluable to advancing research in plant pathology and agricultural technology.
Facebook
TwitterThis repo consists of the datasets used for the TaCo paper. There are four datasets:
Multilingual Alpaca-52K GPT-4 dataset Multilingual Dolly-15K GPT-4 dataset TaCo dataset Multilingual Vicuna Benchmark dataset
We translated the first three datasets using Google Cloud Translation. The TaCo dataset is created by using the TaCo approach as described in our paper, combining the Alpaca-52K and Dolly-15K datasets. If you would like to create the TaCo dataset for a specific language, you can… See the full description on the dataset page: https://huggingface.co/datasets/saillab/taco-datasets.
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
This dataset is designed for detecting and segmenting brain structures and tumor types in MRI images. It includes four classes: brain, glioma, meningioma, and pituitary. Each class represents a distinct area or tumor type within the brain, crucial for medical diagnosis and analysis.
The brain is the central organ of the human nervous system, easily identified in MRI images by its distinct outline, filling the majority of the skull area.
Gliomas are identifiable tumors within the brain, often presenting as irregularly shaped masses.
Meningiomas are typically located near the brain surface and have a somewhat rounded appearance.
The pituitary gland is a small, oval structure located at the brain's base, recognizable by its distinct placement and size.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
Dataset5 Link7 is a dataset for object detection tasks - it contains 5 Diseases annotations for 2,798 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
Conti 2 is a dataset for object detection tasks - it contains Numbers annotations for 293 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
Specific up-to-date data on each of the playable characters in Genshin Impact. Selected data include vision, rarity, release date, character stats and various ascension materials. Other fields may also be added as requested.
All characters belong to HoYoVerse. All character data was scraped from the Genshin Impact Fandom Wiki using the BeautifulSoup package in Python but may still have logical inconsistencies. Note that the .csv file is encoded in the 'UTF-8' format to ensure proper display of special characters.
Version 1 (13 Aug 2024): All playable characters available as of game version 4.8, at the end of Fontaine and right before the Natlan release.
Facebook
TwitterThis dataset was created by Singh Prince Rinku
Released under Other (specified in description)
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
The ARG Database is a huge collection of labeled and unlabeled graphs realized by the MIVIA Group. The aim of this collection is to provide the graph research community with a standard test ground for the benchmarking of graph matching algorithms.
Facebook
TwitterInformation on beneficiary, financial, quality, and cost‑and‑use measures for organizations participating in the Next Generation Accountable Care Organization (NGACO) Model.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This dataset contains NMR spectra obtained for the sample -Carapin F (OC3_DongmoAp_0418_CMS34) Nucleus: 1H NMR Solvent: CDCl3 NMR Probe: Z168773_0003 (CPP1.1 BBO 600S3 BB-H&F-D-05 Z XT) NMR Pulse Sequence: 1d Temperature: 297.9998 Observed Frequency: 600.13 Magnetic Field Strength: 14.095010308659894 Number of Scans: 16 NMR Pulse Sequence: zg30 Spectral Width: 19.8368493381866 Number of Data Points: 65536 Relaxation Delay: 1 Observed Frequency: 600.133705802 Nucleus: 13C NMR Solvent: CDCl3 NMR Probe: Z168773_0003 (CPP1.1 BBO 600S3 BB-H&F-D-05 Z XT) NMR Pulse Sequence: 1d Temperature: 298.0002 Observed Frequency: 150.902808526 Magnetic Field Strength: 14.09286355384592 Number of Scans: 1024 NMR Pulse Sequence: zgpg30 Spectral Width: 236.647117383728 Number of Data Points: 65536 Relaxation Delay: 2 Observed Frequency: 150.917898807 Nucleus: 13C NMR Solvent: CDCl3 NMR Probe: Z168773_0003 (CPP1.1 BBO 600S3 BB-H&F-D-05 Z XT) NMR Pulse Sequence: dept Temperature: 298.0009 Observed Frequency: 150.902808526 Magnetic Field Strength: 14.09286355384592 Number of Scans: 256 NMR Pulse Sequence: deptsp135 Spectral Width: 157.767899965848 Number of Data Points: 65536 Relaxation Delay: 2 Observed Frequency: 150.91488075 Nucleus: 1H,1H NMR Solvent: CDCl3 NMR Probe: Z168773_0003 (CPP1.1 BBO 600S3 BB-H&F-D-05 Z XT) NMR Pulse Sequence: cosy Temperature: 298.0007 Observed Frequency: 600.13,600.13 Magnetic Field Strength: 14.095010308659894 Number of Scans: 2 NMR Pulse Sequence: cosygpppqf Spectral Width: 11.9021176366777,11.9020999999008 Number of Data Points: 2048,128 Relaxation Delay: 2 Observed Frequency: 600.13330072,600.13330072 Nucleus: 1H,13C NMR Solvent: CDCl3 NMR Probe: Z168773_0003 (CPP1.1 BBO 600S3 BB-H&F-D-05 Z XT) NMR Pulse Sequence: hmqc Temperature: 298.0003 Observed Frequency: 600.13,150.902809 Magnetic Field Strength: 14.095010308659894 Number of Scans: 4 NMR Pulse Sequence: hmqcgpqf Spectral Width: 11.9021176366777,164.999999482852 Number of Data Points: 1024,128 Relaxation Delay: 1.5 Observed Frequency: 600.13330072,150.91412671 Nucleus: 1H,13C NMR Solvent: CDCl3 NMR Probe: Z168773_0003 (CPP1.1 BBO 600S3 BB-H&F-D-05 Z XT) NMR Pulse Sequence: hmbc Temperature: 297.9987 Observed Frequency: 600.13,150.902809 Magnetic Field Strength: 14.095010308659894 Number of Scans: 8 NMR Pulse Sequence: hmbcgpndqf Spectral Width: 11.9021176366777,239.999999250963 Number of Data Points: 4096,128 Relaxation Delay: 1.5 Observed Frequency: 600.13330072,150.92016282 Nucleus: 1H,1H NMR Solvent: CDCl3 NMR Probe: Z168773_0003 (CPP1.1 BBO 600S3 BB-H&F-D-05 Z XT) NMR Pulse Sequence: noesy Temperature: 297.9976 Observed Frequency: 600.13,600.13 Magnetic Field Strength: 14.095010308659894 Number of Scans: 4 NMR Pulse Sequence: noesygpphpp Spectral Width: 9.25721228388213,9.25721228388212 Number of Data Points: 2048,54 Relaxation Delay: 1.985664 Observed Frequency: 600.132673334975,600.132673334975 Nucleus: 1H,1H NMR Solvent: CDCl3 NMR Probe: Z168773_0003 (CPP1.1 BBO 600S3 BB-H&F-D-05 Z XT) NMR Pulse Sequence: noesy Temperature: 297.9976 Observed Frequency: 600.13,600.13 Magnetic Field Strength: 14.095010308659894 Number of Scans: 4 NMR Pulse Sequence: noesygpphpp Spectral Width: 9.25721228388213,9.25721228388212 Number of Data Points: 2048,54 Relaxation Delay: 1.985664 Observed Frequency: 600.132673334975,600.132673334975
Facebook
TwitterThe focus of this work is the analysis of different degradation phenomena based on thermal overstress and electrical overstress accelerated aging systems and the use of accelerated aging techniques for prognostics algorithm development. Results on thermal overstress and electrical overstress experiments are presented. In addition, preliminary results toward the development of physics-based degradation models are presented focusing on the electrolyte evaporation failure mechanism. An empirical degradation model based on percentage capacitance loss under electrical overstress is presented and used in: (i) a Bayesian-based implementation of model-based prognostics using a discrete Kalman filter for health state estimation, and (ii) a dynamic system representation of the degradation model for forecasting and remaining useful life (RUL) estimation. A leave-one-out validation methodology is used to assess the validity of the methodology under the small sample size constrain. The results observed on the RUL estimation are consistent through the validation tests comparing relative accuracy and prediction error. It has been observed that the inaccuracy of the model to represent the change in degradation behavior observed at the end of the test data is consistent throughout the validation tests, indicating the need of a more detailed degradation model or the use of an algorithm that could estimate model parameters on-line. Based on the observed degradation process under different stress intensity with rest periods, the need for more sophisticated degradation models is further supported. The current degradation model does not represent the capacitance recovery over rest periods following an accelerated aging stress period.
Facebook
TwitterThe Allen Brain Observatory – Visual Coding is a large-scale, standardized survey of physiological activity across the mouse visual cortex, hippocampus, and thalamus. It includes datasets collected with both two-photon imaging and Neuropixels probes, two complementary techniques for measuring the activity of neurons in vivo. The two-photon imaging dataset features visually evoked calcium responses from GCaMP6-expressing neurons in a range of cortical layers, visual areas, and Cre lines. The Neuropixels dataset features spiking activity from distributed cortical and subcortical brain regions, collected under analogous conditions to the two-photon imaging experiments. We hope that experimentalists and modelers will use these comprehensive, open datasets as a testbed for theories of visual information processing.
Facebook
TwitterHome Depot location dataset — Mexico subset. Verified addresses, coordinates, phones, and operating hours. Licensed via CREHQ Data Store.
Facebook
TwitterThe table concept_relationship is part of the dataset MedAlign, available at https://stanford.redivis.com/datasets/48nr-frxd97exb. It contains 58831134 rows across 8 variables.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
AI Safe Landing 2 is a dataset for instance segmentation tasks - it contains Landing Spot annotations for 931 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
CWHR species range datasets represent the maximum current geographic extent of each species within California. Ranges were originally delineated at a scale of 1:5,000,000 by species-level experts more than 30 years ago and have gradually been revised at a scale of 1:1,000,000. Species occurrence data are used in defining species ranges, but range polygons may extend beyond the limits of extant occurrence data for a particular species. When drawing range boundaries, CDFW seeks to err on the side of commission rather than omission. This means that CDFW may include areas within a range based on expert knowledge or other available information, despite an absence of confirmed occurrences, which may be due to a lack of survey effort. The degree to which a range polygon is extended beyond occurrence data will vary among species, depending upon each species’ vagility, dispersal patterns, and other ecological and life history factors. The boundary line of a range polygon is drawn with consideration of these factors and is aligned with standardized boundaries including watersheds (NHD), ecoregions (USDA), or other ecologically meaningful delineations such as elevation contour lines. While CWHR ranges are meant to represent the current range, once an area has been designated as part of a species’ range in CWHR, it will remain part of the range even if there have been no documented occurrences within recent decades. An area is not removed from the range polygon unless experts indicate that it has not been occupied for a number of years after repeated surveys or is deemed no longer suitable and unlikely to be recolonized. It is important to note that range polygons typically contain areas in which a species is not expected to be found due to the patchy configuration of suitable habitat within a species’ range. In this regard, range polygons are coarse generalizations of where a species may be found. This data is available for download from the CDFW website: https://www.wildlife.ca.gov/Data/CWHR. The following data sources were collated for the purposes of range mapping and species habitat modeling by RADMAP. Each focal taxon’s location data was extracted (when applicable) from the following list of sources. BIOS datasets are bracketed with their “ds” numbers and can be located on CDFW’s BIOS viewer: https://wildlife.ca.gov/Data/BIOS. California Natural Diversity Database, Terrestrial Species Monitoring [ds2826], North American Bat Monitoring Data Portal, VertNet, Breeding Bird Survey, Wildlife Insights, eBird, iNaturalist, other available CDFW or partner data.
Facebook
Twitterdatht/geneva-generated-dataset dataset hosted on Hugging Face and contributed by the HF Datasets community
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
This dataset contains the digitized treatments in Plazi based on the original journal article Nucci, Paulo Ricardo, Melo, Gustavo Augusto Schmidt De (2007): Hermit crabs from Brazil. Family Paguridae (Crustacea: Decapoda: Paguroidea): Genus Pagurus. Zootaxa 1406: 47-59, DOI: 10.5281/zenodo.175515
Abstract
In Brazil, the hermit crab family Paguridae is represented by 11 genera, of which the genus Pagurus is the most speciose, with seven species that occur from the intertidal zone to shallow waters and one from deeper regions. In this paper we present the diagnosis, distribution and some remarks on each species of the genus Pagurus known from Brazil.
Key words: Hermit crabs, Paguridae, Pagurus, Brazil
Facebook
Twitterhttps://www.usa.gov/government-workshttps://www.usa.gov/government-works
The CMS Program Statistics – Medicare Advantage, Physician, Non-Physician Practitioner and Supplier tables provide utilization data for physician, non-physician practitioners, and suppliers, by Medicare Advantage beneficiaries.
For additional information on enrollment, providers, and Medicare use and payment, visit the CMS Program Statistics page.
Below is the list of tables:
MDCR PHYSSUPP MA 1. Medicare Physicians, Non-Physician Practitioners, and Suppliers: Utilization for Medicare Advantage Beneficiaries, by Type of Entitlement, Yearly Trend
MDCR PHYSSUPP MA 2. Medicare Physicians, Non-Physician Practitioners, and Suppliers: Utilization for Medicare Advantage Beneficiaries, by Demographic Characteristics and Medicare-Medicaid Enrollment Status
MDCR PHYSSUPP MA 3. Medicare Physicians, Non-Physician Practitioners, and Suppliers: Utilization for Medicare Advantage Beneficiaries, by Area of Residence
MDCR PHYSSUPP MA 4. Medicare Physicians, Non-Physician Practitioners, and Suppliers: Utilization for Medicare Advantage Beneficiaries, by Place of Service
MDCR PHYSSUPP MA 5. Medicare Physicians, Non-Physician Practitioners, and Suppliers: Utilization for Medicare Advantage Beneficiaries, by Restructured BETOS Classification System
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
VLM-3R Training Data
Training QA data for VLM-3R: vsibench_train/ (VSI-Bench-style tasks) and vstibench_train/ (VSTI-Bench tasks over ScanNet train split).
Erratum (2026-07-13): corrected camera-position ground truth
A bug in the QA generation pipeline (reported by Jacob Yeung, CMU) extracted the camera center from camera-to-world poses using -R.T @ t instead of pose[:3, 3]. Answers in five vstibench_train files depended on the camera's world position and have… See the full description on the dataset page: https://huggingface.co/datasets/Journey9ni/VLM-3R-DATA.