4 datasets found

Underwater Object Detection Dataset
kaggle.com
zip
Updated Feb 12, 2022
Share
Facebook
Twitter
Email
Click to copy link
Link copied
Cite
Slavko Prytula (2022). Underwater Object Detection Dataset [Dataset]. https://www.kaggle.com/datasets/slavkoprytula/aquarium-data-cots
Explore at:
zip(69834944 bytes)Available download formats
Dataset updated
Feb 12, 2022
Authors
Slavko Prytula
License
Attribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically
Description
https://i.imgur.com/s4PgS4X.gif" alt="CreateML Output">

Info

The dataset contains 7 classes of underwater creatures with provided bboxes locations for every animal. The dataset is already split into the train, validation, and test sets.

Data

It includes 638 images. - Creatures are annotated in YOLO v5 PyTorch format

Pre-Processing

The following pre-processing was applied to each image: - Auto-orientation of pixel data (with EXIF-orientation stripping) - Resize to 1024x1024 (Fit within)

Class Breakdown

The following classes are labeled: ['fish', 'jellyfish', 'penguin', 'puffin', 'shark', 'starfish', 'stingray']. Most images contain multiple bounding boxes.

https://i.imgur.com/lFzeXsT.png" alt="Class Balance">
f
Data from: SyMANTIC: An Efficient Symbolic Regression Method for...
acs.figshare.com
zip
Updated Feb 4, 2025
Share
Facebook
Twitter
Email
Click to copy link
Link copied
Cite
Madhav R. Muthyala; Farshud Sorourifar; You Peng; Joel A. Paulson (2025). SyMANTIC: An Efficient Symbolic Regression Method for Interpretable and Parsimonious Model Discovery in Science and Beyond [Dataset]. http://doi.org/10.1021/acs.iecr.4c03503.s001
Explore at:
zipAvailable download formats
Unique identifier
https://doi.org/10.1021/acs.iecr.4c03503.s001
Dataset updated
Feb 4, 2025
Dataset provided by
ACS Publications
Authors
Madhav R. Muthyala; Farshud Sorourifar; You Peng; Joel A. Paulson
License
Attribution-NonCommercial 4.0 (CC BY-NC 4.0)https://creativecommons.org/licenses/by-nc/4.0/
License information was derived automatically
Description
Symbolic regression (SR) is an emerging branch of machine learning focused on discovering simple and interpretable mathematical expressions from data. Although a wide-variety of SR methods have been developed, they often face challenges such as high computational cost, poor scalability with respect to the number of input dimensions, fragility to noise, and an inability to balance accuracy and complexity. This work introduces SyMANTIC, a novel SR algorithm that addresses these challenges. SyMANTIC efficiently identifies (potentially several) low-dimensional descriptors from a large set of candidates (from ∼105 to ∼1010 or more) through a unique combination of mutual information-based feature selection, adaptive feature expansion, and recursively applied l0-based sparse regression. In addition, it employs an information-theoretic measure to produce an approximate set of Pareto-optimal equations, each offering the best-found accuracy for a given complexity. Furthermore, our open-source implementation of SyMANTIC, built on the PyTorch ecosystem, facilitates easy installation and GPU acceleration. We demonstrate the effectiveness of SyMANTIC across a range of problems, including synthetic examples, scientific benchmarks, real-world material property predictions, and chaotic dynamical system identification from small datasets. Extensive comparisons show that SyMANTIC uncovers similar or more accurate models at a fraction of the cost of existing SR methods.
Imbalanced Cifar-10
kaggle.com
zip
Updated Jun 17, 2023
Share
Facebook
Twitter
Email
Click to copy link
Link copied
Cite
Akhil Theerthala (2023). Imbalanced Cifar-10 [Dataset]. https://www.kaggle.com/datasets/akhiltheerthala/imbalanced-cifar-10
Explore at:
zip(807146485 bytes)Available download formats
Dataset updated
Jun 17, 2023
Authors
Akhil Theerthala
Description
This dataset is a modified version of the classic CIFAR 10, deliberately designed to be imbalanced across its classes. CIFAR 10 typically consists of 60,000 32x32 color images in 10 classes, with 5000 images per class in the training set. However, this dataset skews these distributions to create a more challenging environment for developing and testing machine learning algorithms. The distribution can be visualized as follows,

https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F7862887%2Fae7643fe0e58a489901ce121dc2e8262%2FCifar_Imbalanced_data.png?generation=1686732867580792&alt=media" alt="">

The primary purpose of this dataset is to offer researchers and practitioners a platform to develop, test, and enhance algorithms' robustness when faced with class imbalances. It is especially suited for those interested in binary and multi-class imbalance learning, anomaly detection, and other relevant fields.

The imbalance was created synthetically, maintaining the same quality and diversity of the original CIFAR 10 dataset, but with varying degrees of representation for each class. Details of the class distributions are included in the dataset's metadata.

This dataset is beneficial for: - Developing and testing strategies for handling imbalanced datasets. - Investigating the effects of class imbalance on model performance. - Comparing different machine learning algorithms' performance under class imbalance.

Usage Information:

The dataset maintains the same format as the original CIFAR 10 dataset, making it easy to incorporate into existing projects. It is organised in a way such that the dataset can be integrated into PyTorch ImageFolder directly. You can load the dataset in Python using popular libraries like NumPy and PyTorch.

License: This dataset follows the same license terms as the original CIFAR 10 dataset. Please refer to the official CIFAR 10 website for details.

Acknowledgments: We want to acknowledge the creators of the CIFAR 10 dataset. Without their work and willingness to share data, this synthetic imbalanced dataset wouldn't be possible.
n
OpenABC-D: A Large-Scale Dataset For Machine Learning Guided Integrated...
ultraviolet.library.nyu.edu
bin, png, zip
Updated Jul 22, 2025
Share
Facebook
Twitter
Email
Click to copy link
Link copied
Cite
Animesh Basak Chowdhury; Animesh Basak Chowdhury; Benjamin Tan; Ramesh Karri; Siddarth Garg; Benjamin Tan; Ramesh Karri; Siddarth Garg (2025). OpenABC-D: A Large-Scale Dataset For Machine Learning Guided Integrated Circuit Synthesis [Dataset]. http://doi.org/10.58153/mw6q2-a8p15
Explore at:
bin, zip, pngAvailable download formats
Unique identifier
https://doi.org/10.58153/mw6q2-a8p15
Dataset updated
Jul 22, 2025
Dataset provided by
New York University
Authors
Animesh Basak Chowdhury; Animesh Basak Chowdhury; Benjamin Tan; Ramesh Karri; Siddarth Garg; Benjamin Tan; Ramesh Karri; Siddarth Garg
License
https://opensource.org/licenses/BSD-3-Clausehttps://opensource.org/licenses/BSD-3-Clause
Description
OpenABC-D is a large-scale labeled dataset generated by synthesizing open source hardware IPs using state-of-art logic synthesis tool yosys-abc. We consider 29 open-source hardware IP designs collected from various sources (MIT-CEP, IWLS, OpenROAD, OpenPiton etc) and synthesized them with 1500 random synthesis flows (we call them synthesis recipes).
Each synthesis flow has a predefined length L (L=20, in our case). We preserved all AIGs: starting, intermediate and final AIGs with labels like number of nodes, longest path, sequence of atomic synthesis transformations (rewrite, refactor, balance etc.) along with graph statistics, area and delay of final AIG.
We converted the AIGs in pytorch data format that can be directly used by a machine learning engineer lessening the effort of costly labeled data generation and pre-processing. OpenABC-D can be used for a variety of learning tasks on logic synthesis such as
Predicting quality of result (QoR) performance of a *synthesis recipe* on a hardware IP.
Area and delay prediction post techonolgy mapping.
Learn functional and structural features of AIG using self-supervised labels (useful for tasks like RL-based logic synthesis)
Our dataset can easily be used for graph-based machine learning framework like Pytorch-Geometric.
Not seeing a result you expected?
Learn how you can add new datasets to our index.

Facebook

Twitter

Click to copy link

Link copied

Cite

Slavko Prytula (2022). Underwater Object Detection Dataset [Dataset]. https://www.kaggle.com/datasets/slavkoprytula/aquarium-data-cots

Underwater Object Detection Dataset

Yolov5 PyTorch format underwater life dataset for object detection

Explore at:

9 scholarly articles cite this dataset (View in Google Scholar)

zip(69834944 bytes)Available download formats

Dataset updated

Feb 12, 2022

Authors

Slavko Prytula

License

Attribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically

Description

https://i.imgur.com/s4PgS4X.gif" alt="CreateML Output">

Info

The dataset contains 7 classes of underwater creatures with provided bboxes locations for every animal. The dataset is already split into the train, validation, and test sets.

Data

It includes 638 images. - Creatures are annotated in YOLO v5 PyTorch format

Pre-Processing

The following pre-processing was applied to each image: - Auto-orientation of pixel data (with EXIF-orientation stripping) - Resize to 1024x1024 (Fit within)

Class Breakdown

The following classes are labeled: ['fish', 'jellyfish', 'penguin', 'puffin', 'shark', 'starfish', 'stingray']. Most images contain multiple bounding boxes.

https://i.imgur.com/lFzeXsT.png" alt="Class Balance">

Clear search

Close search

Google apps

Main menu

Underwater Object Detection Dataset

Info

Data

Pre-Processing

Class Breakdown

Data from: SyMANTIC: An Efficient Symbolic Regression Method for...

Imbalanced Cifar-10

OpenABC-D: A Large-Scale Dataset For Machine Learning Guided Integrated...

Underwater Object Detection Dataset

Yolov5 PyTorch format underwater life dataset for object detection

Info

Data

Pre-Processing

Class Breakdown