Facebook
TwitterAttribution-NonCommercial-ShareAlike 4.0 (CC BY-NC-SA 4.0)https://creativecommons.org/licenses/by-nc-sa/4.0/
License information was derived automatically
Please ensure to cite the paper when utilizing the dataset in a research study. Refer to the paper link or BibTeX provided below.
This repository contains comprehensive datasets for soil classification and recognition research. The Original Dataset comprises soil images sourced from various online repositories, which have been meticulously cleaned and preprocessed to ensure data quality and consistency. To enhance the dataset's size and diversity, we employed Generative Adversarial Networks (GANs), specifically the CycleGAN architecture, to generate synthetic soil images. This augmented collection is referred to as the CyAUG Dataset. Both datasets are specifically designed to advance research in soil classification and recognition using state-of-the-art deep learning methodologies.
This dataset was curated as part of the research study titled "An advanced artificial intelligence framework integrating ensembled convolutional neural networks and Vision Transformers for precise soil classification with adaptive fuzzy logic-based crop recommendations" by Farhan Sheth, Priya Mathur, Amit Kumar Gupta, and Sandeep Chaurasia, published in Engineering Applications of Artificial Intelligence.
Application produced by this research is available at:
Note: If you are using any part of this project; dataset, code, application, then please cite the work as mentioned in the Citation section below.
Both dataset consists of images of 7 different soil types.
The Soil Classification Dataset is structured to facilitate the classification of various soil types based on images. The dataset includes images of the following soil types:
The dataset is organized into folders, each named after a specific soil type, containing images of that soil type. The images vary in resolution and quality, providing a diverse set of examples for training and testing classification models.
If you are using any of the derived dataset, please cite the following paper:
@article{SHETH2025111425,
title = {An advanced artificial intelligence framework integrating ensembled convolutional neural networks and Vision Transformers for precise soil classification with adaptive fuzzy logic-based crop recommendations},
journal = {Engineering Applications of Artificial Intelligence},
volume = {158},
pages = {111425},
year = {2025},
issn = {0952-1976},
doi = {https://doi.org/10.1016/j.engappai.2025.111425},
url = {https://www.sciencedirect.com/science/article/pii/S0952197625014277},
author = {Farhan Sheth and Priya Mathur and Amit Kumar Gupta and Sandeep Chaurasia},
keywords = {Soil classification, Crop recommendation, Vision transformers, Convolutional neural network, Transfer learning, Fuzzy logic}
}
Facebook
TwitterU.S. Government Workshttps://www.usa.gov/government-works
License information was derived automatically
This dataset is a digital soil survey and generally is the most detailed level of soil geographic data developed by the National Cooperative Soil Survey. The information was prepared by digitizing maps, by compiling information onto a planimetric correct base and digitizing, or by revising digitized maps using remotely sensed and other information.
This dataset consists of georeferenced digital map data and computerized attribute data. The map data are in a soil survey area extent format and include a detailed, field verified inventory of soils and miscellaneous areas that normally occur in a repeatable pattern on the landscape and that can be cartographically shown at the scale mapped. A special soil features layer (point and line features) is optional. This layer displays the location of features too small to delineate at the mapping scale, but they are large enough and contrasting enough to significantly influence use and management. The soil map units are linked to attributes in the National Soil Information System relational database, which gives the proportionate extent of the component soils and their properties.
SSURGO depicts information about the kinds and distribution of soils on the landscape. The soil map and data used in the SSURGO product were prepared by soil scientists as part of the National Cooperative Soil Survey.
Facebook
TwitterAttribution-NonCommercial 4.0 (CC BY-NC 4.0)https://creativecommons.org/licenses/by-nc/4.0/
License information was derived automatically
Brief Description: This dataset contains 1 million simulated soil samples from various locations around the globe. Each sample includes data on soil texture, pH, organic matter content, moisture content, bulk density, nutrient levels (N, P, K), cation exchange capacity, electrical conductivity, color, porosity, and water holding capacity. Designed for environmental scientists, agronomists, and data scientists, this dataset is ideal for research, machine learning models, and educational purposes. Purpose: To provide a comprehensive soil dataset for environmental and agricultural research, including machine learning and data analysis applications. Data Collection Method: Simulated data generated using Python with realistic ranges and distributions based on common soil characteristics.
Usage Examples
Predictive modeling of soil properties.
Classification of soil types based on texture and nutrient content.
Analysis of soil health and fertility across different geographic locations.
File Descriptions
soil_data.csv - The main dataset file containing 1 million rows of soil data across 17 features.
Data Fields
Soil_ID: Unique identifier for each soil sample.
Location_Latitude and Location_Longitude: Geographic coordinates of the soil sample.
Depth_cm: Depth at which the soil sample was collected (cm).
Texture: Soil texture classification (sandy, loamy, clayey).
pH: Soil pH level.
Organic_Matter_%: Percentage of organic matter in the soil.
Moisture_Content_%: Soil moisture content percentage.
Bulk_Density_g/cm³: Soil bulk density (g/cm³).
Nitrogen_N_ppm, Phosphorus_P_ppm, Potassium_K_ppm: Nutrient levels in parts per million (ppm).
Cation_Exchange_Capacity_meq/100g: Soil's ability to hold positively charged ions (meq/100g).
Electrical_Conductivity_dS/m: Soil electrical conductivity (dS/m).
Soil_Color: Color of the soil (brown, red, black, yellow).
Porosity_%: Percentage of pore space in the soil.
Water_Holding_Capacity_%: Soil's water holding capacity percentage.
Acknowledgments
If your dataset generation was inspired by specific studies, data sources, or methodologies, acknowledge them here.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
https://storage.googleapis.com/kagglesdsdata/datasets/1262694/2104731/GridMaps250m_Info.png?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210410%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20210410T121915Z&X-Goog-Expires=172799&X-Goog-SignedHeaders=host&X-Goog-Signature=7d230635b7350fa1b67890294161d9e89660b6343a7e04af766918f88b85d968df94c8053446b05ac8f6871b5548aab08619101442af0289f5d7e00284d48fd93612f66b5598a3ba443256fa28b3f4df537d54516c30e2af3fbfcbd64e406852f9d1875b6dbdf8548d5deb4df4f8dd8331311c7de89bfe7898bde84536008098f03815d099571a3b7fad845ddae94049f877cefec13ec502323879ed51a58a3b5dd055d4cbca9cdfa002c157222743d23678cfab9e658fedf1968a23bc71e56b434f473e96fc08a6411ef7bbc938d93f26e9651f7fd99721aff0876a1abe9e1fc642abe3843869bf30fd56b2cdce02831ac5f9570d9ef64296eee4864b0172ad" alt="IMG">
Maps of clay, silt and sand contents (g kg-1) were predicted at 0-20 cm, 20-60 cm and 60-100 cm depths intervals by random forest regression in Google Earth Engine. Gridded soil information covers a part of the Midwest Brazil, from 12° S to 20° S and from 45° W to 54° W, and is available with 250m resolution. The maps were cross-validated and had Coefficient of Determination ranging from 0.64 to 0.85 at all depth intervals.
Poppiel, Raúl Roberto; Lacerda, Marilusa Pinto Coelho; Safanelli, José Lucas; Rizzo, Rodnei; Pereira de Oliveira Junior, Manuel; Novais, Jean Jesus; Dematte, Jose Alexandre (2020), “250 m-gridded soil texture at multiple depths of Midwest Brazil”, Mendeley Data, V4, doi: 10.17632/52cfcm3xr7.4
Facebook
TwitterThis data set provides soil maps for the United States (US) (including Alaska), Canada, Mexico, and a part of Guatemala. The map information content includes maximum soil depth and eight soil attributes including sand, silt, and clay content, gravel content, organic carbon content, pH, cation exchange capacity, and bulk density for the topsoil layer (0-30 cm) and the subsoil layer (30-100 cm). The spatial resolution is 0.25 degree. The Unified North American Soil Map (UNASM) combined information from the state-of-the-art US General Soil Map (STATSGO2) and Soil Landscape of Canada (SLCs) databases. The area not covered by these data sets was filled by using the Harmonized World Soil Database version 1.21 (HWSD1.21). The Northern Circumpolar Soil Carbon (NCSCD) database was used to provide more accurate and up-to-date soil organic carbon information for the high-latitude permafrost region and was combined with soil organic carbon content derived from the UNASM (Liu et al., 2013). The UNASM data were utilized in the North American Carbon Program (NACP) Multi-Scale Synthesis and Terrestrial Model Intercomparison Project (MsTMIP) as model input driver data (Huntzinger et al., 2013). The driver data were used by 22 terrestrial biosphere models to run baseline and sensitivity simulations. The compilation of these data was facilitated by the NACP Modeling and Synthesis Thematic Data Center (MAST-DC). MAST-DC was a component of the NACP (www.nacarbon.org) designed to support NACP by providing data products and data management services needed for modeling and synthesis activities.
Facebook
TwitterThe U.S. Department of Agriculture, Agriculture and Agri-Food Canada, the Russian Academy of Agricultural Sciences, the University of Copenhagen Institute of Geography, the European Soil Bureau, the University of Manchester Institute of Landscape Ecology, MTT Agrifood Research Finland, and the Agricultural Research Institute Iceland have shared data and expertise in order to develop the Northern and Mid Latitude Soil Database (Cryosol Working Group, 2001). This database was the source of data for the current product. The spatial coverage of the Northern and Mid Latitude Soil Database is the polar and mid-latitude regions of the northern hemisphere: Alaska, Canada, Conterminous United States, Eurasia (except Italy), Greenland, Iceland, Kazakstan, Mexico, Mongolia, Italy, and Svalbard. The Northern and Mid Latitude Soil Database represents the proportion (percentage) of polygon encompassed by the dominant soil or nonsoil. Soils include turbels, orthels, histels, histosols, mollisols, vertisols, aridisols, andisols, entisols, spodosols, inceptisols (and hapludolls), alfisols (cryalf and udalf), natric great groups, aqu-suborders, glaciers, and rocklands. Also included are data on the circumpolar distribution of gelisols (turbels, orthels, and histels), and the ice content (low, medium, or high) of circumpolar soil materials (from the International Permafrost Association, 1997). The resulting maps show the dominant soil of the spatial polygon unless the polygon is over 90 percent rock or ice. Data are in the U.S. soil classification system and includes the distribution of soil types (%) within a map unit (polygon). Data are available in ESRI shapefile format and include the same attribute values with the exception of Italy, which does not contain distribution values.
Facebook
TwitterU.S. Government Workshttps://www.usa.gov/government-works
License information was derived automatically
The dataset consists of three raster GeoTIFF files describing the following soil properties in the US: available water capacity, field capacity, and soil porosity. The input data were obtained from the gridded National Soil Survey Geographic (gNATSGO) Database and the Gridded Soil Survey Geographic (gSSURGO) Database with Soil Data Development tools provided by the Natural Resources Conservation Service. The soil characteristics derived from the databases were Available Water Capacity (AWC), Water Content (one-third bar) (WC), and Bulk Density (one-third bar) (BD) aggregated as weighted average values in the upper 1 m of soil. AWC and WC layers were converted to mm/m to express respectively available water capacity and field capacity in 1 m of soil, and BD layer was used to produce soil porosity raster assuming that the average particle density of soils is equal to 2.65 g/cm3. For each soil property, soil maps with CONUS, Alaska, and Hawaii geographic coverages were derived from s ...
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This data set is a digital soil survey and generally is the most detailed level of soil geographic data developed by the National Cooperative Soil Survey. The information was prepared by digitizing maps, by compiling information onto a planimetric correct base and digitizing, or by revising digitized maps using remotely sensed and other information. This data set consists of georeferenced digital map data and computerized attribute data. The map data are in a soil survey area extent format and include a detailed, field verified inventory of soils and miscellaneous areas that normally occur in a repeatable pattern on the landscape and that can be cartographically shown at the scale mapped. A special soil features layer (point and line features) is optional. This layer displays the location of features too small to delineate at the mapping scale, but they are large enough and contrasting enough to significantly influence use and management. The soil map units are linked to attributes in the National Soil Information System relational database, which gives the proportionate extent of the component soils and their properties.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This dataset contains high-quality images of soil surfaces categorized into three moisture levels—Wet, Moderate, and Dry—captured under natural outdoor lighting in Sirajganj, Bangladesh. Images were collected at seven time intervals (0 min, 30 min, 1 hr, 2 hr, 4 hr, 5 hr, and 7+ hr after saturation) using a Sony Xperia 1 Mark II smartphone. A total of 1,177 raw images were captured, with blurry, noisy, and low-quality photos removed during pre-processing. The dataset reflects real-world agricultural conditions and serves as a benchmark for training machine learning and deep learning models for non-invasive soil moisture classification. Subject Areas: Computer Science, Agriculture Science, AI, Computer Vision, Environmental Monitoring, Pattern Recognition Data Format: JPG images (raw and filtered) Data Collection: Captured using Sony Xperia 1 Mark II under natural outdoor lighting in multiple soil locations. Organized into three labeled categories (Wet, Moderate, Dry) based on time intervals after saturation. Can be split into training and testing sets (recommended 80:20 ratio). Usage Notes: Ideal for developing AI models in soil moisture classification, precision irrigation scheduling, and image-based environmental monitoring. Supports affordable, sensor-free soil analysis for sustainable farming practices, particularly in resource-limited settings.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This is a synthetic dataset designed for training and evaluating machine learning models that classify whether certain soil and climate conditions are compatible for growing a given crop. The data includes environmental and soil features for a variety of crops commonly grown in India.
This dataset simulates the relationship between crop types, soil properties, and climate conditions. The goal is to predict whether a given combination of soil, weather, and crop factors is compatible for cultivation — making it ideal for binary classification tasks.
It is especially useful for: - Crop recommendation systems - Soil-climate compatibility prediction - Educational ML applications in agriculture
soil_climate_crop_data.csv: Main dataset file with synthetic records| Column Name | Description |
|---|---|
Crop_Type | Crop name (19 commonly grown crops) |
Soil_Type | Soil classification (4 common Indian soil types) |
Farm_Size_Acres | Size of the farm (in acres) |
Irrigation_Available | Boolean (Yes/No) representing irrigation access |
Soil_pH | Soil pH level |
Soil_Nitrogen | Nitrogen level in soil (ppm) |
Soil_Organic_Matter | Organic matter content (%) |
Temperature | Average temperature (°C) |
Rainfall | Annual rainfall (mm) |
Humidity | Relative humidity (%) |
Compatible | Binary label: 1 = Compatible, 0 = Not Compatible |
Wheat, Rice, Maize, Soybean, Niger, Urd, Summer Paddy, Gram, Tiwra, Millets, Arhar, Mustard, Jwar, Moong, Kulthi, Groundnut, Masoor, Til, Pea
This dataset is synthetically generated using domain-inspired rules and randomized inputs.
It is not based on real-world data but reflects typical patterns observed in Indian agricultural settings. The labels for compatibility are assigned based on plausible thresholds and combinations of soil, climate, and crop requirements.
License: CC0 1.0 Universal (Public Domain Dedication)
You are free to use, modify, and share this dataset for any purpose without restrictions.
Created by: Rajeev
Open to feedback and collaboration!
Facebook
TwitterOpen Government Licence - Canada 2.0https://open.canada.ca/en/open-government-licence-canada
License information was derived automatically
These Soil Mapping Data Packages include 1. a Soil Map dataset which includes the equivalents to Soil Project Boundaries, Soil Survey Spatial View mapping polygons with attributes from the Soil Name and Layer Files, plus + A Soil Site dataset which includes soil pit site information and detailed soil pit descriptions and any associated lab analyses, and + The Soil Data Dictionary which documents the fields and allowable codes within the data. The Soil Map geodatabase contains the 'best available' data ranging from 1:20,000 scale to 1:250,000 scale with overlapping data removed. The choice of the datasets that remain is based on connectivity to the soil attributes (soil name and layer files), map scale and survey date. (Note: the BC Soil Landscapes of Canada (BCSLC) 1:1,000,000 data has not been included in the Soil_Map or SIFT, but is available from: CANSIS. (A complete soils data package with overlapping soil survey mapping and BCSLC is available on request. Note that the soil survey data with attributes can also be viewed interactively in the [Soil Information Finder Tool](The Soil Map dataset is also available for interactive map viewing or as KMZs from the Soil Information Finder Tool website.
Facebook
Twitterhttps://pasteur.epa.gov/license/sciencehub-license.htmlhttps://pasteur.epa.gov/license/sciencehub-license.html
A workbook of all the soils data collected near Holton, Kansas, in agricultural fields. Laboratory analysis of soil properties was completed by Ward Labs in Kearny Nebraska. Isotope analysis of soils was completed in Integrated Stable Isotope Research Facility operated by US Environmental Protection Agency. The goal of this project was to evaluate if Soil Health Principles can reduce the risk of nitrate leaching from agricultural fields. This effort was a collaborative project between EPA Region 7, EPA Office of Research and Development, and Kansas Department of Health and Environment (KDHE).
Discussion of the project generating these data is available on the KDHE website: https://storymaps.arcgis.com/stories/1efcfe1924fc4daf85a7958c0a41fa5a
It can also be found on the KDHE Watershed Management Section at the end of the What we Do section.
https://www.kdhe.ks.gov/974/Watershed-Management-Section
Facebook
TwitterA global data set of soil types is available at 1-degree latitude by 1-degree longitude resolution. There are 26 soil units based on Zobler?s assessment of FAO Soil Units (Zobler, 1986). The data set was compiled as part of an effort to improve modeling of the hydrologic cycle portion of global climate models. A more extensive version of these data, including 106 soil units as well as soil texture and slope, is available from NCAR, Scientific Computing Division, Data Support Section; the more extensive data set is entitled "Staub and Rosenweig's GISS Soil & Sfc Slope, 1-Deg" [http://www.dss.ucar.edu/datasets/ds770.0/]. A help file prepared by Matthews and Fung (1987) (soil1x1.help) is provided as a companion file. Image of 26 soil types available at 1-degree by 1-degree resolution. Additional documentation from Zobler?s assessment of FAO soil units is available from the NASA Center for Scientific Information
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This is a public compendium of global, regional, national and sub-national soil samples and/or soil profile datasets (points with Observations and Measurements of soil properties and characteristics). Datasets listed here, assuming compatible open license, are afterwards imported into the Global compilation of soil chemical and physical properties and soil classes and eventually used to create a better open soil information across countries. Please feel free to contribute entries. See GitHub repository for more detailed instructions.
To see the most up-to-date version of this compendium please visit: https://soildb.OpenLandMap.org/
To read about the import steps and quality control, refer to the Hengl et al. (2026; ESSD).
Facebook
TwitterThis data set provides gridded data for selected soil parameters derived from data and methods developed by the Global Soil Data Task, an international collaborative project with the objective of making accurate and appropriate data relating to soil properties accessible to the global change research community. The task was coordinated by the International Geosphere-Biosphere Programme (IGBP-DIS). The data in this data set were produced by the International Satellite Land-Surface Climatology Project, Initiative II (ISLSCP II) staff from data obtained from the Oak Ridge National Laboratory Distributed Active Archive Center (ORNL DAAC, http://daac.ornl.gov/). See the related data sets section below. Two-dimensional gridded maps of selected soil parameters, including soil texture, at a 1.0 by 1.0 degree spatial resolution and for two soil depths are provided. All data layers have been adjusted to match the ISLSCP II land/water mask. There are 36 data files with this data set.
Facebook
TwitterThis data set is a digital soil survey and generally is the mostdetailed level of soil geographic data developed by the NationalCooperative Soil Survey. The information was prepared by digitizingmaps, by compiling information onto a planimetric correct baseand digitizing, or by revising digitized maps using remotelysensed and other information.This data set consists of georeferenced digital map data andcomputerized attribute data. The map data are in a soil survey areaextent format and include a detailed, field verified inventoryof soils and miscellaneous areas that normally occur in a repeatablepattern on the landscape and that can be cartographically shown atthe scale mapped. The soil map units are linked to attributes in theNational Soil Information System relational database, which givesthe proportionate extent of the component soils and their properties.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
The Multiclass Soil Image Dataset is a curated collection of soil surface images intended to support research and experimentation in soil type recognition and agricultural analysis. The dataset brings together visually distinct soil categories that commonly appear in agricultural fields and natural environments, providing a diverse visual representation of soil conditions.
This dataset contains 1,378 soil images, each corresponding to a specific soil category. The images capture variations in texture, color, granularity, and surface structure that are characteristic of different soil types. Such diversity makes the dataset valuable for studying soil variability and supporting image-based analysis in agricultural and environmental applications.
The dataset is suitable for academic research, educational purposes, and experimental evaluation related to soil analysis, precision agriculture, land management, and environmental monitoring. By offering a balanced representation of multiple soil classes, it enables comparative studies and benchmarking of image-based soil recognition approaches without embedding assumptions about specific algorithms or modeling techniques.
Alluvial Soil – Images representing fertile soils commonly found along river plains and floodplains.
Black Soil – Images depicting dark-colored soils rich in organic matter and moisture retention.
Cinder Soil – Images showing coarse, granular soils formed from volcanic or cinder-like materials.
Clay Soil – Images illustrating fine-grained soils with dense texture and compact structure.
Laterite Soil – Images representing iron-rich soils typically found in tropical and subtropical regions.
Peat Soil – Images depicting organic-rich soils formed from partially decomposed plant material.
Red Soil – Images showing soils characterized by reddish coloration due to iron content.
Yellow Soil – Images representing soils with yellowish tones commonly observed in humid regions.
Facebook
TwitterThe Global Gridded Surfaces of Selected Soil Characteristics (IGBP-DIS) data set contains 7 data surfaces: soil-carbon density, total nitrogen density, field capacity, wilting point, profile available water capacity, thermal capacity, and bulk density. All the surfaces are global, at a resolution of 5x5 arc-minutes, in ASCIIGRID format for ARC INFO. Each file contains a single ASCII array in a geographic (lat/long) projection. The ascii files consist of header information containing a set of keywords, followed by cell values in row-major order. These data surfaces were generated by the SoilData System, which was developed by the Global Soil Data Task of the International Geosphere-Biosphere Programme (IGBP) Data and Information Services (DIS). The SoilData System generates soil information and maps for geographic regions at soil depths and resolutions selected by the user. Derived surfaces of selected soil characteristics are suitable for modeling and inventory purposes. The data surfaces are also distributed as part of the Global Soil Data Products CD-ROM. The SoilData System uses a statistical bootstrapping approach to link the pedon records in the Global Pedon Database to the FAO/UNESCO Digital Soil Map of the World. It can generate maps and output data sets for a range of original and derived soil parameters, such as carbon and nitrogen density, thermal conductivity, and water-holding capacity, for any part of the world at user-selected depth ranges. The digital output can be at any resolution (in increments of 5').
Facebook
TwitterU.S. Government Workshttps://www.usa.gov/government-works
License information was derived automatically
Soil Landscapes of the United States, or SOLUS, is a national map product developed by the National Cooperative Soil Survey that is focused on providing a consistent set of spatially continuous soil property maps to support large scope soil investigations and land use decisions. SOLUS maps use a digital soil mapping framework that combines multiple sources of soil survey data with environmental covariate data and machine learning. Digital soil mapping is the production of georeferenced soil databases based on the quantitative relationships between soil measurements made in the field or laboratory and environmental data. Numerical models use the quantitative relationships to predict the spatial distribution of either discrete soil classes, such as map units, or continuous soil properties, such as clay content.
SOLUS maps use continuous property mapping, which predicts soil physical or chemical properties in horizontal and vertical dimensions. The soil properties are represented across a continuous range of values. Raster datasets of select soil properties can be predicted at specified depths or depth intervals. Continuous soil property maps such as SOLUS provide critical natural resource information to support environmental researchers and modelers, conservationists, and others making land management decisions. SOLUS will be updated annually with improved data and methodology.
The SOLUS dataset includes 20 different soil properties (listed below) with most properties predicted for seven standard depths (0, 5, 15, 30, 60, 100, and 150 cm).
Properties included in SOLUS100:
Bulk density (oven dry) Calcium carbonate Cation Exchange Capacity (pH 7) Clay Coarse sand Electrical Conductivity (sat. paste) Effective cation exchange capacity Fine sand Gypsum (in <20 mm fraction) Medium sand pH (1:1 method) Rock content Sand Sodium adsorption ratio Silt Soil organic carbon Very coarse sand Very fine sand Depth to bedrock Depth to restriction
Facebook
TwitterSoil texture classes (USDA system) for 6 soil depths (0, 10, 30, 60, 100 and 200 cm) at 250 m Derived from predicted soil texture fractions using the soiltexture package in R. Processing steps are described in detail here. Antarctica is not included. To access and visualize maps outside of Earth Engine, use this page. If you discover a bug, artifact or inconsistency in the LandGIS maps or if you have a question please use the following channels: Technical issues and questions about the code General questions and comments
Facebook
TwitterAttribution-NonCommercial-ShareAlike 4.0 (CC BY-NC-SA 4.0)https://creativecommons.org/licenses/by-nc-sa/4.0/
License information was derived automatically
Please ensure to cite the paper when utilizing the dataset in a research study. Refer to the paper link or BibTeX provided below.
This repository contains comprehensive datasets for soil classification and recognition research. The Original Dataset comprises soil images sourced from various online repositories, which have been meticulously cleaned and preprocessed to ensure data quality and consistency. To enhance the dataset's size and diversity, we employed Generative Adversarial Networks (GANs), specifically the CycleGAN architecture, to generate synthetic soil images. This augmented collection is referred to as the CyAUG Dataset. Both datasets are specifically designed to advance research in soil classification and recognition using state-of-the-art deep learning methodologies.
This dataset was curated as part of the research study titled "An advanced artificial intelligence framework integrating ensembled convolutional neural networks and Vision Transformers for precise soil classification with adaptive fuzzy logic-based crop recommendations" by Farhan Sheth, Priya Mathur, Amit Kumar Gupta, and Sandeep Chaurasia, published in Engineering Applications of Artificial Intelligence.
Application produced by this research is available at:
Note: If you are using any part of this project; dataset, code, application, then please cite the work as mentioned in the Citation section below.
Both dataset consists of images of 7 different soil types.
The Soil Classification Dataset is structured to facilitate the classification of various soil types based on images. The dataset includes images of the following soil types:
The dataset is organized into folders, each named after a specific soil type, containing images of that soil type. The images vary in resolution and quality, providing a diverse set of examples for training and testing classification models.
If you are using any of the derived dataset, please cite the following paper:
@article{SHETH2025111425,
title = {An advanced artificial intelligence framework integrating ensembled convolutional neural networks and Vision Transformers for precise soil classification with adaptive fuzzy logic-based crop recommendations},
journal = {Engineering Applications of Artificial Intelligence},
volume = {158},
pages = {111425},
year = {2025},
issn = {0952-1976},
doi = {https://doi.org/10.1016/j.engappai.2025.111425},
url = {https://www.sciencedirect.com/science/article/pii/S0952197625014277},
author = {Farhan Sheth and Priya Mathur and Amit Kumar Gupta and Sandeep Chaurasia},
keywords = {Soil classification, Crop recommendation, Vision transformers, Convolutional neural network, Transfer learning, Fuzzy logic}
}