Facebook
TwitterThis dataset was created by JAYAPRAKASHPONDY
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This data set is a digital soil survey and generally is the most detailed level of soil geographic data developed by the National Cooperative Soil Survey. The information was prepared by digitizing maps, by compiling information onto a planimetric correct base and digitizing, or by revising digitized maps using remotely sensed and other information. This data set consists of georeferenced digital map data and computerized attribute data. The map data are in a soil survey area extent format and include a detailed, field verified inventory of soils and miscellaneous areas that normally occur in a repeatable pattern on the landscape and that can be cartographically shown at the scale mapped. A special soil features layer (point and line features) is optional. This layer displays the location of features too small to delineate at the mapping scale, but they are large enough and contrasting enough to significantly influence use and management. The soil map units are linked to attributes in the National Soil Information System relational database, which gives the proportionate extent of the component soils and their properties.
Facebook
Twitterhttps://dataverse.harvard.edu/api/datasets/:persistentId/versions/2.7/customlicense?persistentId=doi:10.7910/DVN/1PEEY0https://dataverse.harvard.edu/api/datasets/:persistentId/versions/2.7/customlicense?persistentId=doi:10.7910/DVN/1PEEY0
One of the obstacles in applying advanced crop simulation models such as DSSAT at a grid-based platform is the lack of gridded soil input data at various resolutions. Recently, there has been many efforts in scientific communities to develop spatially continuous soil database across the globe. The most representative example is the SoilGrids 1km released by ISRIC in 2014. In addition recent AfSIS project put a lot of efforts to develop more accurate soil database in Africa at high spatial resolution. Taking advantage of those two available high resolution soil databases (SoilGrids 1km and ISRIC-AfSIS at 1km resolution), this project aims to develop a set of DSSAT compatible soil profiles on 5 arc-minute grid (which is HarvestChoice’s standard grid). Six soil properties (bulk density, organic carbon, percentage of clay and silt, soil pH and cation exchange capacity) available from the original SoilGrids 1km or ISRIC-AfSIS were directly used as DSSAT inputs. We applied a pedo-transfer function to derive some soil hydraulic properties (saturated hydraulic conductivity, soil water content at field capacity, wilting point and saturation) which are critical to simulate crop growth. For other required variables, HarvestChoice’s HC27 database are used as a reference. Final outputs are provided in *.SOL file format (DSSAT soil database) for each country at 5-min resolution. In addition, uncertainty maps for organic carbon and soil water content at wilting points at the top 15 cm soil layers were generated to provide brief idea about accuracy of the final products. The generated soil properties were evaluated by visualizing their global maps and by comparing them with IIASA-IFPRI cropland map and AfSIS-GYGA’s available water content maps.
Facebook
TwitterThe International Soil Reference and Information Centre-World Inventory of Soil Emission Potentials (ISRIC-WISE) international soil profile data set consists of a homogenized, global set of 1,125 soil profiles for use by global modelers. These profiles provided the basis for the Global Pedon Database (GPDB) of the International Geosphere-Biosphere Programme (IGBP) - Data and Information System (DIS). The data set consists of a selection of 665 profiles originating from the Natural Resources Conservation Service (NRCS, Lincoln), 250 profiles obtained from the Food and Agriculture Organization (FAO, Rome), and 210 profiles from the reference collection of the International Soil Reference and Information Centre (ISRIC, Wageningen). All profiles are georeferenced and classified according to the 1974 Legend of the FAO-UNESCO Soil Map (FAC-UNESCO, 1974) of the World, as well as the 1988 Revised Legend of FAO-UNESCO (FAO, 1990). The data set includes information on soil classification, site data, soil horizon data, source of data, and methods used for determining analytical data. The data files are in a comma-delimited format. Data Citation: The data set should be cited as follows: Batjes, N. H. (ed). 2000. Global Soil Profile Data (ISRIC-WISE). Available on-line from the ORNL Distributed Active Archive Center, Oak Ridge National Laboratory, Oak Ridge, Tennessee, U.S.A.
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
This is a dataset download, not a document. The Open button will start the download.Detailed soil units from Soils Surveys covering nonfederal land conducted by the U.S. Natural Resource Conservation Service (NRCS) that differentiates mapped units on the basis of a range of physical, topographic, and chemical properties.
Facebook
TwitterThis data set provides soil maps for the United States (US) (including Alaska), Canada, Mexico, and a part of Guatemala. The map information content includes maximum soil depth and eight soil attributes including sand, silt, and clay content, gravel content, organic carbon content, pH, cation exchange capacity, and bulk density for the topsoil layer (0-30 cm) and the subsoil layer (30-100 cm). The spatial resolution is 0.25 degree.
The Unified North American Soil Map (UNASM) combined information from the state-of-the-art US General Soil Map (STATSGO2) and Soil Landscape of Canada (SLCs) databases. The area not covered by these data sets was filled by using the Harmonized World Soil Database version 1.21 (HWSD1.21). The Northern Circumpolar Soil Carbon (NCSCD) database was used to provide more accurate and up-to-date soil organic carbon information for the high-latitude permafrost region and was combined with soil organic carbon content derived from the UNASM (Liu et al., 2013).
The UNASM data were utilized in the North American Carbon Program (NACP) Multi-Scale Synthesis and Terrestrial Model Intercomparison Project (MsTMIP) as model input driver data (Huntzinger et al., 2013). The driver data were used by 22 terrestrial biosphere models to run baseline and sensitivity simulations. The compilation of these data was facilitated by the NACP Modeling and Synthesis Thematic Data Center (MAST-DC). MAST-DC was a component of the NACP (www.nacarbon.org) designed to support NACP by providing data products and data management services needed for modeling and synthesis activities.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This is a synthetic dataset designed for training and evaluating machine learning models that classify whether certain soil and climate conditions are compatible for growing a given crop. The data includes environmental and soil features for a variety of crops commonly grown in India.
This dataset simulates the relationship between crop types, soil properties, and climate conditions. The goal is to predict whether a given combination of soil, weather, and crop factors is compatible for cultivation — making it ideal for binary classification tasks.
It is especially useful for: - Crop recommendation systems - Soil-climate compatibility prediction - Educational ML applications in agriculture
soil_climate_crop_data.csv: Main dataset file with synthetic records| Column Name | Description |
|---|---|
Crop_Type | Crop name (19 commonly grown crops) |
Soil_Type | Soil classification (4 common Indian soil types) |
Farm_Size_Acres | Size of the farm (in acres) |
Irrigation_Available | Boolean (Yes/No) representing irrigation access |
Soil_pH | Soil pH level |
Soil_Nitrogen | Nitrogen level in soil (ppm) |
Soil_Organic_Matter | Organic matter content (%) |
Temperature | Average temperature (°C) |
Rainfall | Annual rainfall (mm) |
Humidity | Relative humidity (%) |
Compatible | Binary label: 1 = Compatible, 0 = Not Compatible |
Wheat, Rice, Maize, Soybean, Niger, Urd, Summer Paddy, Gram, Tiwra, Millets, Arhar, Mustard, Jwar, Moong, Kulthi, Groundnut, Masoor, Til, Pea
This dataset is synthetically generated using domain-inspired rules and randomized inputs.
It is not based on real-world data but reflects typical patterns observed in Indian agricultural settings. The labels for compatibility are assigned based on plausible thresholds and combinations of soil, climate, and crop requirements.
License: CC0 1.0 Universal (Public Domain Dedication)
You are free to use, modify, and share this dataset for any purpose without restrictions.
Created by: Rajeev
Open to feedback and collaboration!
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
https://storage.googleapis.com/kagglesdsdata/datasets/1262694/2104731/GridMaps250m_Info.png?X-Goog-Algorithm=GOOG4-RSA-SHA256&X-Goog-Credential=databundle-worker-v2%40kaggle-161607.iam.gserviceaccount.com%2F20210410%2Fauto%2Fstorage%2Fgoog4_request&X-Goog-Date=20210410T121915Z&X-Goog-Expires=172799&X-Goog-SignedHeaders=host&X-Goog-Signature=7d230635b7350fa1b67890294161d9e89660b6343a7e04af766918f88b85d968df94c8053446b05ac8f6871b5548aab08619101442af0289f5d7e00284d48fd93612f66b5598a3ba443256fa28b3f4df537d54516c30e2af3fbfcbd64e406852f9d1875b6dbdf8548d5deb4df4f8dd8331311c7de89bfe7898bde84536008098f03815d099571a3b7fad845ddae94049f877cefec13ec502323879ed51a58a3b5dd055d4cbca9cdfa002c157222743d23678cfab9e658fedf1968a23bc71e56b434f473e96fc08a6411ef7bbc938d93f26e9651f7fd99721aff0876a1abe9e1fc642abe3843869bf30fd56b2cdce02831ac5f9570d9ef64296eee4864b0172ad" alt="IMG">
Maps of clay, silt and sand contents (g kg-1) were predicted at 0-20 cm, 20-60 cm and 60-100 cm depths intervals by random forest regression in Google Earth Engine. Gridded soil information covers a part of the Midwest Brazil, from 12° S to 20° S and from 45° W to 54° W, and is available with 250m resolution. The maps were cross-validated and had Coefficient of Determination ranging from 0.64 to 0.85 at all depth intervals.
Poppiel, Raúl Roberto; Lacerda, Marilusa Pinto Coelho; Safanelli, José Lucas; Rizzo, Rodnei; Pereira de Oliveira Junior, Manuel; Novais, Jean Jesus; Dematte, Jose Alexandre (2020), “250 m-gridded soil texture at multiple depths of Midwest Brazil”, Mendeley Data, V4, doi: 10.17632/52cfcm3xr7.4
Facebook
Twitterhttps://pasteur.epa.gov/license/sciencehub-license.htmlhttps://pasteur.epa.gov/license/sciencehub-license.html
A workbook of all the soils data collected near Holton, Kansas, in agricultural fields. Laboratory analysis of soil properties was completed by Ward Labs in Kearny Nebraska. Isotope analysis of soils was completed in Integrated Stable Isotope Research Facility operated by US Environmental Protection Agency. The goal of this project was to evaluate if Soil Health Principles can reduce the risk of nitrate leaching from agricultural fields. This effort was a collaborative project between EPA Region 7, EPA Office of Research and Development, and Kansas Department of Health and Environment (KDHE).
Discussion of the project generating these data is available on the KDHE website: https://storymaps.arcgis.com/stories/1efcfe1924fc4daf85a7958c0a41fa5a
It can also be found on the KDHE Watershed Management Section at the end of the What we Do section.
https://www.kdhe.ks.gov/974/Watershed-Management-Section
Facebook
TwitterOpen Government Licence - Canada 2.0https://open.canada.ca/en/open-government-licence-canada
License information was derived automatically
These Soil Mapping Data Packages include 1. a Soil Map dataset which includes the equivalents to Soil Project Boundaries, Soil Survey Spatial View mapping polygons with attributes from the Soil Name and Layer Files, plus + A Soil Site dataset which includes soil pit site information and detailed soil pit descriptions and any associated lab analyses, and + The Soil Data Dictionary which documents the fields and allowable codes within the data. The Soil Map geodatabase contains the 'best available' data ranging from 1:20,000 scale to 1:250,000 scale with overlapping data removed. The choice of the datasets that remain is based on connectivity to the soil attributes (soil name and layer files), map scale and survey date. (Note: the BC Soil Landscapes of Canada (BCSLC) 1:1,000,000 data has not been included in the Soil_Map or SIFT, but is available from: CANSIS. (A complete soils data package with overlapping soil survey mapping and BCSLC is available on request. Note that the soil survey data with attributes can also be viewed interactively in the [Soil Information Finder Tool](The Soil Map dataset is also available for interactive map viewing or as KMZs from the Soil Information Finder Tool website.
Facebook
Twitterhttps://data.gov.tw/licensehttps://data.gov.tw/license
Provide soil map (coordinate system: TWD97) data file download.
Facebook
TwitterThis data set is a digital soil survey and generally is the most detailed level of soil geographic data developed by the National Cooperative Soil Survey. The information was prepared by digitizing maps, by compiling information onto a planimetric correct base and digitizing, or by revising digitized maps using remotely sensed and other information. This data set consists of georeferenced digital map data and computerized attribute data. The map data are in a soil survey area extent format and include a detailed, field verified inventory of soils and miscellaneous areas that normally occur in a repeatable pattern on the landscape and that can be cartographically shown at the scale mapped. A special soil features layer (point and line features) is optional. This layer displays the location of features too small to delineate at the mapping scale, but they are large enough and contrasting enough to significantly influence use and management. The soil map units are linked to attributes in the National Soil Information System relational database, which gives the proportionate extent of the component soils and their properties.
Facebook
TwitterThe Harmonized World Soil Database version 2.0 (HWSD v2.0) is a unique global soil inventory providing information on the morphological, chemical and physical properties of soils at approximately 1 km resolution. Its main objective is to serve as a basis for prospective studies on agro-ecological zoning, food security and climate change.
The Harmonized World Soil Database (HWSD) was established in 2008 by the International Institute for Applied Systems Analysis (IIASA) and FAO, and in partnership with International Soil Reference and Information Centre (ISRIC), the European Soil Bureau Network (ESBN) and the Institute for Soil Sciences Chinese Academy of Sciences (CAS). The data entry and harmonization within a Geographic Information System (GIS) was carried out at IIASA, with verification of the database undertaken by all partners. HWSD was then updated in 2013 (HWSD v1.2) and in 2023 (HWSD v2.0).
This updated version (HWSD v2.0) is built on the previous versions of HWSD with several improvements on (i) the data source that now includes several national soil databases, (ii) an enhanced number of soil attributes available for seven soil depth layers, instead of two in HWSD v1.2, and (iii) a common soil reference for all soil units (FAO1990 and the World Reference Base for Soil Resources). This contributes to a further harmonization of the database.
The GIS raster image file is linked to the soil attribute database. The HWSD v2.0 soil attribute database provides information on the soil unit composition for each of the near 30 000 soil association mapping units. The HWSD v2.0 Viewer, provided with the database, creates this link automatically and provides direct access to the soil attribute data and the soil association information.
Note: A tutorial for accessing HWSD ver. 2.0 using R (prepared by David Rossiter, June 2023) has been added as an 'associated resource' (NOTE: Needs the SQLite version of HWSD v2 as provided below).
Facebook
TwitterThese data were compiled to demonstrate new predictive mapping approaches and provide comprehensive gridded 30-meter resolution soil property maps for the Colorado River Basin above Hoover Dam. Random forest models related environmental raster layers representing soil forming factors with field samples to render predictive maps that interpolate between sample locations. Maps represented soil pH, texture fractions (sand, silt clay, fine sand, very fine sand), rock, electrical conductivity (ec), gypsum, CaCO3, sodium adsorption ratio (sar), available water capacity (awc), bulk density (dbovendry), erodibility (kwfact), and organic matter (om) at 7 depths (0, 5, 15, 30, 60, 100, and 200 cm) as well as depth to restrictive layer (resdept) and surface rock size and cover. Accuracy and error estimated using a 10-fold cross validation indicated a range of model performances with coefficient of variation (R2) for models ranging from 0.20 to 0.76 with mean of 0.52 and a standard deviation of 0.12. Models of pH, om and ec had the best accuracy (R2 > 0.6). Most texture fractions, CaCO3, and SAR models had R2 values from 0.5-0.6. Models of kwfact, dbovendry, resdept, rock models, gypsum and awc had R2 values from 0.4-0.5 excepting near surface models which tended to perform better. Very fine sands and 200 cm estimates for other models generally performed poorly (R2 from 0.2-0.4), and sample size for the 200 cm models was too low for reliable model building. More than 90% of the soils data used was sampled since 2000, but some older samples are included. Uncertainty estimates were also developed by creating relative prediction intervals, which allow end users to evaluate uncertainty easily.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
📘 Dataset Description — Soil With Crop Recommendations 🌾 Overview
This dataset provides a detailed collection of soil nutrient compositions, environmental conditions, and recommended crops, making it a valuable resource for researchers, data scientists, and agricultural technologists. It is designed for building machine learning models that predict the best crop to grow given specific soil and climate parameters.
📊 Purpose of the Dataset
The dataset aims to support advancements in:
Smart agriculture
Crop recommendation systems
Precision farming
Soil fertility analysis
Sustainable farming practices
It is ideal for classification tasks, EDA, pattern recognition, and ML-based crop advisory solutions.
🧪 Features Included
Typical features include (adjust based on your actual columns):
Feature Description N Nitrogen content in soil (mg/kg) P Phosphorus content in soil (mg/kg) K Potassium content in soil (mg/kg) pH Acidity/alkalinity level of the soil Temperature Ambient temperature (°C) Humidity Relative humidity (%) Rainfall Annual rainfall (mm) Crop Recommended crop for given parameters 🎯 Target Variable
Crop – The crop label recommended based on the soil and climatic conditions. This includes a variety of major Indian crops such as rice, maize, wheat, cotton, millet, sugarcane, cereals, and several others.
💡 Potential Use Cases
Building a crop recommendation ML model
Soil parameter analysis
Feature correlation studies
Climate–soil–crop relationship modeling
Sustainable farming decision support
📂 Format
File type: CSV
Each row represents a unique soil sample with its recommended crop.
🧭 Why This Dataset is Useful
With agriculture becoming increasingly data-driven, this dataset provides a practical foundation for developing intelligent systems that help farmers:
Improve yield
Reduce soil damage
Optimize fertilizer usage
Select the right crop for higher profitability
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
The Southern Great Plains 1997 (SGP97) Hydrology Experiment originated from an interdisciplinary investigation, "Soil Moisture Mapping at Satellite Temporal and Spatial Scales" (PI: Thomas J. Jackson, USDA Agricultural Research Service, Beltsville, MD) selected under the NASA Research Announcement 95-MTPE-03. The core of the 1997 experiment involves the deployment of the L-band Electronically Scanned Thinned Array Radiometer (ESTAR) for daily mapping of surface soil moisture. The region selected for investigation is the best instrumented site for surface soil moisture, hydrology and meteorology in the world. This includes the USDA/ARS Little Washita Watershed, the USDA/ARS facility at El Reno, Oklahoma, the ARM/CART central facility, as well as the Oklahoma Mesonet. The temporal coverage for this dataset is as follows: Begin datetime: 1995-10-01 00:00:00, End datetime: 2001-03-31 23:59:59. The Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) Southern Great Plains (SGP) Soil Texture Data Set is one of the various sub-surface data sets developed for the ARM/GCIP (Global Energy and Water Cycle Experiment (GEWEX) Continental-scale International Project) 1996 Near-Surface Observation (NESOB-96) Data Set. This data set contains a summary table of the percentages of sand, silt, and clay fractions in each soil layer at each of the ARM SWATS (Soil Water and Temperature System) sites at the SGP site. Also included is the corresponding USDA texture class as determined from the "soil triangle". The soil characterizations were perfomed by Oklahoma State University. Resources in this dataset:Resource Title: GeoData catalog record. File Name: Web Page, url: https://geodata.nal.usda.gov/geonetwork/srv/eng/catalog.search#/metadata/SGP97armTexture_JJM_2015-04-23_1409
Facebook
TwitterAttribution 3.0 (CC BY 3.0)https://creativecommons.org/licenses/by/3.0/
License information was derived automatically
The Soil and Terrain database for China primary data (version 1.0), at scale 1:1 million (SOTER_China), was compiled of enhanced soil information within the framework of the FAO's program of Land Degradation Assessment in Drylands (LADA). The primary database was compiled using the SOTER methodology. The SOTER unit delineation was based on a raster format of the soil map of China, correlated and converted to FAO’s Revised Legend (1988), in combination with a SOTER landform characterization derived from Shuttle Radar Topographic Mission (SRTM) 90 m digital elevation model (DEM). Reference profiles for the dominant soil of the SOTER units has been directly linked to the polygons. SOTER forms a part of the ongoing activities of ISRIC, FAO and UNEP to update the world's baseline information on natural resources.The project involved collaboration with national soil institutes from the countries in the region as well as individual experts.
Facebook
TwitterAttribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically
Overview
Precision Liming Soil Datasets (LimeSoDa) is a collection of 31 datasets from a field- and farm-scale soil mapping context. These datasets are "ready-to-use" for modeling purposes, as they include target soil properties and features in a tidy tabular format. Three target soil properties are present in every dataset: (1) soil organic matter (SOM) or soil organic carbon (SOC), (2) pH, and (3) clay content, while the features for modeling are dataset-specific. The primary goal of `LimeSoDa` is to enable more reliable benchmarking of machine learning methods in digital soil mapping and pedometrics. All the associated materials and data from LimeSoDa can be downloaded in this data repository. However, for a more in-depth analysis, we refer to the published paper "LimeSoDa: A Dataset Collection for Benchmarking of Machine Learning Regressors in Digital Soil Mapping" by Schmidinger et al. (2025). You may also use our R and Python package likewise called LimeSoDa.
Citation
Upon usage of datasets from LimeSoDa, please cite our associated paper:
Schmidinger, J., Vogel, S., Barkov, V., Pham, A.-D., Gebbers, R., Tavakoli, H., Correa, J., Tavares, T.R., Filippi, P., Jones, E. J., Lukas, V., Boenecke, E., Ruehlmann, J., Schroeter, I., Kramer, E., Paetzold, S., Kodaira, M., Wadoux, A.M.J.-C., Bragazza, L., Metzger, K., Huang, J., Valente, D.S.M., Safanelli, J.L., Bottega, E.L., Dalmolin, R.S.D., Farkas, C., Steiger, A., Horst, T. Z., Ramirez-Lopez, L., Scholten, T., Stumpf, F., Rosso, P., Costa, M.M., Zandonadi, R.S., Wetterlind, J. & Atzmueller, M. (2025). LimeSoDa: A Dataset Collection for Benchmarking of Machine Learning Regressors in Digital Soil Mapping. https://doi.org/10.48550/arXiv.2502.20139
Facebook
TwitterThis data set consists of soil texture classification data derived from field surveys as part of the Soil Moisture Active Passive Validation Experiment 2012 (SMAPVEX12). The soil texture classification map provides information about vegetation present in the study area.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
The National Soil Database has produced a national database of soil geochemistry including point and spatial distribution maps of major nutrients, major elements, essential trace elements, trace elements of special interest and minor elements. In addition, this study has generated a National Soil Archive, comprising bulk soil samples and a nucleic acids archive each of which represent a valuable resource for future soils research in Ireland. The geographical coherence of the geochemical results was considered to be predominantly underpinned by underlying parent material and glacial geology. Other factors such as soil type, land use, anthropogenic effects and climatic effects were also evident. The coherence between elements, as displayed by multivariate analyses, was evident in this study. Examples included strong relationships between Co, Fe, As, Mn and Cu. This study applied large-scale microbiological analysis of soils for the first time in Ireland and in doing so also investigated microbial community structure in a range of soil types in order to determine the relationship between soil microbiology and chemistry. The results of the microbiological analyses were consistent with geochemical analyses and demonstrated that bacterial community populations appeared to be predominantly determined by soil parent material and soil type.
Facebook
TwitterThis dataset was created by JAYAPRAKASHPONDY