Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
Explore the TCGA Whole Slide Image (WSI) SVS files available on Kaggle, offering detailed visual representations of tissue samples from various cancer types. These high-resolution images provide valuable insights into tumor morphology and tissue architecture, facilitating cancer diagnosis, prognosis, and treatment research. Delve into the rich landscape of cancer biology, leveraging the wealth of information contained within these SVS files to drive innovative advancements in oncology. This is a dataset of WSI images downloaded from the TCGA portal.
Facebook
TwitterAttribution-NonCommercial-ShareAlike 4.0 (CC BY-NC-SA 4.0)https://creativecommons.org/licenses/by-nc-sa/4.0/
License information was derived automatically
Large set of whole-slide-images (WSI) of prostatectomy specimens with various grades of prostate cancer (PCa). More information can be found in the corresponding paper: https://doi.org/10.1038/s41598-018-37257-4
The WSIs in this dataset can be viewed using the open-source software ASAP or Open Slide.
Due to the large size of the complete dataset, the data has been split up in to multiple archives.
The data from the training set:
The data from the test set:
This study was financed by a grant from the Dutch Cancer Society (KWF), grant number KUN 2015-7970.
If you make use of this dataset please cite both the dataset itself and the corresponding paper: https://doi.org/10.1038/s41598-018-37257-4
Facebook
TwitterThe dataset consists of 99 H&E-stained whole slide skin images (WSI) - 49 abnormal and 50 normal cases. All significant abnormal findings identified are outlined and categorized into 13 types such as actinic keratosis, basal cell carcinoma and dermatofibroma. Other tissue components, such as epidermis, adnexal structures, as well as the surgical margin are delineated to create a complete histological map. In total, 16741 separate annotations have been made to segment the different tissue structures and link them to ontological information.
Facebook
Twitterhttps://dataintelo.com/privacy-and-policyhttps://dataintelo.com/privacy-and-policy
Whole Slide Image market valued at $1.29 billion in 2025, projected to reach $3.85 billion by 2034 at 12.8% CAGR. Covers hardware, software, services across telepathology and diagnostics.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Links to code and bioRxiv pre-print:
1. Multi-lens Neural Machine (MLNM) Code
2. An AI-assisted Tool For Efficient Prostate Cancer Diagnosis (bioRxiv Pre-print)
Digitized hematoxylin and eosin (H&E)-stained whole-slide-images (WSIs) of 40 prostatectomy and 59 core needle biopsy specimens were collected from 99 prostate cancer patients at Tan Tock Seng Hospital, Singapore. There were 99 WSIs in total such that each specimen had one WSI. H&E-stained slides were scanned at 40× magnification (specimen-level pixel size 0·25μm × 0·25μm) using Aperio AT2 Slide Scanner (Leica Biosystems). Institutional board review from the hospital were obtained for this study, and all the data were de-identified.
Prostate glandular structures in core needle biopsy slides were manually annotated and classified using the ASAP annotation tool (ASAP). A senior pathologist reviewed 10% of the annotations in each slide, ensuring that some reference annotations were provided to the researcher at different regions of the core. It is to be noted that partial glands appearing at the edges of the biopsy cores were not annotated.
Patches of size 512 × 512 pixels were cropped from whole slide images at resolutions 5×, 10×, 20×, and 40× with an annotated gland centered at each patch. This dataset contains these cropped images.
This dataset is used to train two AI models for Gland Segmentation (99 patients) and Gland Classification (46 patients). Tables 1 and 2 illustrate both gland segmentation and gland classification datasets. We have put the two corresponding sub-datasets as two zip files as follows:
Table 1: The number of slides and patches in training, validation, and test sets for gland segmentation task. There is one H&E stained WSI for each prostatectomy or core needle biopsy specimen.
|
|
#Slides |
|
|
|
|
|
Train |
Valid |
Test |
Total |
|
Prostatectomy |
17 |
8 |
15 |
40 |
|
Biopsy |
26 |
13 |
20 |
59 |
|
Total |
43 |
21 |
35 |
99 |
|
|
#Patches |
|
|
|
|
|
Train |
Valid |
Test |
Total |
|
Prostatectomy |
7795 |
3753 |
7224 |
18772 |
|
Biopsy |
5559 |
4028 |
5981 |
15568 |
|
Total |
13354 |
7781 |
13205 |
34340 |
Table 2: The number of slides and patches in training, validation, and test sets for gland classification task. There is one H&E stained WSI for each prostatectomy or core needle biopsy specimen. The gland classification datasets are the subsets of the gland segmentation datasets. GS: Gleason Score. B: Benign. M: Malignant.
|
|
#Slides (GS 3+3:3+4:4+3) |
|
|
|
|
|
Train |
Valid |
Test |
Total |
|
Biopsy |
10:9:1 |
3:7:0 |
6:10:0 |
19:26:1 |
|
|
#Patches (B:M) |
|
|
|
|
|
Train |
Valid |
Test |
Total |
|
Biopsy |
1557:2277 |
1216:1341 |
1543:2718 |
4316:6336 |
NB: Gland classification folder (gland_classification_dataset.zip) may contain extra patches, labels of which could not be identified from H&E slides. They were not used in the machine learning study.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This dataset contains mouse colon H&E and immunohistochemistry (IHC) whole-slide images in MRXS format. The dataset includes 24 H&E slides and 72 IHC slides for ZO-1, Claudin-1, and MUC2. Slide-level metadata are provided in metadata_public.csv.
Facebook
TwitterThe ACROBAT data set consists of 4,212 whole slide images (WSIs) from 1,153 female primary breast cancer patients. The WSIs in the data set are available at 10X magnification and show tissue sections from breast cancer resection specimens stained with hematoxylin and eosin (H&E) or immunohistochemistry (IHC). For each patient, one WSI of H&E stained tissue and at least one one, and up to four, WSIs of corresponding tissue stained with the routine diagnostic stains ER, PGR, HER2 and KI67 are available. The data set was acquired as part of the CHIME study (chimestudy.se) and its primary purpose was to facilitate the ACROBAT WSI registration challenge (acrobat.grand-challenge.org). The histopathology slides originate from routine diagnostic pathology workflows and were digitised for research purposes at Karolinska Institutet (Stockholm, Sweden). The image acquisition process resembles the routine digital pathology image digitisation workflow, using three different Hamamatsu WSI scanners, specifically one NanoZoomer S360 and two NanoZoomer XR. The WSIs in this data set are accompanied by a data table with one row for each WSI, specifying an anonymised patient ID, the stain or IHC antibody type of each WSI, as well as the magnification and microns per pixel at each available resolution level. Automated registration algorithm performance evaluation is possible through the ACROBAT challenge website based on over 37,000 landmark pair annotations from 13 annotators. While the primary purpose of this data set was the development and evaluation of WSI registration methods, this data set has the potential to facilitate further research in the context of computational pathology, for example in the areas of stain-guided learning, virtual staining, unsupervised learning and stain-independent models.
The data set consists of three subsets, the training, validation and test set, based on the ACROBAT WSI registration challenge. There are 750 cases in the training set, for each of which one H&E WSI and one to four IHC WSIs are available, with 3406 WSIs in total. The validation set consists of 100 cases with 200 WSIs in total and the test set of 303 cases with 606 WSIs in total. Both for the validation and test set, one H&E WSI as well as one randomly selected IHC WSI is available.
WSIs were anonymised by deleting the associated macro images, by generating filenames with random case IDs and by overwriting meta data fields with potentially personal information. Hamamatsu NDPI files were then converted using libvips (libvips.org/). WSIs are available as generic tiled TIFF WSIs (openslide.org/formats/generic-tiff/) at 10X magnification and lower image levels.
The data set is available for download in seven separate ZIP archives, five for the training data (train_part1.zip (71.47 GB), train_part2.zip (70.59 GB), train_part3.zip (75.91 GB), train_part4.zip (71.63 GB) and train_part5.zip (69.09 GB)), one for the validation data (valid.zip 21.79 GB) and one for the test data (test.zip 68.11 GB).
File listings and checksums in SHA1 format are available for checking archive/data integrity when downloading.
While it would be helpful to notify SND of any publications using this data set by sending an email to request@snd.gu.se, please note that this is not required to use the data.
Facebook
TwitterAn open-access histopathology dataset of 1,399 H&E-stained whole-slide images of sentinel lymph nodes from breast cancer patients, created for automated detection and classification of metastases. The dataset includes training and test sets with pixel-level annotations.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Links to code:
Tissue Region Segmentation Code
This dataset comprises high-quality immunohistochemistry (IHC) and Haematoxylin and Eosin (H&E) whole slide images (WSIs) of breast tissues, provided in .svs format.
The data were collected from Bahçeşehir University Medical School and are intended for research in histopathology and computational pathology.
This study was approved by the Bahçeşehir University Clinical Research Institutional Review Board (Approval No: 2022-10/03).
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
Precise detection of invasive cancer on whole-slide images (WSI) is a critical first step in digital pathology tasks of diagnosis and grading. Convolutional neural network (CNN) is the most popular representation learning method for computer vision tasks, which have been successfully applied in digital pathology, including tumor and mitosis detection. However, CNNs are typically only tenable with relatively small image sizes (200x200 pixels). Only recently, Fully convolutional networks (FCN) are able to deal with larger image sizes (500x500 pixels) for semantic segmentation. Hence, the direct application of CNNs to WSI is not computationally feasible because for a WSI, a CNN would require billions or trillions of parameters. To alleviate this issue, this paper presents a novel method, High-throughput Adaptive Sampling for whole-slide Histopathology Image analysis (HASHI), which involves: i) a new efficient adaptive sampling method based on probability gradient and quasi-Monte Carlo sampling, and, ii) a powerful representation learning classifier based on CNNs. We applied HASHI to automated detection of invasive breast cancer on WSI. HASHI was trained and validated using three different data cohorts involving near 500 cases and then independently tested on 195 studies from The Cancer Genome Atlas. The results show that (1) the adaptive sampling method is an effective strategy to deal with WSI without compromising prediction accuracy by obtaining comparative results of a dense sampling (~6 million of samples in 24 hours) with far fewer samples (~2,000 samples in 1 minute), and (2) on an independent test dataset, HASHI is effective and robust to data from multiple sites, scanners, and platforms, achieving an average Dice coefficient of 76%.
Convolutional Neural Network - CS256-FC256 Convolutional Neural Network (CNN) trained for patch-based classification of invasive breast cancer from histopathology digital images. The CNN architecture is 256 units in convolution and pooling layers, 256 units of fully connected layer and 2 units for output classification layer of softmax. The model was trained with Torch7. model_epoch25.net XML annotations by HG of WSIs from CINJ XML region-based annotations by HG pathologist of whole-slide images (WSIs) from CINJ institution data cohort. XML_CINJ_HG.zip XML annotations by MF and NS of WSIs from CINJ XML region-based annotations by MF and NS pathologists of whole-slide images (WSIs) from CINJ institution data cohort. XML_CINJ_MF+NS.zip XML annotations by HG of WSIs from TCGA XML region-based annotations by HG pathologist of whole-slide images (WSIs) from TCGA institution data cohort subset used in the paper. XML_TCGA_HG.zip TCGA scaled images 195 TCGA scaled images used as D_test to test HASHI method. TCGA_imgs_idx5.zip UHCMC/CWRU scaled images 110 UHCMC/CWRU scaled images used as D_2 dataset as part of Whole-Slide Image training data set. CWRU_imgs_idx8.zip CINJ scaled images 40 CINJ scaled images used as D_4 dataset as part of Whole-slide Image validation data set. CINJ_imgs_idx5.zip HUP scaled images Part1 120 of 239 HUP scaled images used as D_1 dataset as part of Whole-Slide Image training data set. HUP_imgs_idx5_Part1.zip HUP scaled images Part2 119 of 239 HUP scaled images used as D_1 dataset as part of Whole-Slide Image training data set. HUP_imgs_idx5_Part2.zip HUP binary masks of annotations HUP binary masks of manual annotations from pathologists. HUP_masks.zip UHCMC/CWRU binary masks of annotations UHCMC/CWRU binary masks of manual annotations from pathologists. CWRU_masks.zip CINJ binary masks of annotations CINJ binary masks of manual annotations from pathologists. CINJ_masks_HG.zip TCGA binary masks of annotations TCGA binary masks of manual annotations from pathologists. TCGA_masks.zip
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
This dataset was created by Lakeshprabhu Thangadurai
Released under Apache 2.0
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
A large annotated dataset, composed of both microscopy (classification task) and whole-slide images (segmentation task), was specifically compiled and made publicly available for the BACH challenge. Following a positive response from the scientific community, a total of 64 submissions, out of 677 registrations, effectively entered the competition. From the submitted algorithms it was possible to push forward the state-of-the-art in terms of accuracy (87%) in automatic classification of breast cancer with histopathological images.
There are two main folders for classification task: train and test. In Photos folder, there are totally four classes: benign, in situ, invasive, and normal. There is also a ground truth csv file for labels. Images are tif format.
Paper: https://arxiv.org/abs/1808.04277
Citation: Aresta, G., Araújo, T., Kwok, S., Chennamsetty, S. S., Safwan, M., Alex, V., ... & Aguiar, P. (2019). Bach: Grand challenge on breast cancer histology images. Medical image analysis, 56, 122-139.
Dataset: https://zenodo.org/record/3632035
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This is a representative sample from the dataset that was used to develop resolution-agnostic convolutional neural networks for tissue segmentation1 in whole-slide histopathology images.
The dataset is composed of two parts: development set and dissimilar set.
Sample images from the development set:
Sample images from the dissimilar set:
Facebook
TwitterThe data set consists of 81 registered whole slide image pairs, a pair represents unstained and H&E stained images of the same tissue sample. In addition to that, it also contains a tissue mask for each whole slide image pair. The samples are used for studying the histological feasibility of AI-driven virtual histopathology staining.
Imaging was performed using Thunder Imager 3D Tissue slide scanner (Leica Microsystems, Wetzlar, Germany) equipped with DMC2900 camera and HC PL APO 40x/0.95 DRY objective with an isotropic pixel resolution of 0.353 µm.
Facebook
Twitterhttps://researchintelo.com/privacy-and-policyhttps://researchintelo.com/privacy-and-policy
According to our latest research, the Global Whole Slide Image Management Systems market size was valued at $1.2 billion in 2024 and is projected to reach $4.7 billion by 2033, expanding at a robust CAGR of 16.5% during the forecast period of 2025–2033. The primary catalyst for this impressive growth trajectory is the rapid adoption of digital pathology solutions, particularly in clinical diagnostics and research, which has significantly increased the demand for efficient whole slide image management systems globally. As healthcare providers and research organizations transition from traditional glass slides to digital formats, the necessity for robust, scalable, and secure image management platforms has become paramount, driving investments and innovation across the sector.
North America currently commands the largest share of the global whole slide image management systems market, accounting for approximately 42% of the total market value in 2024. This dominance is attributed to the region’s mature healthcare infrastructure, early adoption of digital pathology, and supportive regulatory frameworks that encourage technological integration in clinical workflows. The presence of leading industry players, coupled with high investment in healthcare IT and frequent technological upgrades, has solidified North America’s leadership. Furthermore, the region benefits from a strong network of academic and research institutions that continuously drive demand for advanced imaging solutions, further supporting market expansion.
Asia Pacific is poised to be the fastest-growing region, with a projected CAGR of over 19% from 2025 to 2033. The surge in market growth is primarily driven by increasing healthcare expenditure, rapid digital transformation initiatives, and growing awareness of the benefits of digital pathology in countries such as China, India, and Japan. Governments in the region are actively investing in upgrading healthcare infrastructure and promoting the adoption of advanced diagnostic technologies, which is further bolstered by an expanding base of skilled medical professionals and researchers. The influx of foreign direct investment and the establishment of regional manufacturing hubs are also contributing to the accelerated adoption of whole slide image management systems in Asia Pacific.
Emerging economies in Latin America and the Middle East & Africa are witnessing gradual adoption of whole slide image management systems, albeit at a slower pace compared to developed regions. Challenges such as limited access to high-speed internet, budgetary constraints, and a lack of standardized digital pathology protocols have somewhat impeded market penetration. However, localized demand is steadily increasing, particularly in urban centers where healthcare modernization is a priority. Policy reforms aimed at improving healthcare delivery and increasing investments in digital health infrastructure are expected to drive future growth, although overcoming infrastructural and training barriers remains essential for widespread adoption.
| Attributes | Details |
| Report Title | Whole Slide Image Management Systems Market Research Report 2033 |
| By Component | Software, Hardware, Services |
| By Deployment Mode | On-Premises, Cloud-Based |
| By Application | Pathology, Education, Research, Telemedicine, Others |
| By End-User | Hospitals, Diagnostic Laboratories, Academic & Research Institutes, Pharmaceutical & Biotechnology Companies, Others |
| Regions Covered | North America, Europe, Asia Pacific, Latin America and Middle East & Africa |
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
The dataset is comprised of 38 chemically stained Whole slide image samples along with their corresponding ground truth annotated by histopathologists for 12 classes indicating skin layers (Epidermis, Reticular dermis, Papillary dermis, Dermis, Keratin), Skin tissues (Inflammation, Hair follicles, Glands), skin cancer (Basal cell carcinoma, Squamous cell carcinoma, Intraepidermal carcinoma) and background (BKG).
Facebook
TwitterThis is a whole slide image (WSI) dataset for glomeruli segmentation on kidney tissue, in total 88 images. The train-set (58 images) and test-set (32 images) has been used in the Orbit publication (1) to train and test the glomeruli segmentation model (2).
The images are pyramidal tiff images (tiled, jpeg-compression) and can be displayed with Orbit Image Analysis (3).
The file orbit.db is a sqllite database which contains the manual drawn glomeruli annotations for all images, in total 21037 annotations. It can be placed in the user-home folder, then Orbit Image Analysis (3) will detect the database and show the glomeruli annotations in the annotation tab when opening an image. (Orbit will use the md5 hashes of the images for identification.)
For more information on how to train a CNN model or to use the existing model (2) please visit the Orbit deep learning page (4).
(1) Manuel Stritt, Anna K. Stalder, Enrico Vezzali; Orbit Image Analysis: An open-source whole slide image analysis...
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
Digitized H&E-stained whole-slide image from TCGA case TCGA-CC-A1HT (LIHC). Includes 105 representative tiles across 21 histomic features with AI cell segmentation overlays.
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
This dataset contains key characteristics about the data described in the Data Descriptor A completely annotated whole slide image dataset of canine breast cancer to aid human breast cancer research. Contents:
1. human readable metadata summary table in CSV format
2. machine readable metadata file in JSON format
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
Anonymized whole slide image of canine mammary carcinoma, stained with H&E. File is in Aperio SVS format.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
Explore the TCGA Whole Slide Image (WSI) SVS files available on Kaggle, offering detailed visual representations of tissue samples from various cancer types. These high-resolution images provide valuable insights into tumor morphology and tissue architecture, facilitating cancer diagnosis, prognosis, and treatment research. Delve into the rich landscape of cancer biology, leveraging the wealth of information contained within these SVS files to drive innovative advancements in oncology. This is a dataset of WSI images downloaded from the TCGA portal.