Facebook
TwitterAttribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically
This chart shows the 2-Year Impact of Medical Reference Services Quarterly over time and its percentile among journals.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
The dataset contains multi-modal data from over 75,000 open access and de-identified case reports, including metadata, clinical cases, image captions and more than 130,000 images. Images and clinical cases belong to different medical specialties, such as oncology, cardiology, surgery and pathology. The structure of the dataset allows to easily map images with their corresponding article metadata, clinical case, captions and image labels. Details of the data structure can be found in the file data_dictionary.csv.
Almost 100,000 patients and almost 400,000 medical doctors and researchers were involved in the creation of the articles included in this dataset. The citation data of each article can be found in the metadata.parquet file.
Refer to the examples showcased in this GitHub repository to understand how to optimize the use of this dataset.
For a detailed insight about the contents of this dataset, please refer to this data article published in Data In Brief.
Facebook
TwitterThis dataset describes the Release File structure of SNOMED CT UK Drug Extension, referred to as Release Format 2 (RF2). The UK Edition of SNOMED CT is the official source of SNOMED CT for use in UK healthcare systems. The UK Edition is a standalone release that combines the content of both the US Extension and the International release of SNOMED CT
A Simple reference set does not have any addition fields.
Facebook
Twitterhttps://fred.stlouisfed.org/legal/#copyright-public-domainhttps://fred.stlouisfed.org/legal/#copyright-public-domain
Graph and download economic data for Hourly Compensation for Health Care and Social Assistance: Medical and Diagnostic Laboratories (NAICS 6215) in the United States (IPURN6215U121000000) from 1995 to 2025 about diagnostic labs, healthcare, medical, social assistance, compensation, health, NAICS, hours, IP, and USA.
Facebook
Twitterhttps://fred.stlouisfed.org/legal/#copyright-public-domainhttps://fred.stlouisfed.org/legal/#copyright-public-domain
Graph and download economic data for Real Pharmaceutical and Other Medical Products Expenditures (PHMEPRREXHCSA) from 2000 to 2021 about pharmaceuticals, healthcare, medical, health, expenditures, real, and USA.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This dataset contains 1,000 structured, context-specific recommendation records derived from 76 English-language Clinical Practice Guidelines (CPGs) published by the Ministry of Health Malaysia, covering 20 clinical specialty categories. It was developed as a CPG-grounded reference dataset for small language model research, instruction tuning, retrieval-augmented generation, clinical guideline retrieval, and recommendation-generation evaluation. The file extraction_prompt.txt contains the complete prompt template used for candidate inclusion decisions, field mapping and structured draft generation.
The dataset is provided in JSON Lines (JSONL) format, with one independent record per line. Each record contains three fields:
• “instruction”: A standard request for a CPG recommendation. • “input”: A disease or clinical condition and its specific clinical context. • “output”: A concise CPG-grounded recommendation for that context.
Example:
{"instruction": "Provide standard CPG recommendations.", "input": "Disease: retinopathy of prematurity Clinical context: screening examination", "output": "CPG Recommendation: Use binocular indirect ophthalmoscopy or an approved imaging pathway by trained personnel and document zone, stage, extent, and plus disease."}
The records address clinical contexts such as screening, diagnosis, risk assessment, treatment, medication, monitoring, referral, prevention, follow-up, warning signs, and intervention cessation. A disease may appear in multiple records representing different clinical contexts.
The source CPGs were obtained from official Malaysian Ministry of Health websites. Dataset preparation involved identifying recommendation-bearing content, assigning the relevant disease and clinical context, converting the content into a consistent instruction–input–output structure, and checking JSONL validity, duplication, relevance, and source consistency.
The dataset may be used as an experimental ground-truth reference for model evaluation but has not been independently validated as clinical ground truth by qualified clinical experts. It must not replace the original CPGs, professional medical judgement, diagnosis, or patient-specific treatment decisions.
Copyright in the original CPG documents remains with the Malaysian Ministry of Health or the respective rights holders.
Facebook
TwitterReference dataset on what makes an expired or aged domain in the health and medical niche valuable and what its history must prove, presented as general market education about domain SEO and not as medical advice or a valuation of any specific domain. Health and medical content is the first-listed Your Money or Your Life (YMYL) category because inaccurate medical information can cause direct physical harm, so its trust bar runs higher than any niche, higher even than finance: trust is the heaviest weighted component of the quality framework, credentialed and licensed-practitioner authorship plus a medical-review step are expected, and anonymous or generic authorship is filtered out of medical citations. The health trust bar has tightened across a multi-year line: the August 2018 core update was nicknamed the Medic update after more than 41 percent of impacted sites were reported as healthcare-related, and the December 2025 core update (rolled out across 18 days from 11 to 29 December 2025) saw around 67 percent of YMYL sites register ranking declines with health and medical just behind finance among the hardest-hit verticals; the YMYL recovery window runs 6 to 12 months versus 2 to 6 for non-YMYL content, which compresses the rebuild margin an aged health domain provides. A health domain qualifies on three history signals, not its name: genuine prior health content verified through archived snapshots, an on-topic clean backlink profile concentrated in health sources, and clean index standing held until expiry (a name deindexed before expiry is a warning sign while one indexed until expiry indicates relatively clean standing). The prior-use category shapes risk (a former clinic, hospital, or charity carries the strongest continuity and the heaviest impersonation and trademark screen) while topical history plus a clean profile decide transfer. The health sub-verticals shift the topical-continuity equation: clinical and medical-device content sits at the strictest trust bar with licensed-practitioner authorship and medical review and the least drift tolerance; mental health runs close behind as a high-harm sub-vertical (a documented case rebuilt a former mental-health software domain, Domain Rating 38, 659 backlinks from 153 referring domains including bbc.co.uk, techcrunch.com, and crunchbase.com, into the adjacent mental-health and nootropic-supplement space and reached roughly $3,226 in its fourth month by keeping the rebuild in-category); supplements pair the YMYL bar with affiliate-disclosure scrutiny and strong economics (payouts in the triple digits around $140 and above); telehealth and medical-SaaS overlap the software and SaaS backlink ecosystem so continuity is comparatively easy; wellness and fitness span the broadest content base at a lower bar. Mismatch and contamination are most dangerous in health because the medical YMYL trust bar magnifies both the relevance discount on an off-topic history and the risk of a tainted profile (expired domain abuse, repurposing primarily to manipulate, is tested by thematic coherence rather than age or link count; a complete topic change destroys most inherited value while the headline metric barely moves; documented patterns include commercial medical products on a former non-profit medical charity and a lapsed children's cancer-charity domain repurposed into casino content; a documented vetting rule rejects a medical blog later turned into a gambling site). Speed: a fresh domain needs 12 or more months for commercial visibility in a competitive niche (meaningful gains in 3 to 6 months, YMYL recovery and build 6 to 12 months), and niche-focused sites reach Domain Authority 40 about 30 percent faster than generalist sites (roughly 18 versus 26 months); health is among the most competitive and most scrutinized verticals, so the aligned aged-domain head start is most valuable here, and the referring-domain topical mix decides whether the inheritance is real or hollow. Diligence workflow: verify the health topic via archives, audit the backlink topical mix for an on-topic clean share, confirm clean index standing, against a health-calibrated bar with an added impersonation and trademark screen for medical or charity prior-use. Selection versus guarantee: a niche-aware curated catalogue sorts aged domains by genuine health history, screens each health name for an on-topic clean profile, and verifies index standing at intake, so a buyer sources health-appropriate pre-screened inventory; it removes the history guesswork before a listing but does not author the credentialed, reviewed content the trust bar demands and does not and cannot promise a ranking outcome, and health aged-domain projects still carry execution risk.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
By Huggingface Hub [source]
The MedQuad dataset provides a comprehensive source of medical questions and answers for natural language processing. With over 43,000 patient inquiries from real-life situations categorized into 31 distinct types of questions, the dataset offers an invaluable opportunity to research correlations between treatments, chronic diseases, medical protocols and more. Answers provided in this database come not only from doctors but also other healthcare professionals such as nurses and pharmacists, providing a more complete array of responses to help researchers unlock deeper insights within the realm of healthcare. This incredible trove of knowledge is just waiting to be mined - so grab your data mining equipment and get exploring!
For more datasets, click here.
- 🚨 Your notebook can be here! 🚨!
In order to make the most out of this dataset, start by having a look at the column names and understanding what information they offer: qtype (the type of medical question), Question (the question in itself), and Answer (the expert response). The qtype column will help you categorize the dataset according to your desired question topics. Once you have filtered down your criteria as much as possible using qtype, it is time to analyze the data. Start by asking yourself questions such as “What treatments do most patients search for?” or “Are there any correlations between chronic conditions and protocols?” Then use simple queries such as SELECT Answer FROM MedQuad WHERE qtype='Treatment' AND Question LIKE '%pain%' to get closer to answering those questions.
Once you have obtained new insights about healthcare based on the answers provided in this dynmaic data set - now it’s time for action! Use all that newfound understanding about patient needs in order develop educational materials and implement any suggested changes necessary. If more criteria are needed for querying this data set see if MedQuad offers additional columns; sometimes extra columns may be added periodically that could further enhance analysis capabilities; look out for notifications if these happen.
Finally once making an impact with the use case(s) - don't forget proper citation etiquette; give credit where credit is due!
- Developing medical diagnostic tools that use natural language processing (NLP) to better identify and diagnose health conditions in patients.
- Creating predictive models to anticipate treatment options for different medical conditions using machine learning techniques.
- Leveraging the dataset to build chatbots and virtual assistants that are able to answer a broad range of questions about healthcare with expert-level accuracy
If you use this dataset in your research, please credit the original authors. Data Source
License: CC0 1.0 Universal (CC0 1.0) - Public Domain Dedication No Copyright - You can copy, modify, distribute and perform the work, even for commercial purposes, all without asking permission. See Other Information.
File: train.csv | Column name | Description | |:--------------|:------------------------------------------------------| | qtype | The type of medical question. (String) | | Question | The medical question posed by the patient. (String) | | Answer | The expert response to the medical question. (String) |
If you use this dataset in your research, please credit the original authors. If you use this dataset in your research, please credit Huggingface Hub.
Facebook
Twitterhttps://fred.stlouisfed.org/legal/#copyright-public-domainhttps://fred.stlouisfed.org/legal/#copyright-public-domain
Graph and download economic data for Producer Price Index by Commodity: Health Care Services: Medical Laboratory and Diagnostic Imaging Center Care (WPU511102) from Mar 2009 to Jul 2026 about diagnostic labs, medical, healthcare, health, services, commodities, PPI, inflation, price index, indexes, price, and USA.
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
This dataset consists of electronic medical records collected from healthcare information systems and publicly available clinical data repositories. It includes structured clinical attributes related to patient demographics, medical conditions, hospital encounters, and treatment-related information. The dataset is designed to support research on medical data organization, secure access, and relevance-based information retrieval within healthcare environments.
All records are anonymized to protect patient confidentiality, ensuring compliance with ethical data usage standards. The dataset enables exploration of semantic relevance in medical data access, efficient record retrieval, and secure handling of sensitive healthcare information. It is suitable for developing and evaluating healthcare data search systems that emphasize accurate information access while maintaining data privacy and access accountability.
Patient Demographics: gender: Male, Female, Unknown age ethnicity
Hospital Details: hospitalid: Each hospital was given unique id wardid: Ward Id is given in which patient was treated apacheadmissiondx: Disease diagnosed admissionheight: Height of the patients hospitaladmittime24: Admission time to the hospital hospitaladmitsource: Department Source of the admission hospitaldischargeyear: Discharge year from the hospital hospitaldischargetime24: Discharge time from the hospital hospitaldischargelocation: Patient Discharge to which location (Home, Death, Other hospital. etc) hospitaldischargestatus (Alive, Expired)
Hospital Unit Details: unittype: Unit in which admitted unitadmittime24: Time of admision to the Unit unitadmitsource: Department source for the unit unitvisitnumber: No. of times visited unitstaytype: Admit, readmit, etc admissionweight: Weight during the admission dischargeweight: Weight during the Discharge unitdischargetime24: Discharge time from the Unit unitdischargelocation: Patient Discharge to which location (Home, Death, Other hospital. etc) unitdischargestatus: (Alive, Expired)
Column Descriptions
patient_id – Unique identifier assigned to each patient record.
anonymized_patient_id – Privacy-preserving identifier used to protect patient identity.
age – Age of the patient at the time of record entry.
gender – Biological gender of the patient.
admission_type – Type of hospital admission associated with the record.
diagnosis – Medical condition or diagnosis recorded for the patient.
procedure – Clinical procedure or medical intervention documented.
medication – Medication or treatment information associated with the patient.
hospital_stay – Duration or details of the patient’s hospital visit.
clinical_text – Consolidated clinical information representing patient medical context.
search_query – Medical information request used for record retrieval evaluation.
permission_level – Authorized access role associated with the medical record.
access_hash – Cryptographic reference used to record secure access events.
relevance_score – Degree of relevance between a query and a medical record.
Facebook
TwitterA collection of online books and documents in life science and healthcare whose full text can be searched through the Entrez system. Bookshelf provides free online access to books and documents in life science and healthcare.
Facebook
Twitterhttps://www.grandviewresearch.com/horizon//info/terms-of-usehttps://www.grandviewresearch.com/horizon//info/terms-of-use
The medical reference apps market in the UAE is expected to reach a projected revenue of US$ 8.7 million by 2030. A compound annual growth rate of 16.2% is expected of the UAE medical reference apps market from 2025 to 2030.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
This dataset was created by Diện Trần
Released under Apache 2.0
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
This dataset provides detailed laboratory test results, including test types, quantitative and qualitative outcomes, reference ranges, ordering physician information, specimen details, and timestamps. It enables clinical analysis, patient monitoring, and quality assurance in healthcare settings, supporting interoperability and regulatory compliance.
Facebook
TwitterParse measures the brands and sources AI systems cite when answering "As a physician, what's the best medical reference app for my smartphone? UpToDate vs. Epocrates?".
Facebook
Twitterhttps://www.grandviewresearch.com/info/terms-of-usehttps://www.grandviewresearch.com/info/terms-of-use
Market size, estimate, forecast and CAGR for the Medical Reference Apps Market Size Report, 2025-2030.
Facebook
Twitterhttps://fred.stlouisfed.org/legal/#copyright-public-domainhttps://fred.stlouisfed.org/legal/#copyright-public-domain
Graph and download economic data for All Employees: Education and Health Services: General Medical and Surgical Hospitals in Maryland (DISCONTINUED) (SMU24000006562210001SA) from Jan 1990 to Dec 2023 about surgical, hospitals, medical, MD, health, employment, and USA.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
We broke down the giant, lumped Diabetes_Digestive_Kidney category into 7 clear, separate medical fields: - Diabetes - Digestive Diseases - Kidney Diseases - Urological Disease - Liver Disease - Endocrine Disease - Blood Disease
We repaired all dead, old, and broken URLs so your application can display working website citations: NIDDK Links: Replaced dead old .aspx paths with live, modern resource pages. Rare Diseases (GARD): Turned old database ID numbers into working text search links (with hyphens removed). Senior Health: Moved links from the retired nihseniorhealth.gov site directly to active deep-links on the modern National Institute on Aging (NIA) portal.
We cleaned the focus column by removing messy sentences, questions, and filler phrases (like “What I need to know about...” or “Overview of...”). These are now short, precise medical terms that work perfectly as filter buttons in your frontend user interface.
This dataset was scrapped from the MedQuAD repository and then converted it to a csv file.
IT DOES NOT INCLUDE DATA OF MPlus_ADAM_Encyclopedia, MPlus_Drugs, and MPlus_Herbs_Supplements as their answers were removed from the repository due to copyright strike.
topic: The broad medical category mapped from the authoritative source institute (e.g., cancer, Heart_Lung_Blood).
focus: The highly specific disease, syndrome, drug, or clinical condition targeted (e.g., Adult Acute Lymphoblastic Leukemia).
qtype: The precise clinical intent of the question (e.g., symptoms, treatment, exams and tests, stages, outlook).
question: The natural language medical query string.
answer: The authoritative, verified medical text passage answering the query.
url: The official NIH source web URL used for lineage tracking and user interface citations.
cancer: National Cancer Institute (CancerGov)
Genetic_and_Rare_Diseases: Genetic and Rare Diseases Information Center (GARD)
Genetics_Home_Reference: Genetics Home Reference (GHR - Core genetic conditions)
others: MedlinePlus General Health Topics
Diabetes_Digestive_Kidney: National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK)
Neurological_Disorders_Stroke: National Institute of Neurological Disorders and Stroke (NINDS)
seniorHealth: NIH Senior Health Website
Heart_Lung_Blood: National Heart, Lung, and Blood Institute (NHLBI)
Disease_Control_Prevention: Centers for Disease Control and Prevention (CDC)
MedQuAD includes 47,457 medical question-answer pairs created from 12 NIH websites (e.g. cancer.gov, niddk.nih.gov, GARD, MedlinePlus Health Topics). The collection covers 37 question types (e.g. Treatment, Diagnosis, Side Effects) associated with diseases, drugs and other medical entities such as tests.
We included additional annotations in the XML files, that could be used for diverse IR and NLP tasks, such as the question type, the question focus, its syonyms, its UMLS Concept Unique Identifier (CUI) and Semantic Type. We added the category of the question focus (Disease, Drug or Other) in the 4 MedlinePlus collections. All other collections are about diseases.
The paper cited below describes the collection, the construction method as well as its use and evaluation within a medical question answering system.
N.B. We removed the answers from 3 subsets to respect the MedlinePlus copyright (https://medlineplus.gov/copyright.html): (1) A.D.A.M. Medical Encyclopedia, (2) MedlinePlus Drug information, and (3) MedlinePlus Herbal medicine and supplement information. -- We kept all the other information including the URLs in case you want to crawl the answers. Please contact me if you have any questions.
If you use the MedQuAD dataset and/or the collection of 2,479 judged answers, please cite the following paper: "A Question-Entailment Approach to Question Answering". Asma Ben Abacha and Dina Demner-Fushman. BMC Bioinformatics, 2019.
@ARTICLE{BenAbacha-BMC-2019,
author = {Asma {Ben Abacha} and Dina Demner{-}Fushman}, title = {A Question-Entailment Approach to Question Answering}, journal = {{BMC} Bioinform.}, volume = {20}, number = {1}, pages = {511:1--511:23}, year = {2019}, url = {https://bmcbioinformatics.biomedcentral.com/articles/10.1186/s12859-019-3119-4} }
Facebook
TwitterAttribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically
This chart shows the 5-Year Impact of Medical Reference Services Quarterly over time and its percentile among journals.
Facebook
Twitterhttps://www.marketreportanalytics.com/privacy-policyhttps://www.marketreportanalytics.com/privacy-policy
The Clinical Reference Laboratory Services Market is booming, projected to reach [estimated 2033 market size in millions] by 2033, driven by technological advancements, rising chronic diseases, and an aging population. Explore market trends, key players, and regional insights in this comprehensive analysis.
Facebook
TwitterAttribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically
This chart shows the 2-Year Impact of Medical Reference Services Quarterly over time and its percentile among journals.