66 datasets found
  1. Z

    NA12878 WES Benchmark dataset

    • data.niaid.nih.gov
    Updated May 31, 2020
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Pranckeviciene Erinija (2020). NA12878 WES Benchmark dataset [Dataset]. https://data.niaid.nih.gov/resources?id=zenodo_3597726
    Explore at:
    Dataset updated
    May 31, 2020
    Dataset authored and provided by
    Pranckeviciene Erinija
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    This dataset makes available the UCSC Genome Browser (genome.ucsc.edu) GRCh37 genome build public session NA12878 WES Benchmark files in a single dataset so that these files can be used in other applications or genome browsers such as IGV. All genomic variant calls in all VCF files were decomposed and normalized with vt. This dataset contains: Genome in a bottle (GIAB) version 3.3.2 high confidence (HC) variant calls and genomic regions for HapMap individual NA12878 : GIAB_v3.3.2_NA12878-decomposed-normalized.vcf.gz GIAB_v3.3.2_NA12878-decomposed-normalized.vcf.gz.tbi GIAB_v3.3.2_NA12878_HC_regions.bed HapMap individual NA12878 WES variant calls (VCF) and capture regions (BED) from diagnostic laboratories : ARUP whole exome sequencing data (HiSeq 2000) publically available from NCBI GeT-RM Browser converted_ARUP_NA12878_Exome-decomposed-normalized.vcf.gz converted_ARUP_NA12878_Exome-decomposed-normalized.vcf.gz.tbi ARUP_SeqCap_EZ_Exome.bed UCSF whole exome sequencing data (HiSeq 2500) publically available from NCBI GeT-RM Browser converted_UCSF_NA12878_WES_Agilent_V4_Custom-decomposed-normalized.vcf.gz converted_UCSF_NA12878_WES_Agilent_V4_Custom-decomposed-normalized.vcf.gz.tbi UCSF_WES_Agilent_V4_Custom.bed Whole exome data (NextSeq 500) sequenced in CHEO diagnostic laboratory CHEO_NA12878_WES_S1dataset.vcf.gz CHEO_NA12878_WES_S1dataset.vcf.gz.tbi Agilent_CRE_v2.bed Genomic coordinates (BED) of OMIM genes for which a molecular basis of the associated disease is known (as of September 2019) : Omim_Genes.bed

  2. Concordance of genotypes represented in VCF and gVCF files with those...

    • figshare.com
    xls
    Updated Jun 1, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Alberto Ferrarini; Luciano Xumerle; Francesca Griggio; Marianna Garonzi; Chiara Cantaloni; Cesare Centomo; Sergio Marin Vargas; Patrick Descombes; Julien Marquis; Sebastiano Collino; Claudio Franceschi; Paolo Garagnani; Benjamin A. Salisbury; John Max Harvey; Massimo Delledonne (2023). Concordance of genotypes represented in VCF and gVCF files with those detected by the MI RISK Plus kit. [Dataset]. http://doi.org/10.1371/journal.pone.0132180.t001
    Explore at:
    xlsAvailable download formats
    Dataset updated
    Jun 1, 2023
    Dataset provided by
    PLOShttp://plos.org/
    Authors
    Alberto Ferrarini; Luciano Xumerle; Francesca Griggio; Marianna Garonzi; Chiara Cantaloni; Cesare Centomo; Sergio Marin Vargas; Patrick Descombes; Julien Marquis; Sebastiano Collino; Claudio Franceschi; Paolo Garagnani; Benjamin A. Salisbury; John Max Harvey; Massimo Delledonne
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    Concordance of genotypes represented in VCF and gVCF files with those detected by the MI RISK Plus kit.

  3. c

    Research data supporting "Aedes aegypti exome-sequencing and Aedes bromeliae...

    • repository.cam.ac.uk
    bin
    Updated Nov 17, 2016
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Crawford, Jacob; Alves, Joel M; Palmer, William J; Day, Jonathan P; Sylla, Massamba; Ramasamy, Ranjan; Surendran, Sinnathamby N; Black IV, William C.; Pain, Arnab; Jiggins, Francis M (2016). Research data supporting "Aedes aegypti exome-sequencing and Aedes bromeliae whole-genome sequencing" (VCF files containing variants) [Dataset]. http://doi.org/10.17863/CAM.6367
    Explore at:
    bin(1044912528 bytes), bin(139205435 bytes)Available download formats
    Dataset updated
    Nov 17, 2016
    Dataset provided by
    University of Cambridge
    Apollo
    Authors
    Crawford, Jacob; Alves, Joel M; Palmer, William J; Day, Jonathan P; Sylla, Massamba; Ramasamy, Ranjan; Surendran, Sinnathamby N; Black IV, William C.; Pain, Arnab; Jiggins, Francis M
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    This repository contains two VCF files that were generated from the called genotyped obtained from ANGSD. One VCF file contains variants obtained from whole-genome sequencing data from one individual sample of Aedes bromeliae (aedesBromeliae_1sample_Brom.mdp8.vcf.gz ) The other VCF file contains variants obtained from whole-exome sequencing data from 71 individual samples belonging to different populations in Africa and outside of Africa (aedesAegypti_71samples_AEGY.allpop.J18.AllSites.REFpol.CALLED.vcf_.gz). The data is accompanied by the raw sequencing data available in SRA with the accession number SRP092518 and with the BioProject number PRJNA349448 More information about the project is available in the manuscript that is under preparation.

  4. E

    Whole Exome Sequencing of healthy Spanish individuals - VCF file

    • ega-archive.org
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Whole Exome Sequencing of healthy Spanish individuals - VCF file [Dataset]. https://ega-archive.org/datasets/EGAD00001003101
    Explore at:
    License

    https://ega-archive.org/dacs/EGAC00001000222https://ega-archive.org/dacs/EGAC00001000222

    Description

    The need for a detailed catalogue of local variability for the study of rare diseases within the context of the Medical Genome Project motivated the whole exome sequencing of 267 unrelated individuals, representative of the healthy Spanish population.

  5. d

    Data from: Unraveling the genetics of feline hypertrophic cardiomyopathy: A...

    • search.dataone.org
    Updated Jun 27, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Michael Vandewege; Joanna Kaplan; Victor Rivas; Jalena Wouters; Samantha Harris; Kathryn Meurs; Joshua Stern (2025). Unraveling the genetics of feline hypertrophic cardiomyopathy: A multiomics study of 138 cats [Dataset]. http://doi.org/10.5061/dryad.cjsxksnjh
    Explore at:
    Dataset updated
    Jun 27, 2025
    Dataset provided by
    Dryad Digital Repository
    Authors
    Michael Vandewege; Joanna Kaplan; Victor Rivas; Jalena Wouters; Samantha Harris; Kathryn Meurs; Joshua Stern
    Description

    Hypertrophic cardiomyopathy (HCM) is the most common inherited cardiac disease in cats, often leading to congestive heart failure, arterial thromboembolism, and sudden cardiac death. The genetics of feline HCM are poorly understood, and limited genetic discoveries remain breed or family-specific. We aimed to identify novel causative or disease-modifying variants in a large cohort of cats reflective of the general cat population. In a second cohort, we sought to characterize transcriptomic differences between HCM-affected cats and healthy controls. DNA was isolated from 138 domestic cats (109 HCM and 29 controls). No single or combination of variants of high, moderate, or modifying impact were identified in genome-wide analysis to cause or modify the disease severity of HCM. Several rare high and moderate-impact variants in genes associated with human HCM were detected in diseased cats. In a second cohort, left ventricular (LV), interventricular septal (IVS), and left atrial (LA) tissues..., WGS data generation A total of 1-2 mL of whole blood were collected from the cephalic, saphenous, or jugular vein into EDTA blood collection tubes. DNA was either isolated from whole blood or from buffy coats after whole blood centrifugation at 2000 rpm for 15 minutes. Genomic DNA isolation was performed using commercially available kits (Gentra Puregene Blood kit, QIAGEN, Hilden Germany; ArchivePure;5Prime) and by following the respective manufacturer’s protocol. High-quality unfragmented DNA was selected by a combination of 1% agarose gel visualization and spectrophotometric confirmation (a 260/280 ratio of ~1.8 and a concentration of > 50 ng/uL; NanoDrop One/One, Thermofisher, Waltham, GA, USA). Samples were stored at -20°C until ready for shipment to Theragen Bio Co., Ltd, Gyeonggi-do, Republic of Korea for WGS. Paired-end DNA libraries were generated with a TruSeq DNA Nano library prep kit. Samples were then pooled and sequenced at ~30x coverage on the Illumina NovaSeq6000 platf..., # Unraveling the genetics of feline hypertrophic cardiomyopathy: A multiomics study of 138 cats

    Dataset DOI: 10.5061/dryad.cjsxksnjh

    Description of the data and file structure

    Data available

    1. A population level vcf of polymorphic SNP and indel variants were called among 138 domestic cats with and without hypertrophic cardiomyopathy (HCM). The VCF was generated by mapping paired wgs fastq reads to the Fca126 reference genome with bwa mem and calling variants through GATK4 best practices. Variant annotations were generated with Ensembl's VEP based on Fca126 gene and exon boundaries.  The vcf file contains meta-information lines, followed by a header line specifying fixed fields per sample and subsequent data lines detail variants at genomic positions. The fixed fields include chromosome (CHROM), position (POS), identifier (ID), the reference base(s) (REF), alternate base(s) (ALT), quality (QUAL), filter status (FILTER), and additional information ...,

  6. Comparison of the number of dbSNP, ClinVar and GWAScat sites represented...

    • plos.figshare.com
    xls
    Updated Jun 3, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Alberto Ferrarini; Luciano Xumerle; Francesca Griggio; Marianna Garonzi; Chiara Cantaloni; Cesare Centomo; Sergio Marin Vargas; Patrick Descombes; Julien Marquis; Sebastiano Collino; Claudio Franceschi; Paolo Garagnani; Benjamin A. Salisbury; John Max Harvey; Massimo Delledonne (2023). Comparison of the number of dbSNP, ClinVar and GWAScat sites represented using VCF, gVCF and eVCF files. [Dataset]. http://doi.org/10.1371/journal.pone.0132180.t007
    Explore at:
    xlsAvailable download formats
    Dataset updated
    Jun 3, 2023
    Dataset provided by
    PLOShttp://plos.org/
    Authors
    Alberto Ferrarini; Luciano Xumerle; Francesca Griggio; Marianna Garonzi; Chiara Cantaloni; Cesare Centomo; Sergio Marin Vargas; Patrick Descombes; Julien Marquis; Sebastiano Collino; Claudio Franceschi; Paolo Garagnani; Benjamin A. Salisbury; John Max Harvey; Massimo Delledonne
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    Comparison of the number of dbSNP, ClinVar and GWAScat sites represented using VCF, gVCF and eVCF files.

  7. d

    Annotated VCF of 192 Verticillium dahliae isolates

    • search.dataone.org
    Updated Jul 16, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Benjamin Mimee; Joel Lafond-Lapalme; Mario Tenuta (2025). Annotated VCF of 192 Verticillium dahliae isolates [Dataset]. http://doi.org/10.5061/dryad.g79cnp5v0
    Explore at:
    Dataset updated
    Jul 16, 2025
    Dataset provided by
    Dryad Digital Repository
    Authors
    Benjamin Mimee; Joel Lafond-Lapalme; Mario Tenuta
    Time period covered
    Jan 1, 2023
    Description

    Verticillium dahliae is an important soil-borne pathogen causing Verticillium wilt. It is also the primary causal agent of the Potato Early Dying, a disease complex involving the root-lesion nematode. Here, we report the whole-genome sequencing of 192 isolates of V. dahliae originating from the major potato production areas across Canada. Our results yielded a resource of 277,010 genetic variations that will be useful for genetic analyses and revealed the presence of two major lineages, both present in all provinces but exhibiting differences in regional prevalence., Filtered WGS reads (fastp) aligned on Verticillium dahliae reference (https://www.ncbi.nlm.nih.gov/assembly/95341/GCA_000150675.2) with BWA. VCF called with freebayes v1.3.6 and annotated with snpeff.,

  8. S

    Exome sequencing data of the patients with liver disease

    • scidb.cn
    Updated Apr 10, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Kenan Moral; Gulsum Kayhan; Tarik Duzenli; Sinan Sari; Mehmet Cindoruk; Nergiz Ekmen (2025). Exome sequencing data of the patients with liver disease [Dataset]. http://doi.org/10.57760/sciencedb.23199
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Apr 10, 2025
    Dataset provided by
    Science Data Bank
    Authors
    Kenan Moral; Gulsum Kayhan; Tarik Duzenli; Sinan Sari; Mehmet Cindoruk; Nergiz Ekmen
    License

    Open Database License (ODbL) v1.0https://www.opendatacommons.org/licenses/odbl/1.0/
    License information was derived automatically

    Description

    Exome sequencing data (VCF files) of the nine adult patients with liver diseases.For exome sequencing, the library was prepared using Illumina DNA Prep with Exome 2.5 Enrichment product and sequenced on a NovaSeq 6000 instrument (Illumina, San Diego, CA). Reads were aligned to the GRCh38 Human Reference Genome using the Illumina DRAGEN Bio-IT Platform v3.9.

  9. BETTER Synthetic Healthcare Dataset

    • data.europa.eu
    unknown
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Zenodo, BETTER Synthetic Healthcare Dataset [Dataset]. https://data.europa.eu/data/datasets/oai-zenodo-org-20069305?locale=ro
    Explore at:
    unknown(148013017)Available download formats
    Dataset authored and provided by
    Zenodohttp://zenodo.org/
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    This synthetic dataset release has been generated in the context of the Better Project, funded by the European Union’s Horizon program and the UK Research and Innovation Program. The synthetic datasets mimic healthcare multimodal data (clinical and genomic) from different hospitals. It is intended to be reused for algorithm and pipeline testing for healthcare data management and analysis. The release includes healthcare data combining tabular data (in csv, tsv and xlsx format) and genomic resources in VCF format (Version 4.2 Specification). Three use cases are represented, with subfolders simulating data from distinct hospitals. Content: Use Case 1 condition: Hypotonia synthetic datasets: VCF files and tabular data (xlsx, tsv) description: For each of two hospitals, it is provided tabular data and VCF resources. Tabular data provides information on genomic, phenotypic, biological and clinical data (baseline and dynamic) according to domain rules. Regarding VCF resources, each dataset contains 300 patients, 50% of whom are affected by hypotonia and have a homozygote mutation overlapping the JUN gene region in the hg19 genome (chr1:59246463-59249719), and other random variants distributed throughout the genome. Use Case 2 condition: Inherited Retinal Diseases synthetic datasets: VCF files and tabular data (xlsx) description: For each of two hospitals, it is provided tabular data and VCF resources. Tabular data provides information on electrophysiological and visual field data according to standard clinical eye examination protocols. The dynamic clinical data has been generated to mimic patients showing different clinical progression depending on the affected gene. Measures are generated for both eyes, which may present differences in progression. Regarding VCF resources, each dataset contains 200 patients, of which 50% have mutations overlapping gene ABCA4 and 50% with mutations overlapping gene RPGR. Details on included variants for each gene extracted from ClinVar can be found in the ANNEX below. Use Case 3 condition: Autism Spectrum Disorder synthetic datasets: tabular data (csv and tsv) description: The generated dataset simulates patients with ASD, and provides information on demographic variables, neurodevelopmental and psychiatric comorbidities, family psychiatric history, bullying-related variables, abuse-related variables, sexual and gender-related descriptors, passive ideas of death, self-harm indicators and suicidal-intention-related features. Datasets have been generated with synthetic biases such as correlations and clusters to provide richness to the generated data. No VCF resources are included for this use case. ANNEX (UC2) The variants included in the VCF files for the ABCA4 gene are: ##fileformat=VCFv4.2 ##FORMAT=

  10. Consensus VCF

    • figshare.com
    gz
    Updated Aug 7, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Christina Cuomo (2023). Consensus VCF [Dataset]. http://doi.org/10.6084/m9.figshare.12951845.v1
    Explore at:
    gzAvailable download formats
    Dataset updated
    Aug 7, 2023
    Dataset provided by
    figshare
    Figsharehttp://figshare.com/
    Authors
    Christina Cuomo
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    VCF file of sites called across all pipelines

  11. Genotyping of known SNPs from ClinVar using the VCF and gVCF file formats...

    • plos.figshare.com
    xls
    Updated Jun 3, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Alberto Ferrarini; Luciano Xumerle; Francesca Griggio; Marianna Garonzi; Chiara Cantaloni; Cesare Centomo; Sergio Marin Vargas; Patrick Descombes; Julien Marquis; Sebastiano Collino; Claudio Franceschi; Paolo Garagnani; Benjamin A. Salisbury; John Max Harvey; Massimo Delledonne (2023). Genotyping of known SNPs from ClinVar using the VCF and gVCF file formats and the number of homozygous reference sites and no-calls based on WGS data. [Dataset]. http://doi.org/10.1371/journal.pone.0132180.t004
    Explore at:
    xlsAvailable download formats
    Dataset updated
    Jun 3, 2023
    Dataset provided by
    PLOShttp://plos.org/
    Authors
    Alberto Ferrarini; Luciano Xumerle; Francesca Griggio; Marianna Garonzi; Chiara Cantaloni; Cesare Centomo; Sergio Marin Vargas; Patrick Descombes; Julien Marquis; Sebastiano Collino; Claudio Franceschi; Paolo Garagnani; Benjamin A. Salisbury; John Max Harvey; Massimo Delledonne
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    Genotyping of known SNPs from ClinVar using the VCF and gVCF file formats and the number of homozygous reference sites and no-calls based on WGS data.

  12. n

    PhenoDB

    • neuinfo.org
    Updated Aug 21, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    (2026). PhenoDB [Dataset]. http://identifiers.org/RRID:SCR_016551
    Explore at:
    Dataset updated
    Aug 21, 2026
    Description

    Database for phenotype genotype associations for humans. Used by clinical researchers to store standardized phenotypic information, diagnosis, and pedigree data and then run analyses on VCF files from individuals, families or cohorts with suspected Mendelian disease.

  13. d

    Data from: Acquired dysfunction of CFTR underlies cystic fibrosis-like...

    • search.dataone.org
    Updated Jul 20, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Jody Gookin (2024). Acquired dysfunction of CFTR underlies cystic fibrosis-like disease of the canine gallbladder [Dataset]. http://doi.org/10.5061/dryad.2rbnzs7xq
    Explore at:
    Dataset updated
    Jul 20, 2024
    Dataset provided by
    Dryad Digital Repository
    Authors
    Jody Gookin
    Description

    Mucocele formation in dogs is a unique and enigmatic muco-obstructive disease of the gallbladder caused by amassment of abnormal mucus that bears striking pathological similarity to cystic fibrosis. We investigated the role of CFTR in the pathogenesis of this disease. The location and frequency of disease-associated variants in the coding region of CFTR was compared using whole genome sequence data from 2,642 dogs representing breeds at low-risk, high-risk, or with confirmed disease. Expression, localization, and ion transport activity of CFTR was quantified in control and mucocele gallbladders by NanoString, Western blotting, immunofluorescence imaging, and studies in Ussing chambers. Our results establish significant loss of CFTR-dependent anion secretion by mucocele gallbladder mucosa. A significantly lower quantity of CFTR protein was demonstrated relative to E-cadherin in mucocele compared to control gallbladder mucosa. Immunofluorescence identified CFTR along the apical membrane o..., We used the Whole Animal Genome Sequencing (WAGS) pipeline to identify short nucleotide variants in a dataset of 2,642 dogs encompassing both private and public resources including 1,971 genomes from the Dog10K project. Briefly, the WAGS pipeline used Burrows-Wheeler Alignment tool-MEM to map paired-end reads to the UU_Cfam_GSD_1.0 reference genome. Variant calling was executed with Genome Analysis Toolkit (GATK4), and Ensembl’s Variant Effect Predictor (VEP, RRID:SCR_007931) predicted variant annotations and consequences. From the resulting VEP-processed VCF file, we extracted CFTR genic variants plus variants within 1Kb of the flanking sequence that passed filters. Subsequently, non-reference allele frequencies were calculated for each variant within the control, risk, and affected dog groups. , , # Acquired dysfunction of CFTR underlies cystic fibrosis-like disease of the canine gallbladder.

    https://doi.org/10.5061/dryad.2rbnzs7xq

    This dataset includes supplementary materials for the manuscript entitled Acquired dysfunction of CFTR underlies cystic fibrosis-like disease of the canine gallbladder.

    Description of the data and file structure

    Supplemental Figure S1 illustrates sample procurement and appearance of gallbladder from each of 9 dogs having mucosal RNA extracted for targeted gene expression analysis. Samples of lumen mucosa were obtained by excision from regions devoid of mucus or from which mucus could be gently removed. During sampling (panel A) and after removal of sample (panel B). Remaining panels show each of 9 individual mucocele gallbladders used for mucosal RNA sample collection. Pictures are immediately post-cholecystectomy followed by opening of the gallbladder to expose the lumen.

    **Supplemental Table S1...

  14. d

    Annotated VCF of 192 Verticillium dahliae isolates

    • datadryad.org
    zip
    Updated Feb 2, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Benjamin Mimee; Joel Lafond-Lapalme; Mario Tenuta (2023). Annotated VCF of 192 Verticillium dahliae isolates [Dataset]. http://doi.org/10.5061/dryad.g79cnp5v0
    Explore at:
    zipAvailable download formats
    Dataset updated
    Feb 2, 2023
    Dataset provided by
    Dryad
    Authors
    Benjamin Mimee; Joel Lafond-Lapalme; Mario Tenuta
    Time period covered
    Feb 1, 2023
    Description

    Filtered WGS reads (fastp) aligned on Verticillium dahliae reference (https://www.ncbi.nlm.nih.gov/assembly/95341/GCA_000150675.2) with BWA. VCF called with freebayes v1.3.6 and annotated with snpeff.

  15. F

    The Use of Non-Variant Sites to Improve the Clinical Assessment of...

    • datasetcatalog.nlm.nih.gov
    Updated Jul 6, 2015
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Ferrarini, Alberto; Xumerle, Luciano; Griggio, Francesca; Garonzi, Marianna; Cantaloni, Chiara; Centomo, Cesare; Marin Vargas, Sergio; Descombes, Patrick; Marquis, Julien; Collino, Sebastiano; Franceschi, Claudio; Garagnani, Paolo; A. Salisbury, Benjamin; Max Harvey, John; Delledonne, Massimo (2015). The Use of Non-Variant Sites to Improve the Clinical Assessment of Whole-Genome Sequence Data [Dataset]. http://doi.org/10.1371/journal.pone.0132180
    Explore at:
    Dataset updated
    Jul 6, 2015
    Authors
    Ferrarini, Alberto; Xumerle, Luciano; Griggio, Francesca; Garonzi, Marianna; Cantaloni, Chiara; Centomo, Cesare; Marin Vargas, Sergio; Descombes, Patrick; Marquis, Julien; Collino, Sebastiano; Franceschi, Claudio; Garagnani, Paolo; A. Salisbury, Benjamin; Max Harvey, John; Delledonne, Massimo
    Description

    Genetic testing, which is now a routine part of clinical practice and disease management protocols, is often based on the assessment of small panels of variants or genes. On the other hand, continuous improvements in the speed and per-base costs of sequencing have now made whole exome sequencing (WES) and whole genome sequencing (WGS) viable strategies for targeted or complete genetic analysis, respectively. Standard WGS/WES data analytical workflows generally rely on calling of sequence variants respect to the reference genome sequence. However, the reference genome sequence contains a large number of sites represented by rare alleles, by known pathogenic alleles and by alleles strongly associated to disease by GWAS. It’s thus critical, for clinical applications of WGS and WES, to interpret whether non-variant sites are homozygous for the reference allele or if the corresponding genotype cannot be reliably called. Here we show that an alternative analytical approach based on the analysis of both variant and non-variant sites from WGS data allows to genotype more than 92% of sites corresponding to known SNPs compared to 6% genotyped by standard variant analysis. These include homozygous reference sites of clinical interest, thus leading to a broad and comprehensive characterization of variation necessary to an accurate evaluation of disease risk. Altogether, our findings indicate that characterization of both variant and non-variant clinically informative sites in the genome is necessary to allow an accurate clinical assessment of a personal genome. Finally, we propose a highly efficient extended VCF (eVCF) file format which allows to store genotype calls for sites of clinical interest while remaining compatible with current variant interpretation software.

  16. Genotyping of known SNPs from dbSNP 141 using the VCF and gVCF file formats...

    • plos.figshare.com
    xls
    Updated Jun 3, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Alberto Ferrarini; Luciano Xumerle; Francesca Griggio; Marianna Garonzi; Chiara Cantaloni; Cesare Centomo; Sergio Marin Vargas; Patrick Descombes; Julien Marquis; Sebastiano Collino; Claudio Franceschi; Paolo Garagnani; Benjamin A. Salisbury; John Max Harvey; Massimo Delledonne (2023). Genotyping of known SNPs from dbSNP 141 using the VCF and gVCF file formats and the number of homozygous reference sites and no-calls based on WGS data. [Dataset]. http://doi.org/10.1371/journal.pone.0132180.t003
    Explore at:
    xlsAvailable download formats
    Dataset updated
    Jun 3, 2023
    Dataset provided by
    PLOShttp://plos.org/
    Authors
    Alberto Ferrarini; Luciano Xumerle; Francesca Griggio; Marianna Garonzi; Chiara Cantaloni; Cesare Centomo; Sergio Marin Vargas; Patrick Descombes; Julien Marquis; Sebastiano Collino; Claudio Franceschi; Paolo Garagnani; Benjamin A. Salisbury; John Max Harvey; Massimo Delledonne
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    Genotyping of known SNPs from dbSNP 141 using the VCF and gVCF file formats and the number of homozygous reference sites and no-calls based on WGS data.

  17. d

    PhenoDB

    • dknet.org
    Updated Jul 28, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    (2026). PhenoDB [Dataset]. http://identifiers.org/RRID:SCR_016551
    Explore at:
    Dataset updated
    Jul 28, 2026
    Description

    Database for phenotype genotype associations for humans. Used by clinical researchers to store standardized phenotypic information, diagnosis, and pedigree data and then run analyses on VCF files from individuals, families or cohorts with suspected Mendelian disease.

  18. Human Variant Annotation Datasets

    • console.cloud.google.com
    Updated Jul 22, 2020
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    https://console.cloud.google.com/marketplace/browse?filter=partner:BigQuery%20Public%20Data (2020). Human Variant Annotation Datasets [Dataset]. https://console.cloud.google.com/marketplace/product/bigquery-public-data/human-variant-annotation-public
    Explore at:
    Dataset updated
    Jul 22, 2020
    Dataset provided by
    Googlehttp://google.com/
    BigQueryhttps://cloud.google.com/bigquery
    License

    Open Database License (ODbL) v1.0https://www.opendatacommons.org/licenses/odbl/1.0/
    License information was derived automatically

    Description

    These datasets are important to genomics researchers because they characterize several aspects of what the scientific community has learned to date about human sequence variants. Making this human annotation data freely available in GCP will enable researchers to focus less on data movement and management tasks associated with procuring this data and instead make immediate use of the data to better understand the clinical relevance of particular variant such as disease causing or protective variants (ClinVar), search a catalog of SNPs that have been identified in the human genome (dbSNP), and discover how frequently a particular variant occurs across the human population (1000Genomes, ESP, ExAC, gnomAD) This human annotation dataset contains both a mirror of the original Variant Call Files (VCF) files from NCBI, NHLBI Exome Sequencing Project (ESP) and ensembl as Google Cloud Storage (GCS) objects. In addition, these human sequence variants have also been translated into a particular variant table format and made available in Google BigQuery giving researchers the ability to use cloud technology and code repositories such as the Verily Life Sciences Annotation Toolkit to perform analyses in parallel. This public dataset is hosted in Google BigQuery and is included in BigQuery's 1TB/mo of free tier processing. This means that each user receives 1TB of free BigQuery processing every month, which can be used to run queries on this public dataset. Watch this short video to learn how to get started quickly using BigQuery to access public datasets. What is BigQuery . This public dataset is hosted in Google Cloud Storage and available free to use. Use this quick start guide to quickly learn how to access public datasets on Google Cloud Storage.

  19. r

    PhenoDB

    • rrid.site
    Updated Jan 29, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    (2022). PhenoDB [Dataset]. http://identifiers.org/RRID:SCR_016551
    Explore at:
    Dataset updated
    Jan 29, 2022
    Description

    Database for phenotype genotype associations for humans. Used by clinical researchers to store standardized phenotypic information, diagnosis, and pedigree data and then run analyses on VCF files from individuals, families or cohorts with suspected Mendelian disease.

  20. d

    Data and code from: A multifaceted approach reveals complex genomic...

    • search.dataone.org
    Updated Dec 4, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Samantha Lucy Rita Capel; Devaughn L. Fraser; Kenneth A. Field; DeeAnn M. Reeder; Amy L. Russell; Peter H. Sudmant; Juan Manuel Vazquez; Maarten J. Vonhof; Thomas M. Lilley; Michael R. Buchalski (2025). Data and code from: A multifaceted approach reveals complex genomic mediation of white-nose syndrome resistance in the little brown bat (Myotis lucifugus) [Dataset]. http://doi.org/10.5061/dryad.ncjsxkt66
    Explore at:
    Dataset updated
    Dec 4, 2025
    Dataset provided by
    Dryad Digital Repository
    Authors
    Samantha Lucy Rita Capel; Devaughn L. Fraser; Kenneth A. Field; DeeAnn M. Reeder; Amy L. Russell; Peter H. Sudmant; Juan Manuel Vazquez; Maarten J. Vonhof; Thomas M. Lilley; Michael R. Buchalski
    Description

    Novel pathogens have become a major challenge faced by wildlife in the Anthropocene. White-nose syndrome (WNS), a fungal pathogen, has decimated bat populations across North America over the last two decades. Demographic and physiological evidence of resistance in one heavily affected species, Myotis lucifugus, has prompted multiple attempts to delineate the genomic underpinnings, but they show little congruence in their findings. This may be due, in part, to the limitations of the genomic resources utilized and/or analytical approaches employed. Here, we performed high-coverage whole-genome resequencing of M. lucifugus sampled prior to (n = 29) and 10 years after the arrival of WNS (n = 30), aligned to a new reference genome to identify signatures of selection associated with pathogen resistance. Using 41.9 million SNPs, we implemented a combination of hard and soft sweep detection analyses, leading to discovery of 405 genes with robust evidence of selection. Of these, 241 (59.5 %) wer..., , , # README: A multifaceted approach reveals complex genomic mediation of white-nose syndrome resistance in the little brown bat (Myotis lucifugus)

    Dataset DOI: 10.5061/dryad.ncjsxkt66

    Description of the data and file structure

    This dataset contains all analytical code for this manuscript, along with associated data files, two versions of the final VCF, and selection statistic output files. The genome annotation used with this data is available through https://github.com/docmanny/myotis-gene-annotations.

    Data Files

    VCFs:

    • gatk.snp.qual_hard_filtered_autosomes.vcf.gz -- full set of filtered SNPs
    • gatk.snp.qual_hard_filtered_autosomes_thin.vcf.gz -- filtered SNPs thinned by 10 Kbp distance

    SNP Calling Pipeline Input Files:

    • RG_info.tsv -- read group information for variant calling
    • all_samples.txt -- sample IDs for all individual FASTQ files
    • **all_...,
Share
FacebookFacebook
TwitterTwitter
Email
Click to copy link
Link copied
Close
Cite
Pranckeviciene Erinija (2020). NA12878 WES Benchmark dataset [Dataset]. https://data.niaid.nih.gov/resources?id=zenodo_3597726

NA12878 WES Benchmark dataset

Explore at:
Dataset updated
May 31, 2020
Dataset authored and provided by
Pranckeviciene Erinija
License

Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically

Description

This dataset makes available the UCSC Genome Browser (genome.ucsc.edu) GRCh37 genome build public session NA12878 WES Benchmark files in a single dataset so that these files can be used in other applications or genome browsers such as IGV. All genomic variant calls in all VCF files were decomposed and normalized with vt. This dataset contains: Genome in a bottle (GIAB) version 3.3.2 high confidence (HC) variant calls and genomic regions for HapMap individual NA12878 : GIAB_v3.3.2_NA12878-decomposed-normalized.vcf.gz GIAB_v3.3.2_NA12878-decomposed-normalized.vcf.gz.tbi GIAB_v3.3.2_NA12878_HC_regions.bed HapMap individual NA12878 WES variant calls (VCF) and capture regions (BED) from diagnostic laboratories : ARUP whole exome sequencing data (HiSeq 2000) publically available from NCBI GeT-RM Browser converted_ARUP_NA12878_Exome-decomposed-normalized.vcf.gz converted_ARUP_NA12878_Exome-decomposed-normalized.vcf.gz.tbi ARUP_SeqCap_EZ_Exome.bed UCSF whole exome sequencing data (HiSeq 2500) publically available from NCBI GeT-RM Browser converted_UCSF_NA12878_WES_Agilent_V4_Custom-decomposed-normalized.vcf.gz converted_UCSF_NA12878_WES_Agilent_V4_Custom-decomposed-normalized.vcf.gz.tbi UCSF_WES_Agilent_V4_Custom.bed Whole exome data (NextSeq 500) sequenced in CHEO diagnostic laboratory CHEO_NA12878_WES_S1dataset.vcf.gz CHEO_NA12878_WES_S1dataset.vcf.gz.tbi Agilent_CRE_v2.bed Genomic coordinates (BED) of OMIM genes for which a molecular basis of the associated disease is known (as of September 2019) : Omim_Genes.bed

Search
Clear search
Close search
Google apps
Main menu