Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This dataset makes available the UCSC Genome Browser (genome.ucsc.edu) GRCh37 genome build public session NA12878 WES Benchmark files in a single dataset so that these files can be used in other applications or genome browsers such as IGV. All genomic variant calls in all VCF files were decomposed and normalized with vt. This dataset contains: Genome in a bottle (GIAB) version 3.3.2 high confidence (HC) variant calls and genomic regions for HapMap individual NA12878 : GIAB_v3.3.2_NA12878-decomposed-normalized.vcf.gz GIAB_v3.3.2_NA12878-decomposed-normalized.vcf.gz.tbi GIAB_v3.3.2_NA12878_HC_regions.bed HapMap individual NA12878 WES variant calls (VCF) and capture regions (BED) from diagnostic laboratories : ARUP whole exome sequencing data (HiSeq 2000) publically available from NCBI GeT-RM Browser converted_ARUP_NA12878_Exome-decomposed-normalized.vcf.gz converted_ARUP_NA12878_Exome-decomposed-normalized.vcf.gz.tbi ARUP_SeqCap_EZ_Exome.bed UCSF whole exome sequencing data (HiSeq 2500) publically available from NCBI GeT-RM Browser converted_UCSF_NA12878_WES_Agilent_V4_Custom-decomposed-normalized.vcf.gz converted_UCSF_NA12878_WES_Agilent_V4_Custom-decomposed-normalized.vcf.gz.tbi UCSF_WES_Agilent_V4_Custom.bed Whole exome data (NextSeq 500) sequenced in CHEO diagnostic laboratory CHEO_NA12878_WES_S1dataset.vcf.gz CHEO_NA12878_WES_S1dataset.vcf.gz.tbi Agilent_CRE_v2.bed Genomic coordinates (BED) of OMIM genes for which a molecular basis of the associated disease is known (as of September 2019) : Omim_Genes.bed
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Concordance of genotypes represented in VCF and gVCF files with those detected by the MI RISK Plus kit.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This repository contains two VCF files that were generated from the called genotyped obtained from ANGSD. One VCF file contains variants obtained from whole-genome sequencing data from one individual sample of Aedes bromeliae (aedesBromeliae_1sample_Brom.mdp8.vcf.gz ) The other VCF file contains variants obtained from whole-exome sequencing data from 71 individual samples belonging to different populations in Africa and outside of Africa (aedesAegypti_71samples_AEGY.allpop.J18.AllSites.REFpol.CALLED.vcf_.gz). The data is accompanied by the raw sequencing data available in SRA with the accession number SRP092518 and with the BioProject number PRJNA349448 More information about the project is available in the manuscript that is under preparation.
Facebook
Twitterhttps://ega-archive.org/dacs/EGAC00001000222https://ega-archive.org/dacs/EGAC00001000222
The need for a detailed catalogue of local variability for the study of rare diseases within the context of the Medical Genome Project motivated the whole exome sequencing of 267 unrelated individuals, representative of the healthy Spanish population.
Facebook
TwitterHypertrophic cardiomyopathy (HCM) is the most common inherited cardiac disease in cats, often leading to congestive heart failure, arterial thromboembolism, and sudden cardiac death. The genetics of feline HCM are poorly understood, and limited genetic discoveries remain breed or family-specific. We aimed to identify novel causative or disease-modifying variants in a large cohort of cats reflective of the general cat population. In a second cohort, we sought to characterize transcriptomic differences between HCM-affected cats and healthy controls. DNA was isolated from 138 domestic cats (109 HCM and 29 controls). No single or combination of variants of high, moderate, or modifying impact were identified in genome-wide analysis to cause or modify the disease severity of HCM. Several rare high and moderate-impact variants in genes associated with human HCM were detected in diseased cats. In a second cohort, left ventricular (LV), interventricular septal (IVS), and left atrial (LA) tissues..., WGS data generation A total of 1-2 mL of whole blood were collected from the cephalic, saphenous, or jugular vein into EDTA blood collection tubes. DNA was either isolated from whole blood or from buffy coats after whole blood centrifugation at 2000 rpm for 15 minutes. Genomic DNA isolation was performed using commercially available kits (Gentra Puregene Blood kit, QIAGEN, Hilden Germany; ArchivePure;5Prime) and by following the respective manufacturer’s protocol. High-quality unfragmented DNA was selected by a combination of 1% agarose gel visualization and spectrophotometric confirmation (a 260/280 ratio of ~1.8 and a concentration of > 50 ng/uL; NanoDrop One/One, Thermofisher, Waltham, GA, USA). Samples were stored at -20°C until ready for shipment to Theragen Bio Co., Ltd, Gyeonggi-do, Republic of Korea for WGS. Paired-end DNA libraries were generated with a TruSeq DNA Nano library prep kit. Samples were then pooled and sequenced at ~30x coverage on the Illumina NovaSeq6000 platf..., # Unraveling the genetics of feline hypertrophic cardiomyopathy: A multiomics study of 138 cats
Dataset DOI: 10.5061/dryad.cjsxksnjh
1. A population level vcf of polymorphic SNP and indel variants were called among 138 domestic cats with and without hypertrophic cardiomyopathy (HCM). The VCF was generated by mapping paired wgs fastq reads to the Fca126 reference genome with bwa mem and calling variants through GATK4 best practices. Variant annotations were generated with Ensembl's VEP based on Fca126 gene and exon boundaries.  The vcf file contains meta-information lines, followed by a header line specifying fixed fields per sample and subsequent data lines detail variants at genomic positions. The fixed fields include chromosome (CHROM), position (POS), identifier (ID), the reference base(s) (REF), alternate base(s) (ALT), quality (QUAL), filter status (FILTER), and additional information ...,
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Comparison of the number of dbSNP, ClinVar and GWAScat sites represented using VCF, gVCF and eVCF files.
Facebook
TwitterVerticillium dahliae is an important soil-borne pathogen causing Verticillium wilt. It is also the primary causal agent of the Potato Early Dying, a disease complex involving the root-lesion nematode. Here, we report the whole-genome sequencing of 192 isolates of V. dahliae originating from the major potato production areas across Canada. Our results yielded a resource of 277,010 genetic variations that will be useful for genetic analyses and revealed the presence of two major lineages, both present in all provinces but exhibiting differences in regional prevalence., Filtered WGS reads (fastp) aligned on Verticillium dahliae reference (https://www.ncbi.nlm.nih.gov/assembly/95341/GCA_000150675.2) with BWA. VCF called with freebayes v1.3.6 and annotated with snpeff.,
Facebook
TwitterOpen Database License (ODbL) v1.0https://www.opendatacommons.org/licenses/odbl/1.0/
License information was derived automatically
Exome sequencing data (VCF files) of the nine adult patients with liver diseases.For exome sequencing, the library was prepared using Illumina DNA Prep with Exome 2.5 Enrichment product and sequenced on a NovaSeq 6000 instrument (Illumina, San Diego, CA). Reads were aligned to the GRCh38 Human Reference Genome using the Illumina DRAGEN Bio-IT Platform v3.9.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This synthetic dataset release has been generated in the context of the Better Project, funded by the European Union’s Horizon program and the UK Research and Innovation Program. The synthetic datasets mimic healthcare multimodal data (clinical and genomic) from different hospitals. It is intended to be reused for algorithm and pipeline testing for healthcare data management and analysis. The release includes healthcare data combining tabular data (in csv, tsv and xlsx format) and genomic resources in VCF format (Version 4.2 Specification). Three use cases are represented, with subfolders simulating data from distinct hospitals. Content: Use Case 1 condition: Hypotonia synthetic datasets: VCF files and tabular data (xlsx, tsv) description: For each of two hospitals, it is provided tabular data and VCF resources. Tabular data provides information on genomic, phenotypic, biological and clinical data (baseline and dynamic) according to domain rules. Regarding VCF resources, each dataset contains 300 patients, 50% of whom are affected by hypotonia and have a homozygote mutation overlapping the JUN gene region in the hg19 genome (chr1:59246463-59249719), and other random variants distributed throughout the genome. Use Case 2 condition: Inherited Retinal Diseases synthetic datasets: VCF files and tabular data (xlsx) description: For each of two hospitals, it is provided tabular data and VCF resources. Tabular data provides information on electrophysiological and visual field data according to standard clinical eye examination protocols. The dynamic clinical data has been generated to mimic patients showing different clinical progression depending on the affected gene. Measures are generated for both eyes, which may present differences in progression. Regarding VCF resources, each dataset contains 200 patients, of which 50% have mutations overlapping gene ABCA4 and 50% with mutations overlapping gene RPGR. Details on included variants for each gene extracted from ClinVar can be found in the ANNEX below. Use Case 3 condition: Autism Spectrum Disorder synthetic datasets: tabular data (csv and tsv) description: The generated dataset simulates patients with ASD, and provides information on demographic variables, neurodevelopmental and psychiatric comorbidities, family psychiatric history, bullying-related variables, abuse-related variables, sexual and gender-related descriptors, passive ideas of death, self-harm indicators and suicidal-intention-related features. Datasets have been generated with synthetic biases such as correlations and clusters to provide richness to the generated data. No VCF resources are included for this use case. ANNEX (UC2) The variants included in the VCF files for the ABCA4 gene are: ##fileformat=VCFv4.2 ##FORMAT=
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
VCF file of sites called across all pipelines
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Genotyping of known SNPs from ClinVar using the VCF and gVCF file formats and the number of homozygous reference sites and no-calls based on WGS data.
Facebook
TwitterDatabase for phenotype genotype associations for humans. Used by clinical researchers to store standardized phenotypic information, diagnosis, and pedigree data and then run analyses on VCF files from individuals, families or cohorts with suspected Mendelian disease.
Facebook
TwitterMucocele formation in dogs is a unique and enigmatic muco-obstructive disease of the gallbladder caused by amassment of abnormal mucus that bears striking pathological similarity to cystic fibrosis. We investigated the role of CFTR in the pathogenesis of this disease. The location and frequency of disease-associated variants in the coding region of CFTR was compared using whole genome sequence data from 2,642 dogs representing breeds at low-risk, high-risk, or with confirmed disease. Expression, localization, and ion transport activity of CFTR was quantified in control and mucocele gallbladders by NanoString, Western blotting, immunofluorescence imaging, and studies in Ussing chambers. Our results establish significant loss of CFTR-dependent anion secretion by mucocele gallbladder mucosa. A significantly lower quantity of CFTR protein was demonstrated relative to E-cadherin in mucocele compared to control gallbladder mucosa. Immunofluorescence identified CFTR along the apical membrane o..., We used the Whole Animal Genome Sequencing (WAGS) pipeline to identify short nucleotide variants in a dataset of 2,642 dogs encompassing both private and public resources including 1,971 genomes from the Dog10K project. Briefly, the WAGS pipeline used Burrows-Wheeler Alignment tool-MEM to map paired-end reads to the UU_Cfam_GSD_1.0 reference genome. Variant calling was executed with Genome Analysis Toolkit (GATK4), and Ensembl’s Variant Effect Predictor (VEP, RRID:SCR_007931) predicted variant annotations and consequences. From the resulting VEP-processed VCF file, we extracted CFTR genic variants plus variants within 1Kb of the flanking sequence that passed filters. Subsequently, non-reference allele frequencies were calculated for each variant within the control, risk, and affected dog groups. , , # Acquired dysfunction of CFTR underlies cystic fibrosis-like disease of the canine gallbladder.
https://doi.org/10.5061/dryad.2rbnzs7xq
This dataset includes supplementary materials for the manuscript entitled Acquired dysfunction of CFTR underlies cystic fibrosis-like disease of the canine gallbladder.
Supplemental Figure S1 illustrates sample procurement and appearance of gallbladder from each of 9 dogs having mucosal RNA extracted for targeted gene expression analysis. Samples of lumen mucosa were obtained by excision from regions devoid of mucus or from which mucus could be gently removed. During sampling (panel A) and after removal of sample (panel B). Remaining panels show each of 9 individual mucocele gallbladders used for mucosal RNA sample collection. Pictures are immediately post-cholecystectomy followed by opening of the gallbladder to expose the lumen.
**Supplemental Table S1...
Facebook
TwitterFiltered WGS reads (fastp) aligned on Verticillium dahliae reference (https://www.ncbi.nlm.nih.gov/assembly/95341/GCA_000150675.2) with BWA. VCF called with freebayes v1.3.6 and annotated with snpeff.
Facebook
TwitterGenetic testing, which is now a routine part of clinical practice and disease management protocols, is often based on the assessment of small panels of variants or genes. On the other hand, continuous improvements in the speed and per-base costs of sequencing have now made whole exome sequencing (WES) and whole genome sequencing (WGS) viable strategies for targeted or complete genetic analysis, respectively. Standard WGS/WES data analytical workflows generally rely on calling of sequence variants respect to the reference genome sequence. However, the reference genome sequence contains a large number of sites represented by rare alleles, by known pathogenic alleles and by alleles strongly associated to disease by GWAS. It’s thus critical, for clinical applications of WGS and WES, to interpret whether non-variant sites are homozygous for the reference allele or if the corresponding genotype cannot be reliably called. Here we show that an alternative analytical approach based on the analysis of both variant and non-variant sites from WGS data allows to genotype more than 92% of sites corresponding to known SNPs compared to 6% genotyped by standard variant analysis. These include homozygous reference sites of clinical interest, thus leading to a broad and comprehensive characterization of variation necessary to an accurate evaluation of disease risk. Altogether, our findings indicate that characterization of both variant and non-variant clinically informative sites in the genome is necessary to allow an accurate clinical assessment of a personal genome. Finally, we propose a highly efficient extended VCF (eVCF) file format which allows to store genotype calls for sites of clinical interest while remaining compatible with current variant interpretation software.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Genotyping of known SNPs from dbSNP 141 using the VCF and gVCF file formats and the number of homozygous reference sites and no-calls based on WGS data.
Facebook
TwitterDatabase for phenotype genotype associations for humans. Used by clinical researchers to store standardized phenotypic information, diagnosis, and pedigree data and then run analyses on VCF files from individuals, families or cohorts with suspected Mendelian disease.
Facebook
TwitterOpen Database License (ODbL) v1.0https://www.opendatacommons.org/licenses/odbl/1.0/
License information was derived automatically
These datasets are important to genomics researchers because they characterize several aspects of what the scientific community has learned to date about human sequence variants. Making this human annotation data freely available in GCP will enable researchers to focus less on data movement and management tasks associated with procuring this data and instead make immediate use of the data to better understand the clinical relevance of particular variant such as disease causing or protective variants (ClinVar), search a catalog of SNPs that have been identified in the human genome (dbSNP), and discover how frequently a particular variant occurs across the human population (1000Genomes, ESP, ExAC, gnomAD) This human annotation dataset contains both a mirror of the original Variant Call Files (VCF) files from NCBI, NHLBI Exome Sequencing Project (ESP) and ensembl as Google Cloud Storage (GCS) objects. In addition, these human sequence variants have also been translated into a particular variant table format and made available in Google BigQuery giving researchers the ability to use cloud technology and code repositories such as the Verily Life Sciences Annotation Toolkit to perform analyses in parallel. This public dataset is hosted in Google BigQuery and is included in BigQuery's 1TB/mo of free tier processing. This means that each user receives 1TB of free BigQuery processing every month, which can be used to run queries on this public dataset. Watch this short video to learn how to get started quickly using BigQuery to access public datasets. What is BigQuery . This public dataset is hosted in Google Cloud Storage and available free to use. Use this quick start guide to quickly learn how to access public datasets on Google Cloud Storage.
Facebook
TwitterDatabase for phenotype genotype associations for humans. Used by clinical researchers to store standardized phenotypic information, diagnosis, and pedigree data and then run analyses on VCF files from individuals, families or cohorts with suspected Mendelian disease.
Facebook
TwitterNovel pathogens have become a major challenge faced by wildlife in the Anthropocene. White-nose syndrome (WNS), a fungal pathogen, has decimated bat populations across North America over the last two decades. Demographic and physiological evidence of resistance in one heavily affected species, Myotis lucifugus, has prompted multiple attempts to delineate the genomic underpinnings, but they show little congruence in their findings. This may be due, in part, to the limitations of the genomic resources utilized and/or analytical approaches employed. Here, we performed high-coverage whole-genome resequencing of M. lucifugus sampled prior to (n = 29) and 10 years after the arrival of WNS (n = 30), aligned to a new reference genome to identify signatures of selection associated with pathogen resistance. Using 41.9 million SNPs, we implemented a combination of hard and soft sweep detection analyses, leading to discovery of 405 genes with robust evidence of selection. Of these, 241 (59.5 %) wer..., , , # README: A multifaceted approach reveals complex genomic mediation of white-nose syndrome resistance in the little brown bat (Myotis lucifugus)
Dataset DOI: 10.5061/dryad.ncjsxkt66
This dataset contains all analytical code for this manuscript, along with associated data files, two versions of the final VCF, and selection statistic output files. The genome annotation used with this data is available through https://github.com/docmanny/myotis-gene-annotations.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This dataset makes available the UCSC Genome Browser (genome.ucsc.edu) GRCh37 genome build public session NA12878 WES Benchmark files in a single dataset so that these files can be used in other applications or genome browsers such as IGV. All genomic variant calls in all VCF files were decomposed and normalized with vt. This dataset contains: Genome in a bottle (GIAB) version 3.3.2 high confidence (HC) variant calls and genomic regions for HapMap individual NA12878 : GIAB_v3.3.2_NA12878-decomposed-normalized.vcf.gz GIAB_v3.3.2_NA12878-decomposed-normalized.vcf.gz.tbi GIAB_v3.3.2_NA12878_HC_regions.bed HapMap individual NA12878 WES variant calls (VCF) and capture regions (BED) from diagnostic laboratories : ARUP whole exome sequencing data (HiSeq 2000) publically available from NCBI GeT-RM Browser converted_ARUP_NA12878_Exome-decomposed-normalized.vcf.gz converted_ARUP_NA12878_Exome-decomposed-normalized.vcf.gz.tbi ARUP_SeqCap_EZ_Exome.bed UCSF whole exome sequencing data (HiSeq 2500) publically available from NCBI GeT-RM Browser converted_UCSF_NA12878_WES_Agilent_V4_Custom-decomposed-normalized.vcf.gz converted_UCSF_NA12878_WES_Agilent_V4_Custom-decomposed-normalized.vcf.gz.tbi UCSF_WES_Agilent_V4_Custom.bed Whole exome data (NextSeq 500) sequenced in CHEO diagnostic laboratory CHEO_NA12878_WES_S1dataset.vcf.gz CHEO_NA12878_WES_S1dataset.vcf.gz.tbi Agilent_CRE_v2.bed Genomic coordinates (BED) of OMIM genes for which a molecular basis of the associated disease is known (as of September 2019) : Omim_Genes.bed