Exploring massive, genome scale datasets with the genometricorr package
Cope, Leslie M.
Mironov, Andrey A.
Makeev, Vsevolod J.
Wheelan, Sarah J.
KAUST DepartmentComputational Bioscience Research Center (CBRC)
Permanent link to this recordhttp://hdl.handle.net/10754/325275
MetadataShow full item record
AbstractWe have created a statistically grounded tool for determining the correlation of genomewide data with other datasets or known biological features, intended to guide biological exploration of high-dimensional datasets, rather than providing immediate answers. The software enables several biologically motivated approaches to these data and here we describe the rationale and implementation for each approach. Our models and statistics are implemented in an R package that efficiently calculates the spatial correlation between two sets of genomic intervals (data and/or annotated features), for use as a metric of functional interaction. The software handles any type of pointwise or interval data and instead of running analyses with predefined metrics, it computes the significance and direction of several types of spatial association; this is intended to suggest potentially relevant relationships between the datasets. Availability and implementation: The package, GenometriCorr, can be freely downloaded at http://genometricorr.sourceforge.net/. Installation guidelines and examples are available from the sourceforge repository. The package is pending submission to Bioconductor. © 2012 Favorov et al.
CitationFavorov A, Mularoni L, Cope LM, Medvedeva Y, Mironov AA, et al. (2012) Exploring Massive, Genome Scale Datasets with the GenometriCorr Package. PLoS Comput Biol 8: e1002529. doi:10.1371/journal.pcbi.1002529.
PublisherPublic Library of Science (PLoS)
JournalPLoS Computational Biology
PubMed Central IDPMC3364938
- Girafe--an R/Bioconductor package for functional exploration of aligned next-generation sequencing reads.
- Authors: Toedling J, Ciaudo C, Voinnet O, Heard E, Barillot E
- Issue date: 2010 Nov 15
- The Integrated Genome Browser: free software for distribution and exploration of genome-scale datasets.
- Authors: Nicol JW, Helt GA, Blanchard SG Jr, Raja A, Loraine AE
- Issue date: 2009 Oct 15
- rtracklayer: an R package for interfacing with genome browsers.
- Authors: Lawrence M, Gentleman R, Carey V
- Issue date: 2009 Jul 15
- GOrilla: a tool for discovery and visualization of enriched GO terms in ranked gene lists.
- Authors: Eden E, Navon R, Steinfeld I, Lipson D, Yakhini Z
- Issue date: 2009 Feb 3
- Rintact: enabling computational analysis of molecular interaction data from the IntAct repository.
- Authors: Chiang T, Li N, Orchard S, Kerrien S, Hermjakob H, Gentleman R, Huber W
- Issue date: 2008 Apr 15
Showing items related by title, author, creator and subject.
Long- and short-term selective forces on malaria parasite genomesNygaard, Sanne; Braunstein, Alexander; Malsen, Gareth; Van Dongen, Stijn; Gardner, Paul P.; Krogh, Anders; Otto, Thomas D.; Pain, Arnab; Berriman, Matthew; McAuliffe, Jon; Dermitzakis, Emmanouil T.; Jeffares, Daniel C. (PLoS Genetics, Public Library of Science (PLoS), 2010-09-09) [Article]Plasmodium parasites, the causal agents of malaria, result in more than 1 million deaths annually. Plasmodium are unicellular eukaryotes with small ~23 Mb genomes encoding ~5200 protein-coding genes. The protein-coding genes comprise about half of these genomes. Although evolutionary processes have a significant impact on malaria control, the selective pressures within Plasmodium genomes are poorly understood, particularly in the non-protein-coding portion of the genome. We use evolutionary methods to describe selective processes in both the coding and non-coding regions of these genomes. Based on genome alignments of seven Plasmodium species, we show that protein-coding, intergenic and intronic regions are all subject to purifying selection and we identify 670 conserved non-genic elements. We then use genome-wide polymorphism data from P. falciparum to describe short-term selective processes in this species and identify some candidate genes for balancing (diversifying) selection. Our analyses suggest that there are many functional elements in the non-genic regions of these genomes and that adaptive evolution has occurred more frequently in the protein-coding regions of the genome. © 2010 Nygaard et al.
Pivotal role of the muscle-contraction pathway in cryptorchidism and evidence for genomic connections with cardiomyopathy pathways in RASopathiesCannistraci, Carlo; Ogorevc, Jernej; Zorc, Minja; Ravasi, Timothy; Dovc, Peter; Kunej, Tanja (BMC Medical Genomics, Springer Nature, 2013-02-14) [Article]Background: Cryptorchidism is the most frequent congenital disorder in male children; however the genetic causes of cryptorchidism remain poorly investigated. Comparative integratomics combined with systems biology approach was employed to elucidate genetic factors and molecular pathways underlying testis descent. Methods. Literature mining was performed to collect genomic loci associated with cryptorchidism in seven mammalian species. Information regarding the collected candidate genes was stored in MySQL relational database. Genomic view of the loci was presented using Flash GViewer web tool (http://gmod.org/wiki/Flashgviewer/). DAVID Bioinformatics Resources 6.7 was used for pathway enrichment analysis. Cytoscape plug-in PiNGO 1.11 was employed for protein-network-based prediction of novel candidate genes. Relevant protein-protein interactions were confirmed and visualized using the STRING database (version 9.0). Results. The developed cryptorchidism gene atlas includes 217 candidate loci (genes, regions involved in chromosomal mutations, and copy number variations) identified at the genomic, transcriptomic, and proteomic level. Human orthologs of the collected candidate loci were presented using a genomic map viewer. The cryptorchidism gene atlas is freely available online: http://www.integratomics-time.com/cryptorchidism/. Pathway analysis suggested the presence of twelve enriched pathways associated with the list of 179 literature-derived candidate genes. Additionally, a list of 43 network-predicted novel candidate genes was significantly associated with four enriched pathways. Joint pathway analysis of the collected and predicted candidate genes revealed the pivotal importance of the muscle-contraction pathway in cryptorchidism and evidence for genomic associations with cardiomyopathy pathways in RASopathies. Conclusions: The developed gene atlas represents an important resource for the scientific community researching genetics of cryptorchidism. The collected data will further facilitate development of novel genetic markers and could be of interest for functional studies in animals and human. The proposed network-based systems biology approach elucidates molecular mechanisms underlying co-presence of cryptorchidism and cardiomyopathy in RASopathies. Such approach could also aid in molecular explanation of co-presence of diverse and apparently unrelated clinical manifestations in other syndromes. 2013 Cannistraci et al.; licensee BioMed Central Ltd.
Bacterial niche-specific genome expansion is coupled with highly frequent gene disruptions in deep-sea sedimentsWang, Yong; Yang, Jiang Ke; Lee, On On; Li, Tie Gang; Al-Suwailem, Abdulaziz M.; Danchin, Antoine; Qian, Pei-Yuan (PLoS ONE, Public Library of Science (PLoS), 2011-12-21) [Article]The complexity and dynamics of microbial metagenomes may be evaluated by genome size, gene duplication and the disruption rate between lineages. In this study, we pyrosequenced the metagenomes of microbes obtained from the brine and sediment of a deep-sea brine pool in the Red Sea to explore the possible genomic adaptations of the microbes in response to environmental changes. The microbes from the brine and sediments (both surface and deep layers) of the Atlantis II Deep brine pool had similar communities whereas the effective genome size varied from 7.4 Mb in the brine to more than 9 Mb in the sediment. This genome expansion in the sediment samples was due to gene duplication as evidenced by enrichment of the homologs. The duplicated genes were highly disrupted, on average by 47.6% and 70% for the surface and deep layers of the Atlantis II Deep sediment samples, respectively. The disruptive effects appeared to be mainly due to point mutations and frameshifts. In contrast, the homologs from the Atlantis II Deep brine sample were highly conserved and they maintained relatively small copy numbers. Likely, the adaptation of the microbes in the sediments was coupled with pseudogenizations and possibly functional diversifications of the paralogs in the expanded genomes. The maintenance of the pseudogenes in the large genomes is discussed. © 2011 Wang et al.