Reference. Structured reviews for data and knowledge-driven research
Hypothesis generation is a critical step in research and a cornerstone in the rare disease field. Research is most efficient when those hypotheses are based on the entirety of knowledge known to date. Systematic review articles are commonly used in biomedicine to summarize existing knowledge and contextualize experimental data. But the information contained within review articles is typically only expressed as free-text, which is difficult to use computationally. Researchers struggle to navigate, collect and remix prior knowledge as it is scattered in several silos without seamless integration and access. This lack of a structured information framework hinders research by both experimental and computational scientists. To better organize knowledge and data, we built a structured review article that is specifically focused on NGLY1 Deficiency, an ultra-rare genetic disease first reported in 2012. We represented this structured review as a knowledge graph and then stored this knowledge graph in a Neo4j database to simplify dissemination, querying and visualization of the network. Relative to free-text, this structured review better promotes the principles of findability, accessibility, interoperability and reusability (FAIR). In collaboration with domain experts in NGLY1 Deficiency, we demonstrate how this resource can improve the efficiency and comprehensiveness of hypothesis generation. We also developed a read–write interface that allows domain experts to contribute FAIR structured knowledge to this community resource. In contrast to traditional free-text review articles, this structured review exists as a living knowledge graph that is curated by humans and accessible to computational analyses. Finally, we have generalized this workflow into modular and repurposable components that can be applied to other domain areas. This NGLY1 Deficiency-focused network is publicly available at http://ngly1graph.org/. Availability and implementation Database URL: http://ngly1graph.org/. Network data files are at: https://github.com/SuLab/ngly1-graph and source code at: https://github.com/SuLab/bioknowledge-reviewer. Contact asu@scripps.edu
Cite
Cites 61 works (0 here)
External (61)
- The BioGRID interaction database: 2019 update (2019)
- The gene ontology resource: 20 years and still GOing strong (2019)
- Reproducibility and replicability of systematic reviews (2019)
- N-Glycanase 1 Regulates Aquaporins Independent of Its Enzymatic Activity (2019)
- Protein sequence editing of SKN-1A/Nrf1 by peptide:N-glycanase controls proteasome gene expression (2019)
- Expansion of the human phenotype ontology (HPO) knowledge base and resources (2019)
- STRING v11: protein–protein association networks with increased coverage, supporting functional discovery in genome-wide experimental datasets (2019)
- UniProt: a worldwide hub of protein knowledge (2019)
- BioGraph: a web application and a graph database for querying and analyzing bioinformatics resources (2018)
- NBDC RDF portal: a comprehensive repository for semantic data in life sciences (2018)
- Transcriptome and functional analysis in a drosophila model of NGLY1 deficiency provides insight into therapeutic approaches (2018)
- HMDB 4.0: the human metabolome database for (2018)
- The reactome pathway knowledgebase (2018)
- Towards FAIRer biological knowledge networks using a hybrid linked data and graph database approach (2018)
- The monarch initiative: an integrative data and analytic platform connecting phenotypes to genotypes across species (2017)
- Inhibition of NGLY1 inactivates the transcription factor Nrf1 and potentiates proteasome inhibitor cytotoxicity (2017)
- Prospective phenotyping of NGLY1-CDDG, the first congenital disorder of deglycosylation (2017)
- KEGG: new perspectives on genomes, pathways, diseases and drugs (2017)
- The ChEMBL database in 2017 (2017)
- Aberrant control of NF-κB in cancer permits transcriptional and phenotypic plasticity, to curtail dependence on host tissue: molecular mode (2017)
- Wikidata as a semantic framework for the gene wiki initiative (2016)
- High-performance web services for querying gene and variant annotation (2016)
- Representing and querying disease networks using graph databases (2016)
- Proteasome dysfunction triggers activation of SKN-1A/Nrf1 by the aspartic protease DDI-1 (2016)
- ChEBI in 2016: improved services and an expanding collection of metabolites (2016)
- Endothelial Aquaporin-1 (AQP1) expression is regulated by transcription factor Mef2c (2016)
- The FAIR guiding principles for scientific data management and stewardship (2016)
- KaBOB: ontology-based semantic integration of biomedical databases (2015)
- The molecular signatures database Hallmark gene set collection (2015)
- TRRUST: a reference database of human transcriptional regulatory interactions (2015)
- Disease ontology 2015 update: an expanded and updated database of human diseases for linking biomedical knowledge through disease data (2015)
- Dynamic aberrant NF-κB spurs tumorigenesis: a new model encompassing the microenvironment (2015)
- Biomedical question answering using semantic relations (2015)
- The EBI RDF platform: linked open data for the life sciences (2014)
- The application of the open pharmacological concepts triple store (open PHACTS) to support drug discovery research (2014)
- Living systematic reviews: an emerging opportunity to narrow the evidence-practice gap (2014)
- Mutations in NGLY1 cause an inherited disorder of the endoplasmic reticulum–associated degradation pathway (2014)
- BioGPS and MyGene.Info: organizing online, gene-centric information (2013)
- Clinical application of exome sequencing in undiagnosed genetic conditions (2012)
- An integrated encyclopedia of DNA elements in the human genome (2012)
- Circuitry and dynamics of human transcription factor regulatory networks (2012)
- SemMedDB: a PubMed-scale repository of biomedical semantic predications (2012)
- Wikidata: a new platform for collaborative data collection (2012)
- HyQue: evaluating hypotheses using semantic web technologies (2011)
- BioGraph: unsupervised biomedical knowledge discovery via automated hypothesis generation (2011)
- EpiphaNet: an interactive tool to support biomedical discoveries (2010)
- Biomedical discovery acceleration, with applications to craniofacial development (2009)
- Human protein reference database—2009 update (2009)
- The mammalian phenotype ontology: enabling robust annotation and comparative analysis (2009)
- TRED: a transcriptional regulatory element database, new entries and other development (2007)
- Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles (2005)
- Systematic discovery of regulatory motifs in human promoters and 3′ UTRs by comparison of several mammals (2005)
- The interaction of domain knowledge and linguistic structure in natural language processing: interpreting hypernymic propositions in biomedical text (2003)
- NF-κB functions in synaptic signaling and behavior (2003)
- Hepatitis C virus NS5A protein binds TBP and p53, inhibiting their DNA binding and p53 interactions with TBP and ERCC3 (2002)
- Negative regulation of bcl-2 expression by p53 in hematopoietic cells (2001)
- Evidence for the involvement of TNF and NF-κB in hippocampal synaptic plasticity (2000)
- Pax-6 interactions with TATA-box-binding protein and retinoblastoma protein (1999)
- Spatial attention deficits in patients with acquired or developmental cerebellar abnormality (1999)
- Transactivation by the human cytomegalovirus IE2 86-kilodalton protein requires a domain that binds to both the TATA box-binding protein and the retinoblastoma protein (1994)
- Wild-type p53 binds to the TATA-binding protein and represses transcription (1992)