Drug target identification for idiopathic pulmonary fibrosis (IPF) via geometric deep learning on a knowledge graph

Machine learning
Target identification
Knowledge Graph
Neural Network
A STRING GCN recovers known clinical-stage idiopathic pulmonary fibrosis (IPF) targets at AUPRC 0.322, versus 0.305 for XGBoost and 0.098 for a GCN without edges.
Published

2026-08-16

14 min

1 Goals

Back at my data-scientist role at Idorsia, one of our main mandates was to identify new drug targets for focus diseases. We used a mix of methods: simple differential expression between disease and healthy samples, gradient boosting, and more advanced geometric models and autoencoders.

I wanted to take those learnings to idiopathic pulmonary fibrosis (IPF). It is relatively rare, yet a high-priority disease: progressive lung scarring [1], historically a median survival of about 3 to 5 years after diagnosis [2], a severe quality-of-life burden, high cost per patient [3], and only a handful of antifibrotics that slow decline [4,5]. Lung transplantation is the one potentially curative option. The medical need is obvious; the biology is hard (epithelial injury, fibroblasts, TGF-β, extracellular matrix) [6]. IPF was a natural disease in which to ask whether a protein knowledge graph, used as structure in a GCN [7], recovers more information than simpler models on the same public features. In other words: does the extra complexity pay off?

Infographic summarizing the IPF ranker: STRING-GCN AUPRC 0.322, XGBoost 0.305, GCN without edges 0.098, Open Targets 0.329 as the positive control.

STRING edges earn their keep for IPF target ranking

2 Data

Piece Freeze
Disease IPF (EFO_0000768)
Universe 4,283 Open Targets associated genes [8]
Labels 139 genes with a clinical-stage mechanism of action (phase ≥ 1)
Chance AUPRC 0.032 (prevalence)
Graph STRING v12, combined score ≥ 700 [9]

I used 18 public features, one row per gene. Each feature is numeric and rank-scaled toward [0, 1]; 0 means no evidence in that feature.

Genetics. Open Targets genetic association for IPF [8], plus a sparse GWAS Catalog flag for IPF and interstitial-lung-disease genes [10].

Expression and tissue. A GEO meta-analysis of IPF versus control lung series [11], the Open Targets RNA-expression association [8], lung specificity from GTEx [12] and the Human Protein Atlas [13], and a combined feature that takes the max of those lung-expression views.

Pathways and regulation. Membership in Reactome pathways enriched among differential-expression seeds [14]; a transcription-factor feature whose targets match the GEO signature, from CollecTRI [15]; Open Targets animal-model evidence [8].

STRING neighborhood. Tabular geometry on the same protein network the GCN later uses [9]: how many neighbors are IPF-changed genes, whether the gene sits in a community of those genes, a network walk from them, centrality in IPF gene modules, and the same module idea seeded by genetics. The GCN still mixes those features along edges.

Perturbation. LINCS L1000 knockdown and overexpression connectivity versus the IPF DE signature [16], and DepMap CRISPR essentiality averaged over lung cancer lines [17]. That last feature is a lung prior, not an IPF prior.

The model input is those 18 features. The labels are a known-drug proxy: 139 genes with a clinical-stage mechanism of action (phase ≥ 1). Genes without that label stay unlabeled (positive-unlabeled ranking). I kept Open Targets clinical association out of the features; it defines the labels.

The implementation for this freeze is at github.com/alanzos/ipf-target-ranker.

3 Methods

3.1 Models

I trained every model on the same 18 features. Graph models also saw STRING edges. That isolates architecture from feature engineering.

Model What it is Strength Limit
XGBoost Gradient-boosted trees on the 18 features. Inductive: a held-out gene is predicted from its row alone. Strong default on mixed, sparse tabular features; early stopping is straightforward. No message passing. STRING geometry only enters through neighborhood features I already computed.
Tabular neural net Two-layer net (64 then 32), dropout 0.3. Nonlinear tabular control without trees. Few labeled genes; overfits more easily than trees on this matrix.
STRING-GCN Two GCN layers, hidden 32, using STRING confidence as the edge weight. Each gene takes a degree-normalized average of its neighbors [7]. Cheap mixer; matches homophilous PPI; uses the combined STRING score. Neighbors are mixed uniformly (up to degree and edge weight). Training is transductive.
STRING-GAT Two GAT layers, 16 hidden units × 4 heads [18]. Can down-weight unhelpful neighbors. More parameters; this freeze drops edge weights; ~111 labeled genes per train fold and ~82k edges. Attention is extra capacity that may not pay.
GCN without edges Same GCN, empty edge set (self-loops only). The control that shows whether edges add anything. Not a production ranker.
GCN on degree only STRING edges, features reduced to degree. Hub check: tests whether ranking is high-degree proteins first. Throws away the 18 features.
Shuffled labels Same XGBoost, labels permuted. Sanity check that AUPRC is not a metric bug. Not a ranker.
Open Targets Frozen Open Targets overall association. Not trained. The positive control. Combines every Open Targets evidence type, including the clinical evidence that defines the labels [8]. That is the strongest public Open Targets ranking I can put on the same genes. The clinical evidence that defines the labels sits inside this score.

I compared STRING-GCN with GCN without edges and with XGBoost. GAT is the attention alternative on the same graph.

3.2 Training, validation, and testing

The unit is a gene, not a sample. I split with repeated stratified 5-fold CV (RepeatedStratifiedKFold, 5 splits × 5 repeats, seed 42).

Stratified. There are only 139 labeled genes in 4,283 (about 3%). A random fold can easily get too few labeled genes for AUPRC to mean anything. Stratifying on the labels keeps that prevalence in every test fold (about 28 labeled genes among ~857 genes).

5-fold. In one pass, the universe is partitioned into five test sets. Each gene is held out as test once. The other four fifths are the outer pool for that fold, then split 80/20 into train and early-stopping validation, so a gene is not in the training loss on all four of those folds.

Repeated. AUPRC on ~28 labeled genes is noisy. I shuffled and repeated the 5-fold partition five times, still with seed 42. That gives 25 test folds. The numbers I report are mean ± sd over those 25 evaluations, not a single lucky split. Across repeats a gene is held out more than once; the illustration ranking later averages those held-out predictions.

Train / validation / test inside each fold. The outer test genes (~857) are metrics only. The remaining genes split 80/20 into train (~2,740) and validation (~686). Validation is early stopping only: XGBoost watches AUPRC for 40 rounds; GCN and GAT watch full-graph validation AUPRC with patience 15; the tabular neural net uses patience 12. Test genes never enter the loss, the stopper, or the scaler fit. I froze hyperparameters before this confirmation run. The inner split is not a nested search.

XGBoost and the tabular neural net are inductive: a held-out gene is predicted from its features alone. The GNN is transductive: one forward pass on the full 4,283-gene graph each epoch, with binary cross-entropy only on train genes. I masked test labels. Test features still flow on STRING edges, so a test gene can be predicted from its neighbors. That is why I needed GCN without edges: it is the comparison that answers whether the graph helped.

Class imbalance is handled with a class weight from the training counts. Unlabeled genes stay unlabeled; I did not downsample them away.

3.3 Benchmark metrics

I used AUPRC on held-out genes as the comparison, mean ± sd over the 25 test folds. That is the usual cross-validation summary: one AUPRC per held-out slice, then the mean. Chance AUPRC equals prevalence, about 0.032. With 3% labeled genes, AUROC can look fine while the top of the list is junk [19]. Accuracy is worse: a model that ranks nobody still looks 97% correct.

The comparisons I care about are STRING-GCN versus GCN without edges (whether edges helped), versus XGBoost (whether geometry beat trees on the same features), versus shuffled labels (whether anything sat above chance), and versus Open Targets (the positive control).

Open Targets overall already mixes genetics, expression, animal models, text, and known-drug / clinical evidence into one association [8]. The clinical piece is the same evidence that defines the labels, which I kept out of the 18 features. That is why it is the positive control: the strongest public Open Targets ranking on these genes, clinical evidence included.

For each trained model I paired the 25 fold AUPRCs with the Open Targets AUPRC on the same test genes and ran a two-sided Wilcoxon signed-rank test on those differences [20]. Seven models, so I Holm-adjusted the p-values [21]. The AUPRC plot writes those tests after each mean ± sd: n.s. if p ≥ 0.05, one star if p < 0.05, two stars if p < 0.01, three stars if p < 0.001.

3.4 Feature importance

I ranked the 18 features by how much the trained STRING-GCN depends on each one. On the same 25 test folds as the confirmation run, I trained STRING-GCN, recorded held-out AUPRC, then shuffled one feature on every gene and recorded held-out AUPRC again. The drop is baseline minus shuffled. I average those drops over the 25 folds. The trained weights stay fixed; only that feature is shuffled. Because training is transductive, the shuffle hits every gene, so neighbors lose that feature too.

4 Results

4.1 Model performance

Each of the 25 test folds has its own AUPRC on the held-out genes. The comparison is the mean ± sd of those 25 numbers. Stars after each number are Holm-adjusted Wilcoxon tests against Open Targets.

Horizontal pill bars of mean AUPRC across 25 test folds. A white marker sits at the mean; a rounded whisker is plus or minus one standard deviation. Holm stars sit after each mean plus or minus sd: STRING-GCN and XGBoost are n.s., STRING-GAT is one star, the remaining four are three stars. A dashed line marks chance at 0.032. A legend defines n.s., one star, two stars, and three stars.

AUPRC across 25 test folds, with Holm-adjusted Wilcoxon vs Open Targets

Horizontal pill bars of mean AUPRC across 25 test folds. A light marker sits at the mean; a rounded whisker is plus or minus one standard deviation. Holm stars sit after each mean plus or minus sd: STRING-GCN and XGBoost are n.s., STRING-GAT is one star, the remaining four are three stars. A dashed line marks chance at 0.032. A legend defines n.s., one star, two stars, and three stars.

AUPRC across 25 test folds, with Holm-adjusted Wilcoxon vs Open Targets

STRING-GCN was the trained winner. It beat GCN without edges in 25/25 folds (mean AUPRC lift about 0.22). The protein knowledge graph added information: the same architecture, without STRING edges, fell near 0.10. GCN on degree only stayed near 0.09, so the lift is neighborhood mixing of the features, not a ranking of hubs.

It tied XGBoost (GCN ahead on 16/25 folds, overlapping SDs). Trees already see precomputed STRING neighborhood features; the GCN still gained by mixing those features along the actual edges. The tabular neural net lagged both (0.165): a small net overfits this 3% labeled matrix more than depth-4 trees.

GAT lost to GCN on 21/25 folds. The remaining information is homophilous STRING mixing: a gene looks like its neighbors. A degree-normalized average matches that bias. With ~111 labeled genes per train fold and 82k edges, GCN was the mixer that fit. GAT still beat GCN without edges, so attention was running; it was the worse inductive bias here.

STRING-GCN matched the positive control on AUPRC (n.s., Δ -0.008). XGBoost did too (n.s., Δ -0.025). STRING-GAT sat below Open Targets (0.274, Holm p < 0.05). The remaining controls sat well below (Holm p < 0.001).

Shuffled labels is the sanity check: same XGBoost, labels permuted. AUPRC fell to 0.038, which is chance (0.032). The metric is not inventing a ranking from the 3% prevalence.

What STRING-GCN recovered at 0.322 is therefore a real ranking of known clinical-stage mechanisms, from the 18 features that omit the clinical evidence. Open Targets overall is close (0.329) because that public score includes that evidence.

4.2 Feature importance

I ranked the 18 features by mean AUPRC drop in the winning STRING-GCN: how much held-out AUPRC falls after shuffling that feature on every gene (neighbors lose it too).

Horizontal pill bars of mean AUPRC drop after shuffling each of the 18 features in the trained STRING-GCN. Genetics-linked gene modules lead at 0.052, then matching transcription factors at 0.050 and lung-cell essentiality at 0.045.

STRING-GCN feature importance as mean AUPRC drop

Horizontal pill bars of mean AUPRC drop after shuffling each of the 18 features in the trained STRING-GCN. Genetics-linked gene modules lead at 0.052, then matching transcription factors at 0.050 and lung-cell essentiality at 0.045.

STRING-GCN feature importance as mean AUPRC drop

Genetics-linked gene modules, matching transcription factors, and lung-cell essentiality are the three features the trained GCN leans on (drops 0.052, 0.050, 0.045). STRING neighborhood features come next: IPF-changed genes as neighbors, community, the network walk, and module centrality. GEO lung expression, GWAS flags, animal models, and Open Targets RNA sit at the bottom (drops near 0).

4.3 Target ranking

I looked at out-of-fold STRING-GCN predictions for illustration (each gene’s prediction is the mean of the folds that held it out). Known means the gene is one of the 139 clinical-stage mechanisms I used as labels (phase ≥ 1). Novel means it is not: a hypothesis on this freeze, not a recovered training example.

Gene Known or novel STRING-GCN Note
TGFB1 Known 0.61 Canonical fibrosis biology [6]
TNF Known 0.71 Neighborhood lifts a known mechanism
PDGFRB Known 0.58 Nintedanib-family target [5]
MUC5B Novel 0.33 Famous IPF risk gene; stays mid-rank [22]
COX7B Novel 0.77 Top novel hit: respiratory-chain hub

The model recovered known clinical mechanisms in the TGF-β, TNF, and PDGFR families without putting clinical association into the features. TNF was the clearest neighborhood lift: XGBoost predicted 0.47; the GCN predicted 0.71.

MUC5B is the genetic signature of IPF risk and is Novel in this table: it is not a clinical-stage mechanism under the definition of the labels. Mid-rank is the honest outcome: a risk gene is not the same decision as a drug mechanism.

Those high novel ranks (COX7B, other respiratory-chain subunits, HRAS, MAPK1) are the genes I would prioritize for a chemist or a biologist. They are shaped by the same STRING geometry that recovered the known targets.

5 Conclusions

I started this freeze with a question from target identification: on public IPF data, does a protein knowledge graph earn its keep, or do trees on the same 18 features already do the job? On the confirmation metric, mean AUPRC across 25 test folds, the graph earned it. STRING-GCN recovered known clinical-stage mechanisms at 0.322. XGBoost on the same features reached 0.305. The same GCN with the STRING edges deleted fell to 0.098, in every fold. The lift is neighborhood mixing of the features, not a more fashionable architecture and not a ranking of hubs.

That trained ranking sat next to Open Targets overall at 0.329. Open Targets is the positive control: it already combines every evidence type in that database, including the clinical evidence that defines the labels, which I kept out of the 18 features. I am not offering a better Open Targets. I am offering a GCN that, from genetics, tissue, pathways, perturbation, and STRING, recovers a neighborhood-shaped trace of the same mechanisms, close enough that a Wilcoxon test cannot tell the two lists apart.

What the GCN used matches that picture. It leans on genetics-linked modules, transcription factors whose targets match the IPF signature, and lung-cell essentiality, then on STRING neighborhood. It does not lean on a GEO volcano plot. Known mechanisms in the TGF-β, TNF, and PDGFR families come back without a known-drug column. High novel ranks such as COX7B are the next conversation with a chemist: hypotheses shaped by the same geometry, not recovered labels. MUC5B staying mid-rank is the other honest outcome. A famous risk gene is not a drug mechanism.

This freeze is still a laptop experiment: public features, one disease, a known-drug proxy as the labels, a transductive GNN. DepMap is essentiality in lung cancer lines, not IPF tissue. If I ran the same question inside a company I would train on older evidence and test on clinical-stage genes that only appeared later, use an indication-specific graph, and add chemist labels.

6 References

1.
Ganesh Raghu, Harold R Collard, Jim J Egan, Fernando J Martinez, Jürgen Behr, Kevin K Brown, Thomas V Colby, Jean-François Cordier, Kevin R Flaherty, Joseph A Lasky, David A Lynch, Jay H Ryu, Jeffrey J Swigris, Athol U Wells, Julio Ancochea, Demosthenes Bouros, Carlos Carvalho, Ulrich Costabel, Masahito Ebina, David M Hansell, Takeshi Johkoh, Dong Soon Kim, Talmadge E King, Yasuhiro Kondoh, Jeffrey Myers, Nestor L Müller, Andrew G Nicholson, Luca Richeldi, Moisés Selman, Rosalind F Dudden, Barbara S Griss, Shandra L Protzko, Holger J Schünemann. An official ATS/ERS/JRS/ALAT statement: Idiopathic pulmonary fibrosis: Evidence-based guidelines for diagnosis and management. American Journal of Respiratory and Critical Care Medicine. 2011;183(6):788–824. doi:10.1164/rccm.2009-040GL
2.
Brett Ley, Harold R Collard, Talmadge E King. Clinical course and prediction of survival in idiopathic pulmonary fibrosis. American Journal of Respiratory and Critical Care Medicine. 2011;183(4):431–40. doi:10.1164/rccm.201006-0894CI
3.
Joseph Yang, Rebecca Fee, Andrea Steffens, Lisa Le, Bonnie H Bui, Bijan J Borah. Real-world health care resource utilization and economic burden among patients with idiopathic pulmonary fibrosis in commercially insured and Medicare Advantage populations in the United States. Journal of Health Economics and Outcomes Research. 2026;13(1):200–9. doi:10.36469/001c.161503
4.
Talmadge E King, Williamson Z Bradford, Socorro Castro-Bernardini, Elizabeth A Fagan, Ian Glaspole, Marilyn K Glassberg, Eduard Gorina, Peter M Hopkins, David Kardatzke, Lisa Lancaster, David J Lederer, Steven D Nathan, Carlos A Pereira, Steven A Sahn, Robert Sussman, Jeffrey J Swigris, Paul W Noble. A phase 3 trial of pirfenidone in patients with idiopathic pulmonary fibrosis. The New England Journal of Medicine. 2014;370(22):2083–92. doi:10.1056/NEJMoa1402582
5.
Luca Richeldi, Roland M du Bois, Ganesh Raghu, Arata Azuma, Kevin K Brown, Ulrich Costabel, Vincent Cottin, Kevin R Flaherty, David M Hansell, Yoshikazu Inoue, Dong Soon Kim, Martin Kolb, Andrew G Nicholson, Paul W Noble, Moisés Selman, Hiroyuki Taniguchi, Michèle Brun, Florence Le Maulf, Mannaïg Girard, Susanne Stowasser, Rozsa Schlenker-Herceg, Bernd Disse, Harold R Collard. Efficacy and safety of nintedanib in idiopathic pulmonary fibrosis. The New England Journal of Medicine. 2014;370(22):2071–82. doi:10.1056/NEJMoa1402584
6.
Isis E Fernandez, Oliver Eickelberg. The impact of TGF-β on lung fibrosis: From targeting to biomarkers. Proceedings of the American Thoracic Society. 2012;9(3):111–6. doi:10.1513/pats.201203-023AW
7.
Thomas N Kipf, Max Welling. Semi-supervised classification with graph convolutional networks. In: International conference on learning representations [Internet]. 2017. Available from: https://openreview.net/forum?id=SJU4ayYgl
8.
David Ochoa, Andrew Hercules, Miguel Carmona, Daniel Suveges, Jarrod Baker, Cinzia Malangone, Irene Lopez, Alfredo Miranda, Carlos Cruz-Castillo, Luca Fumis, Manuel Bernal Llinares, Kirill Tsukanov, Helena Cornu, Konstantinos D Tsirigos, Olesya Razuvayevskaya, Annalisa Buniello, Jeremy Schwartzentruber, Mohd Anisul Karim, Bruno Ariano, Ricardo Esteban Martinez Osorio, Javier Ferrer, Xiangyu Ge, Sandra Machlitt-Northen, Asier Gonzalez-Uriarte, Shyamasree Saha, Santosh Tirunagari, Chintan Mehta, Juan María Roldán-Romero, Stuart Horswell, Sarah Young, Maya Ghoussaini, David G Hulcoop, Ian Dunham, Ellen M McDonagh. The next-generation Open Targets Platform: Reimagined, redesigned, rebuilt. Nucleic Acids Research. 2023;51(D1):D1353–9. doi:10.1093/nar/gkac1046
9.
Damian Szklarczyk, Rebecca Kirsch, Mikaela Koutrouli, Katerina Nastou, Farrokh Mehryary, Radja Hachilif, Annika L Gable, Tao Fang, Nadezhda T Doncheva, Sampo Pyysalo, Peer Bork, Lars J Jensen, Christian von Mering. The STRING database in 2023: Protein–protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Research. 2023;51(D1):D638–46. doi:10.1093/nar/gkac1000
10.
Elliot Sollis, Abayomi Mosaku, Ala Abid, Annalisa Buniello, Maria Cerezo, Laurent Gil, Tudor Groza, Osman Güneş, Peggy Hall, James Hayhurst, Arwa Ibrahim, Yue Ji, Sajo John, Elizabeth Lewis, Jacqueline A L MacArthur, Aoife McMahon, David Osumi-Sutherland, Kalliope Panoutsopoulou, Zoë Pendlington, Santhi Ramachandran, Ray Stefancsik, Jonathan Stewart, Patricia Whetzel, Robert Wilson, Lucia Hindorff, Fiona Cunningham, Samuel A Lambert, Michael Inouye, Helen Parkinson, Laura W Harris. The NHGRI-EBI GWAS catalog: Knowledgebase and deposition resource. Nucleic Acids Research. 2023;51(D1):D977–85. doi:10.1093/nar/gkac1010
11.
Tanya Barrett, Stephen E Wilhite, Pierre Ledoux, Carlos Evangelista, Irene F Kim, Maxim Tomashevsky, Kimberly A Marshall, Katherine H Phillippy, Patti M Sherman, Michelle Holko, Andrey Yefanov, Hyeseung Lee, Naigong Zhang, Cynthia L Robertson, Nadezhda Serova, Sean Davis, Alexandra Soboleva. NCBI GEO: Archive for functional genomics data sets—update. Nucleic Acids Research. 2013;41(D1):D991–5. doi:10.1093/nar/gks1193
12.
GTEx Consortium. The GTEx consortium atlas of genetic regulatory effects across human tissues. Science. 2020;369(6509):1318–30. doi:10.1126/science.aaz1776
13.
Mathias Uhlén, Linn Fagerberg, Björn M Hallström, Cecilia Lindskog, Per Oksvold, Adil Mardinoglu, Åsa Sivertsson, Caroline Kampf, Evelina Sjöstedt, Anna Asplund, IngMarie Olsson, Karolina Edlund, Emma Lundberg, Sanjay Navani, Cristina Al-Khalili Szigyarto, Jacob Odeberg, Dijana Djureinovic, Jenny Ottosson Takanen, Sophia Hober, Tove Alm, Per-Henrik Edqvist, Holger Berling, Hanna Tegel, Jan Mulder, Johan Rockberg, Peter Nilsson, Jochen M Schwenk, Marica Hamsten, Kalle von Feilitzen, Mattias Forsberg, Lukas Persson, Fredric Johansson, Martin Zwahlen, Gunnar von Heijne, Jens Nielsen, Fredrik Pontén. Tissue-based map of the human proteome. Science. 2015;347(6220):1260419. doi:10.1126/science.1260419
14.
Marc Gillespie, Bijay Jassal, Ralf Stephan, Marija Milacic, Karen Rothfels, Andrea Senff-Ribeiro, Johannes Griss, Cristoffer Sevilla, Lisa Matthews, Chuqiao Gong, Chuan Deng, Thawfeek Varusai, Eliot Ragueneau, Yusra Haider, Bruce May, Veronica Shamovsky, Joel Weiser, Timothy Brunson, Nasim Sanati, Liam Beckman, Xiang Shao, Antonio Fabregat, Konstantinos Sidiropoulos, Julieth Murillo, Guilherme Viteri, Justin Cook, Solomon Shorser, Gary Bader, Emek Demir, Chris Sander, Robin Haw, Guanming Wu, Lincoln Stein, Henning Hermjakob, Peter D’Eustachio. The reactome pathway knowledgebase 2022. Nucleic Acids Research. 2022;50(D1):D687–92. doi:10.1093/nar/gkab1028
15.
Sophia Müller-Dott, Eirini Tsirvouli, Miguel Vazquez, Ricardo O Ramirez Flores, Pau Badia-i-Mompel, Robin Fallegger, Dénes Türei, Astrid Lægreid, Julio Saez-Rodriguez. Expanding the coverage of regulons from high-confidence prior knowledge for accurate estimation of transcription factor activities. Nucleic Acids Research. 2023;51(20):10934–49. doi:10.1093/nar/gkad841
16.
Aravind Subramanian, Rajiv Narayan, Steven M Corsello, David D Peck, Ted E Natoli, Xiaodong Lu, Joshua Gould, John F Davis, Andrew A Tubelli, Jacob K Asiedu, David L Lahr, Jodi E Hirschman, Zihan Liu, Melanie Donahue, Bina Julian, Mariya Khan, David Wadden, Ian C Smith, Daniel Lam, Arthur Liberzon, Courtney Toder, Mukta Bagul, Marek Orzechowski, Oana M Enache, Federica Piccioni, Sarah A Johnson, Nicholas J Lyons, Alice H Berger, Alykhan F Shamji, Angela N Brooks, Anita Vrcic, Corey Flynn, Jacqueline Rosains, David Y Takeda, Roger Hu, Desiree Davison, Justin Lamb, Kristin Ardlie, Larson Hogstrom, Peyton Greenside, Nathanael S Gray, Paul A Clemons, Serena Silver, Xiaoyun Wu, Wen-Ning Zhao, Willis Read-Button, Xiaohua Wu, Stephen J Haggarty, Lucienne V Ronco, Jesse S Boehm, Stuart L Schreiber, John G Doench, Joshua A Bittker, David E Root, Bang Wong, Todd R Golub. A next generation connectivity map: L1000 platform and the first 1,000,000 profiles. Cell. 2017;171(6):1437–52. doi:10.1016/j.cell.2017.10.049
17.
Aviad Tsherniak, Francisca Vazquez, Phil G Montgomery, Barbara A Weir, Gregory Kryukov, Glenn S Cowley, Stanley Gill, William F Harrington, Sasha Pantel, John M Krill-Burger, Robin M Meyers, Levi Ali, Amy Goodale, Yenarae Lee, Guozhi Jiang, Jessica Hsiao, William F J Gerath, Sara Howell, Erin Merkel, Mahmoud Ghandi, Levi A Garraway, David E Root, Todd R Golub, Jesse S Boehm, William C Hahn. Defining a cancer dependency map. Cell. 2017;170(3):564–76. doi:10.1016/j.cell.2017.06.010
18.
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, Yoshua Bengio. Graph attention networks. In: International conference on learning representations [Internet]. 2018. Available from: https://openreview.net/forum?id=rJXMpikCZ
19.
Takaya Saito, Marc Rehmsmeier. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLOS ONE. 2015;10(3):e0118432. doi:10.1371/journal.pone.0118432
20.
Frank Wilcoxon. Individual comparisons by ranking methods. Biometrics Bulletin. 1945;1(6):80–3. doi:10.2307/3001968
21.
Sture Holm. A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics. 1979;6(2):65–70.
22.
Max A Seibold, Anastasia L Wise, Marcy C Speer, Mark P Steele, Kevin K Brown, James E Loyd, Tasha E Fingerlin, Weimin Zhang, Gunnar Gudmundsson, Steve D Groshong, Christopher M Evans, Stavros Garantziotis, Kenneth B Adler, Burton F Dickey, Roland M du Bois, Ivana V Yang, Ahcene Herron, Dolly Kervitsky, Janet L Talbert, Cheryl Markin, Junesoo Park, Anne L Crews, Susan H Slifer, Scott Auerbach, Michelle G Roy, Jia Lin, Craig E Hennessy, Marvin I Schwarz, David A Schwartz. A common MUC5B promoter polymorphism and pulmonary fibrosis. The New England Journal of Medicine. 2011;364(16):1503–12. doi:10.1056/NEJMoa1013660
Back to top