Evaluation Choices in Gene Regulatory Network Inference: An Empirical Analysis on DREAM4

This article has 0 evaluations Published on
Read the full article Related papers
This article on Sciety

Abstract

Claims that one gene regulatory network (GRN) inference method outperforms another depend on both the data and the evaluation protocol. We examine this dependence in a small, controlled benchmark using the five 100-gene networks from the DREAM4 In Silico Network Challenge. Four transparent estimators—absolute Pearson correlation, shrinkage partial correlation, histogram mutual information, and target-wise Extra Trees—were fitted to 25, 50, or 100 multifactorial expression profiles. The same predictions were evaluated over either all 9,900 directed non-self edges or an oracle candidate set restricted to genes with at least one true outgoing edge. Across ten reproducible subsampling runs, performance at 100 profiles remained near the random-prevalence baseline: median ratios of area under the precision– recall curve (AUPRC) to edge prevalence ranged from 0.91 to 1.05 under the all-edge protocol. Restricting candidates increased raw AUPRC because prevalence rose, but did not improve prevalence-normalized AUPRC. More importantly, the restriction reversed 95 of 900 matched pairwise method comparisons (10.6%); pair-specific reversal fractions ranged from 3.3% to 18.0%. Changing sample size reversed 26.3–32.3% of aggregated pairwise comparisons. These results do not rank modern GRN methods. They show, within a deliberately narrow low-signal setting, that relative order is fragile and that candidate-universe and prevalence choices must be reported alongside benchmark scores.

Related articles

Related articles are currently not available for this article.