Every factor gets a verdict.
Run backtests for realized performance or launch the full validation battery—gate, walk-forward, Monte Carlo, regime simulation, and deflated Sharpe—from one place.
Loading…
Every authored version change.
Each entry is an immutable, authored change to a strategy definition—what changed, why, what was expected, and whether the test result accepted or rejected it.
Loading lineage…
What earns trust—and what breaks it.
No single metric promotes a strategy. The lab asks whether skill exists, survives new data and changing regimes, controls downside, and rests on enough evidence.
IC / ICIR
Information coefficient (IC) measures whether a factor's rankings predict later returns. ICIR compares average IC with its variation, showing whether that skill is stable rather than occasional.
OOS Sharpe / Sortino / Calmar
Out-of-sample Sharpe weighs return against all volatility; Sortino focuses on harmful downside volatility; Calmar compares return with maximum drawdown. Together they show different forms of risk-adjusted performance on held-back data.
Maximum drawdown
The worst peak-to-trough loss in the tested equity curve. It answers how deep the strategy fell before recovering or the test ended.
Deflated Sharpe
A multiple-testing-corrected check on whether the observed Sharpe is distinguishable from luck after accounting for how many ideas were tried.
Risk of ruin
The simulated chance of breaching the configured loss boundary. The verdict uses the less flattering result from plain and regime-switching Monte Carlo tests.
Walk-forward verdict
STABLE means the edge held across rolling out-of-sample windows. WEAK means no reliable edge was found, but there was no strong in-sample edge to expose as overfit. OVERFIT means apparent in-sample skill failed to replicate out of sample.
Trial count
The number of historical ideas tested. More trials create more chances for a lucky result, so uncorrected performance becomes less convincing as this count rises.
A–F grade tiers
A is strongest, B clears the required evidence bar, C is mixed or weak, D is materially below the bar, and F is failed or insufficient evidence. The overall grade is the WORST of edge, robustness, risk, and sample size—not an average—so one strength cannot hide a fatal weakness.