Evidence
Same model and position · retrieved passages supplied · first output
These limitations are not unique to retrosynthesis prediction; benchmarks of temporal distribution shift in other domains likewise report substantial performance drops when evaluation moves from in-distribution to out-of-distribution data, underscoring the need for evaluation protocols that reflect realistic deployment conditions.
Passages supplied to the Evidence version
Wild-Time: A Benchmark of in-the-Wild Distribution Shift over Time
To address this gap, we curate Wild-Time, a benchmark of 5 datasets that reflect temporal distribution shifts arising in a variety of real-world applications, including patient prognosis and news classification. On these datasets, we systematically benchmark 13 prior approaches, including methods in…
Read full passage excerpt
To address this gap, we curate Wild-Time, a benchmark of 5 datasets that reflect temporal distribution shifts arising in a variety of real-world applications, including patient prognosis and news classification. On these datasets, we systematically benchmark 13 prior approaches, including methods in domain generalization, continual learning, self-supervised learning, and ensemble learning. We use two evaluation strategies: evaluation with a fixed time split (Eval-Fix) and evaluation with a data stream (Eval-Stream). Eval-Fix, our primary evaluation strategy, aims to provide a simple evaluation protocol, while Eval-Stream is more realistic for certain real-world applications. Under both evaluation strategies, we observe an average performance drop of 20% from in-distribution to out-of-distribution data.