Evidence
同一模型与写作位置 · 提供检索文献片段 · 首次输出
These limitations are not unique to retrosynthesis prediction; benchmarks of temporal distribution shift in other domains likewise report substantial performance drops when evaluation moves from in-distribution to out-of-distribution data, underscoring the need for evaluation protocols that reflect realistic deployment conditions.
提供给 Evidence 版本的文献片段
Wild-Time: A Benchmark of in-the-Wild Distribution Shift over Time
To address this gap, we curate Wild-Time, a benchmark of 5 datasets that reflect temporal distribution shifts arising in a variety of real-world applications, including patient prognosis and news classification. On these datasets, we systematically benchmark 13 prior approaches, including methods in…
展开完整文献摘录
To address this gap, we curate Wild-Time, a benchmark of 5 datasets that reflect temporal distribution shifts arising in a variety of real-world applications, including patient prognosis and news classification. On these datasets, we systematically benchmark 13 prior approaches, including methods in domain generalization, continual learning, self-supervised learning, and ensemble learning. We use two evaluation strategies: evaluation with a fixed time split (Eval-Fix) and evaluation with a data stream (Eval-Stream). Eval-Fix, our primary evaluation strategy, aims to provide a simple evaluation protocol, while Eval-Stream is more realistic for certain real-world applications. Under both evaluation strategies, we observe an average performance drop of 20% from in-distribution to out-of-distribution data.