WeatherBench2

Model Introduction

WeatherBench2 is an evaluation benchmark for the next generation of data-driven global weather models. It covers deterministic, ensemble-probabilistic, bias, and spectral diagnostics.

Paper: WeatherBench 2: A Benchmark for the Next Generation of Data-Driven Global Weather Models
https://arxiv.org/abs/2308.15560

Model Description

The benchmark was proposed by teams from Google Research, Google DeepMind, and ECMWF. It uses 2020 global forecasts from ERA5, IFS, and multiple data-driven systems. It supports deterministic, probabilistic, bias, and spatial-scale evaluation of global forecasts from one to fourteen days.

Use Cases

Use Case Description
Deterministic evaluation Compute RMSE, ACC, bias, and SEEPS.
Probabilistic evaluation Compute CRPS and spread-skill ratio.
Ensemble diagnosis Compare ensemble means, spread, and skill.
ModelScope/OneCode execution Validate data, training, inference, evaluation, and visualization.
Multi-GPU training Validate a compact baseline through torchrun.

Usage Instructions

hf download OneScience-Group/WeatherBench2 --local-dir ./WeatherBench2
cd WeatherBench2

Environment Dependencies

Hardware Requirements

  • A GPU or DCU is recommended.
  • A CPU can be used for connectivity validation with the default small-sample configuration.
  • DCU users should install DTK 25.04.2 or a compatible OneScience-recommended version first.

DCU Environment

# Activate DTK and Conda first
conda create -n onescience311 python=3.11 -y
conda activate onescience311
pip install onescience[earth-dcu] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai

GPU Environment

# Activate Conda first
conda create -n onescience311 python=3.11 -y libstdcxx-ng=12 libgcc-ng=12 gcc_linux-64=12 gxx_linux-64=12
conda activate onescience311
pip install onescience[earth-gpu] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai

Training Data

WeatherBench2 evaluates 2020 global forecasts from ERA5, IFS, and data-driven systems on a common 1.5-degree grid. Synthetic data retain eight headline variables and the ensemble dimension while reducing times and grid size.

python scripts/fake_data.py

Training

For single-process training, use:

python scripts/train.py

For multi-process training, use:

torchrun --standalone --nproc_per_node=2 scripts/train.py

Training results are saved to:

result/checkpoints/weatherbench2.pt
result/training/metrics.json

Trained Weights

No weights are bundled under weight/. WeatherBench2 is a benchmark rather than a single pretrained model, so there is no unified official weight artifact.

Inference

python scripts/inference.py

Inference generates an ensemble with shape [8,8,8,24,48]. Results are saved to:

result/output/predictions.npz

Evaluation and Visualization

python scripts/result.py

Evaluation reports RMSE, CRPS, and spread-skill ratio and creates a spatial error map. Results are saved to:

result/evaluation/metrics.json
result/evaluation/comparison.png

Official OneScience Information

Citation and License

This repository is an independent engineering reproduction of the public WeatherBench2 specifications.

The WeatherBench2 evaluation code, ERA5, IFS, and forecast data from participating systems remain subject to the licenses and data-use terms of their respective source projects.

Downloads last month
15
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for OneScience-Group/WeatherBench2