Using Generative AI to Measure Uncertainty in Weather Predictions
In December 1972, at the American Association for the Advancement of Science meeting in Washington, D.C., MIT meteorology professor Ed Lorenz presented a talk titled, “Does the Flap of a Butterfly’s Wings in Brazil Set Off a Tornado in Texas?”, which played a significant role in coining the term “butterfly effect.” This discussion built upon his earlier, groundbreaking 1963 paper, where he explored the potential for “very-long-range weather prediction” and illustrated how errors in initial conditions can grow exponentially over time when integrated using numerical weather prediction models. This exponential error growth, referred to as chaos, leads to a deterministic predictability limit that hampers the effectiveness of individual forecasts in decision-making, as these forecasts do not adequately convey the inherent uncertainty of weather conditions. This challenge is especially pronounced when predicting extreme weather events like hurricanes, heatwaves, or floods.
To address the limitations of deterministic forecasts, weather agencies globally have begun to issue probabilistic forecasts. These forecasts are derived from ensembles of deterministic predictions, where each forecast includes synthetic noise in the initial conditions and incorporates stochastic behavior in the physical processes. By leveraging the rapid error growth inherent in weather models, the forecasts within an ensemble are intentionally diverse: initial uncertainties are calibrated to produce variations that are as distinct as possible, while stochastic processes in the weather model introduce additional differences during the simulation. The growth of errors is lessened through averaging the forecasts in the ensemble, and the range of variability among these forecasts effectively quantifies the uncertainty of weather conditions.
While reliable, the generation of these probabilistic forecasts is computationally demanding. It requires the execution of sophisticated numerical weather models on substantial supercomputers multiple times. As a result, many operational weather forecasts can only afford to produce approximately 10–50 ensemble members for each forecast cycle. This limitation poses a challenge for users focused on the likelihood of rare but impactful weather events, which generally necessitate significantly larger ensembles to examine probabilities accurately over longer time frames. For example, to project the likelihood of events with just a 1% probability of occurrence and a relative error of less than 10%, a 10,000-member ensemble would be essential. Understanding the probability of such extreme events can be critical for areas such as emergency management preparation or for sectors involved in energy trading.
