Abstract
Reliable subseasonal-to-seasonal (S2S) precipitation forecasts during the West African monsoon are critical for Senegal, where the economy heavily depends on rain-fed agriculture. However, general circulation models used in current S2S forecasting systems struggle to represent the complex atmospheric and oceanic mechanisms driving monsoon rainfall variability. This study evaluates six machine learning models, including ridge regression, linear regression, random forest, support vector machine, AdaBoost, and multilayer perceptron, for forecasting weekly precipitation anomalies during the monsoon season (June to September, 1982 to 2019). We combine high-resolution precipitation estimates from ground and satellite observations with atmospheric and oceanic reanalysis products, and apply a non-filtering method to extract intraseasonal signals as predictors, enabling real-time applicability. Results show that integrating atmospheric and oceanic predictors significantly enhances forecast skill compared to using individual variables. Ridge regression consistently outperforms all other models and surpasses state-of-the-art dynamical S2S prediction systems across all forecast lead times. These findings highlight the strong potential of machine learning to complement dynamical models for operational S2S precipitation forecasting in West Africa, offering computationally efficient and skillful predictions valuable for climate risk anticipation, agricultural planning, and water resource management. The proposed methodology can be extended to other regions with different climatic characteristics.
1 Introduction
Extreme floods and droughts are becoming more widespread in the current context of climate change, causing considerable economic damage and threatening livelihoods, particularly in vulnerable regions such as Africa, according to Chapter 9 (Africa) of the IPCC Sixth Assessment Report (Masson-Delmotte et al., 2021; Trisos et al., 2022). For West Africa in general, and Senegal in particular, reliable subseasonal-to-seasonal (S2S) forecasts of rainfall during the West African monsoon (WAM) season can provide actionable data for informed decision-making to mitigate the harmful impact of extreme events on the vulnerable population, whose economy and subsistence are mainly based on rain-fed agriculture. Indeed, S2S predictability of rainfall in this region is crucial for crop management, disaster risk reduction and food security. This S2S forecasting time scale, which typically ranges from two weeks to two months, helps bridge the gap between weather forecasting and climate prediction (Hung et al., 2009; Brunet et al., 2010; Vitart et al., 2012). However, rainfall variability associated with the WAM is highly complex due to the combined impact of local mesoscale convective systems and remote sea surface temperature (SST) forcing (Suárez-Moreno et al., 2018; Biasutti, 2019), which contributes to preventing skillful S2S predictability (Vigaud and Giannini, 2019). From a societal perspective, decision-making in the framework of agriculture, food security, water resources, risk management and health care would greatly benefit from improved S2S forecasts. However, this time scale has long been considered a “predictability desert,” being much less studied compared to medium-range and seasonal forecasting (Robertson et al., 2020). Recent studies have shown that predictability at the S2S time scale could be enhanced through various factors, including more complete and reliable observational networks, a better understanding and representation of intra-seasonal atmospheric processes, and improved initialization and assimilation of land, ocean, cryosphere and stratosphere components in general circulation models (GCMs)-based prediction systems (Pegion et al., 2019). Accordingly, several initiatives have been carried out, such as the S2S prediction project and the Subseasonal eXperiment (SubX; now the Subseasonal Consortium, SubC), to provide GCM-derived S2S rainfall forecasts up to 60 days in advance (Vitart et al., 2017; Pegion et al., 2019). However, such forecasts are still far from skillful (de Andrade et al., 2019), as GCMs struggle to represent key atmospheric processes, mainly those on a small scale (Vitart and Robertson, 2019). As a result, post-processing is often required to improve the accuracy and reliability of forecasts (Li et al., 2022). Some studies used, for instance, Bayesian joint probability to post-process S2S rainfall forecasts over different regions, resulting in improved skill and reliability compared to the raw forecasts (Schepen et al., 2012; Li et al., 2021). Moreover, a new spatial correction method proposed by Vigaud et al. (2020) showed improved intra-seasonal rainfall estimates derived from multimodel ensembles. Nevertheless, the accuracy of these post-processed estimates degrades sharply for lead times (i.e., the difference in time between the onset of an observed event and the issuance of the forecast of such event) beyond 10–14 days (Li et al., 2022).
Alternative approaches for S2S rainfall forecasting include machine learning (ML) models that operate from mathematical relationships between rainfall data and preceding indices of atmospheric and/or oceanic variability. Although GCMs have traditionally excelled in short- and medium-term forecasting, and ML models have historically been reserved for longer timescales (Abbot and Marohasy, 2014; Tuel and Eltahir, 2018), recent advancements in ML, such as with models like GraphCast, have demonstrated unprecedented accuracy in medium-term weather forecasting. Many data-driven algorithms, such as multiple linear regression, maximum covariance analysis or canonical correlation analysis, among others, have been devised for seasonal rainfall prediction based on the assumption that seasonal anomalies are caused by slow variations in SST, snow cover and other boundary conditions (Barnston and Smith, 1996; Hwang et al., 2001; Eden et al., 2015; Suárez-Moreno and Rodrı́guez-Fonseca, 2015). Additional examples include an empirical cluster-based method to predict winter precipitation anomalies over European and Mediterranean regions using SST, geopotential height, sea level pressure, snow cover extent, and sea ice concentration as predictors (Totz et al., 2017). A random forest-based model, whose predictors were extracted from SST data, was developed to predict seasonal precipitation in Central and Southern Asia (Gerlitz et al., 2016). However, there is a lack of studies focusing on intra-seasonal (or S2S) forecasting in West Africa. Despite these advances, ML-based subseasonal rainfall forecasting remains largely unexplored in West Africa. Few studies have evaluated the skill of general circulation models (GCMs) in predicting sub-seasonal precipitation over this region, and forecast quality generally decreases beyond two weeks (de Andrade et al., 2021). Despite recent advances in machine learning for weather and climate prediction, ML-based subseasonal rainfall forecasting remains largely unexplored in West Africa. In particular, few studies have investigated the combined contribution of atmospheric and oceanic intraseasonal predictors for weekly rainfall prediction over Senegal. Furthermore, many existing approaches rely on filtering techniques to isolate intraseasonal variability, which may introduce information from future observations and limit their suitability for real-time forecasting applications. To address these gaps, this study develops a machine learning framework based on intraseasonal atmospheric and oceanic predictors extracted using a non-filtering methodology that can be implemented operationally. The proposed framework is also benchmarked against leading operational dynamical S2S forecasting systems (ECMWF, UKMO, and NCEP), providing a comprehensive assessment of its added value for subseasonal rainfall prediction over Senegal.
This study takes a step forward in developing an ML-based intra-seasonal rainfall forecasting system for Senegal. By exploiting the statistical links between rainfall variability and large-scale atmospheric and oceanic fields, we evaluate the performance of six ML approaches in predicting weekly precipitation anomalies during the WAM season. The remainder of this paper is organized as follows. Section 2 describes the data and methodology, including the non-filtering extraction of intraseasonal predictors and the ML algorithms employed. Section 3 presents the forecasting capabilities of the models, evaluated through cross-validation and comparison with dynamical S2S systems. Section 4 discusses the physical interpretability of the predictor hierarchy, methodological limitations, and implications for operational forecasting, before concluding in Section 5.
2 Data and methodology
2.1 Data
2.1.1 Geographical information
Senegal is located in the westernmost part of West Africa, between latitudes 12°00’N and 17°00’N and longitudes 11°00’W and 18°00’W (red box on Figure 1). It is characterized by two main seasons: a dry season (from November to May) marked by the predominance of maritime trade winds blowing from the north to north-west along the coast, and the continental harmattan blowing from the northeast in the interior; and a rainy season (from May–June to October), dominated by the southwesterly flow that characterizes the WAM (Fall et al., 2020). The maximum rainfall is received in August and September, coinciding with the period during which the intertropical convergence zone (ITCZ) reaches its northernmost position above Senegal (Sane et al., 2018). Figure 1 illustrates the spatial distribution of the annual cumulative precipitation in West Africa, along with the area of interest. Precipitation exhibits a distinct meridional gradient, characterized by higher values over southern areas and a northward decrease (Faye et al., 2024).
2.1.2 Target variable: precipitation data
Daily precipitation data are derived from the Climate Hazards group InfraRed Precipitation with Stations (CHIRPS) dataset (Funk et al., 2015). CHIRPS combines ground observation data and satellite observations to produce daily and monthly gridded rainfall estimates worldwide from 1981 to the present. CHIRPS has a high spatial resolution (5 Km), exhibiting more realistic spatial patterns and greater accuracy over land than other observational datasets. A study by Tarnavsky et al. (2014) compared CHIRPS with Tropical Rainfall Measuring Mission (TRMM) data (Simpson et al., 1988) for the West African region. The results showed that the two datasets are well correlated, but CHIRPS tends to better capture short-duration rainfall events. Likewise, comparing with the European Centre for Medium-Range Weather Forecasts (ECMWF) dataset, Funk et al. (2015) showed that CHIRPS is more realistic in reproducing lower rainfall intensities. Recently, Faye et al. (2024) showed that CHIRPS data, compared to those from the National Hydrological and Meteorological Services (NHMS), provide a more accurate representation of the total annual rainfall distribution, as well as the onset and cessation of the rainy season in Senegal.
2.1.3 Atmospheric and oceanic variables
To represent the intraseasonal atmospheric oscillation, outgoing longwave radiation (OLR) and zonal winds (U) in the upper (U200 hPa) and lower (U850 hPa) troposphere are used. Although several indices, including the real-time multivariate Madden-Julian Oscillation (MJO) index (RMM) (Wheeler and Hendon, 2004) and the boreal summer intraseasonal oscillation (BSISO) index (Lee et al., 2013), have been proposed to track the sub-seasonal oscillation propagation, these may not cover the patterns that could be important for subseasonal rainfall in some regions (Li et al., 2022). In addition, correlations with geopotential height (H) in the lower (H850), middle (H500), and upper (H200) troposphere are also analyzed. Leung and Qian (2017) showed that H at 850, 500, and 200 hPa are able to reflect the MJO structure as well as the zonal wind. Daily OLR data used in this study are provided by the National Oceanic and Atmospheric Administration (NOAA) on a 2.5° resolution global grid (Liebmann and Smith, 1996). The OLR data is derived from high-resolution infrared sounders and is valuable for a wide range of applications. The daily averaged U850, U200, H850, H500 and H200 data are derived from the fifth generation of the ECMWF reanalysis (ERA5) dataset. Reanalysis products combine model data with observations from across the world into a globally complete and consistent dataset using the laws of physics (Hersbach et al., 2020). This principle, called data assimilation, is based on the method used by numerical weather prediction centres, in which given a certain time interval a previous forecast is combined with newly available observations in an optimal way to produce a new best estimate of the state of the atmosphere, called analysis, from which an updated, improved forecast is issued. The ERA5 data used in this study have a horizontal resolution of 0.25°.
Regarding SST, we use the NOAA Daily Optimum Interpolation Sea Surface Temperature (OISST) (Huang et al., 2021) which is a long-term climate record integrating observations from different platforms (satellites, ships, buoys and Argo floats) on a regular global grid with a latitude-longitude resolution of 0.25°. The dataset is interpolated to fill gaps on the grid and create a spatially complete SST map. To be consistent with the temporal coverage of the OISST, we focus on the period 1982–2019. To focus on large-scale features and increase computational efficiency, the ERA5 reanalysis and SST data are interpolated to 2.5° × 2.5° spatial resolution.
The study focuses on the June–September (JJAS) period, which corresponds to the core of the West African monsoon (WAM) season in Senegal. This seasonal window was chosen because the WAM accounts for the vast majority of annual rainfall over the country, making it the most critical period for rain-fed agriculture, water resource management, and hydroclimatic risk assessment (Faye et al., 2024). The analysis period spans 1982 to 2019. The start year 1982 is determined by the availability of the NOAA Daily Optimum Interpolation Sea Surface Temperature (OISST) dataset (Huang et al., 2021), which provides the earliest consistent global SST record used in this study. The end year 2019 is constrained by the availability of the outgoing longwave radiation (OLR) dataset in the version accessed, which was available through 2019. To ensure temporal consistency across all datasets, including CHIRPS precipitation, ERA5 reanalysis, OLR, and OISST, the common analysis period was restricted to 1982–2019, yielding 38 JJAS seasons for model training and validation.
2.1.4 S2S hindcasts
The precipitation hindcasts (or reforecasts) from the S2S database were assessed for the models of the ECMWF, the Met Office (UKMO), and the National Centers for Environmental Prediction (NCEP). These forecasts exhibit various configurations, including forecast range, spatial resolution, frequency, period, ensemble size, and coupling effects (Table 1) (Vitart et al., 2017; de Andrade et al., 2021). Furthermore, ECMWF and UKMO reforecasts are generated gradually by updating their model versions according to near real-time forecasts, whereas in the NCEP model, reforecasts are based on a fixed date for a given model version. We analyzed ECMWF and UKMO reforecasts corresponding to model version dates of the year 2018. Four start dates per month were chosen based on UKMO initializations (1st, 9th, 17th, and 25th of each month), as indicated by de Andrade et al. (2021). The closest available start date is selected to account for non-aligned ECMWF initializations. These discrepancies in initialization times limited a fully consistent multimodel intercomparison. For a fair evaluation, six perturbed members in addition to the control member are used, and the ensemble mean is computed for each model. To ensure a fair comparison between models, we select six perturbed members, in addition to the control, to obtain the ensemble average for each model. For the NCEP model, which has a total of three perturbations, we have used three additional perturbations by modifying the initial states. This procedure allows all models to have at least seven members in their ensemble. Since the subseasonal time scale exceeds the limit of weather prediction, a weekly time frame was used to represent the subseasonal forecast range more adequately. Weekly precipitation was obtained considering four accumulation lead times: days 5–11 (week 1), 12–18 (week 2), 19–25 (week 3), and 26–32 (week 4). The results of our ML-based models will be compared to those from these S2S dynamical prediction systems (see Table 1) to assess their performance. The study period selected for comparison extends from 1999 to 2010, determined by the availability of NCEP model reforecasts.
| Model | Forecast length | Spatial resolution | Hindcast frequency | Hindcast period | Ensemble size | Ocean coupled | Sea ice coupled |
|---|---|---|---|---|---|---|---|
| ECMWF | 46 days | Tco639/319 L91 | Two per week | Past 20 years | 11 | Yes | No |
| UKMO | 60 days | N216 L85 | Four per month | 1993–2016 | 7 | Yes | Yes |
| NCEP | 44 days | T126 L64 | Daily | 1999–2010 | 4 + 3a | Yes | Yes |
The main features of the three S2S operational models and their hindcasts (de Andrade et al., 2021).
Three more perturbed members, extracted from 1-day lag after initializations, were added to the NCEP ensemble size.
Having described the observational, reanalysis, and hindcast datasets, we now outline the methodology used to prepare these data for machine learning. A critical step is the extraction of intraseasonal signals (10–60 day variability) from the raw fields, as these subseasonal fluctuations contain the bulk of S2S predictability for West African rainfall.
2.2 Methodology
2.2.1 Extraction of intraseasonal signals
In this part, we discuss the extraction of significant intraseasonal signals, important for S2S precipitation (Li et al., 2022). The raw daily data of atmospheric variables (U850, U200, OLR, H850, H500 and H200), daily precipitation, and SST contain high-frequency noise. Bandpass filtering methods, such as the fast Fourier transform, are commonly used to isolate the intraseasonal scale (10- to 60-day signals) (Zhang, 2005). However, such traditional approaches are not suitable for real-time applications, as they require information beyond the current date, leading to a look-ahead effect.
As proposed by Li et al. (2022), we use a non-filtering method to extract 10- to 60-day signals from atmospheric and oceanic variables and CHIRPS precipitation. Compared to traditional methods of extracting intraseasonal signal, this approach could be used for real-time applications. The climatological annual cycle of the raw daily data is first removed by subtracting a 90-day low-pass filtered climatological component, as follows:
where is the daily data. is the corresponding climatological 90-day low-pass filtered component derived by the Lanczos filtering method (Duchon, 1979). The period during which the low-pass filter is applied spans from 1982 to 2019.
In the second step, low-frequency signals of more than 60 days are removed by subtracting the 30-day moving average as follows:
where is the running mean of the 30 days of .
A weekly mean is performed on (for atmospheric and oceanic signals) to remove higher frequency signals, as follows:
As a result, the signal derived represents the 10- to 60-day signal of . The daily intraseasonal signals are averaged into weekly data to further reduce noise and improve predictability (Li et al., 2022). Hereinafter, the 10- to 60-day accumulated weekly precipitation signal will be referred to as the weekly precipitation anomaly.
Once the intraseasonal signals have been isolated from the raw atmospheric, oceanic, and precipitation fields, the next step is to identify which regions and variables contribute most to predictive skill. Accordingly, we define the predictors by quantifying the lagged statistical relationships between these intraseasonal signals and Senegalese rainfall anomalies.
2.2.2 Definition of predictors
Potential predictor regions were identified from lagged correlations between weekly precipitation anomalies (WPA) over Senegal and intraseasonal atmospheric and oceanic signals during June–September for the period 1982–2019. The objective is to identify large-scale circulation and sea-surface temperature patterns that exhibit statistically significant relationships with Senegalese rainfall at subseasonal lead times. Correlations were computed for weekly averages of atmospheric and oceanic intraseasonal signals at lead times ranging from 0 to 5 weeks prior to the target rainfall week.
For example, predicting the accumulated weekly precipitation for the period from June 1 to 7, 2010. In this case, the mean weekly intraseasonal oscillation (ISO) signals of the atmospheric field for the periods from June 1 to 7 (week 0), May 25 to 31 (week 1), May 18 to 24 (week 2), May 11 to 17 (week 3), May 4 to 10 (week 4), and April 27 to May 3 (week 5) are used as predictors to generate precipitation forecasts at different lead times.
Figure 2 depicts the correlation between the previous 10 to 60 days mean weekly signals of OLR, U850, and U200 and WPA over Senegal at different temporal lags. At weeks 5 and 4, significantly negative correlated OLR signals are primarily located above the Philippine Sea, the Bay of Bengal. These signals seem to propagate towards East Africa at week 4. OLR positive anomalies remain near the above Africa, and the Indian Ocean at week 3 to week 2. At weeks 1 and 0, significant negative OLR correlations become more pronounced over Africa, particularly over Central and West Africa. These negative correlations indicate that below-normal OLR values (enhanced cloudiness and deep convection) are associated with above-normal rainfall over Senegal. The strengthening of these negative correlations closer to the target week suggests that convective activity over tropical Africa acts as a precursor of enhanced rainfall conditions in Senegal. The high negative correlation over Senegal at week 0 confirms the results found in the region during the boreal summer. Accordingly, studies have shown that negative (positive) OLR anomalies are indicative of more (less) cloud coverage and hence enhanced (suppressed) convective precipitation in the region (Mohino et al., 2011; Janicot et al., 2008). Thus, when precipitation occurs, a decrease in outgoing longwave radiation in the atmosphere is observed. This phenomenon is attributed to cloud formation and the onset of atmospheric convection, two processes closely linked to precipitation in West Africa. Convective clouds reflect a portion of the infrared radiation back into space, thereby reducing the amount of infrared radiation detected by satellites (Janicot et al., 2008). For U850, positive correlations correspond to enhanced westerly anomalies, whereas negative correlations correspond to enhanced easterly anomalies (at weeks 4 and 5). Positive correlations identified over the Indian Ocean and East Africa suggest that stronger low-level westerlies are associated with wetter conditions over Senegal, likely through enhanced moisture transport toward West Africa. In contrast, negative correlations over the western Pacific indicate that stronger easterly anomalies in this region are associated with increased rainfall over Senegal. These U850 hPa signals are located between the Philippine Sea and the Horn of Africa at week 3 to week 2 and then seem to be found toward Central Africa at week 2. A more westward progression is noted for week 0 to week 1, with a spatial distribution of U850 signals more concentrated between subtropical latitudes of both hemispheres at week 0, indicating more robust statistical relationships in this area.
The U850 ocean–continent dipole is also observed between the Atlantic Ocean and the African continent, and may be related to the onset of the West African summer monsoon system or to low-level trade winds advecting moisture from the ocean toward the continent. In this context, Sultan et al. (2003) show that the boreal summer circulation dipole between ocean and continent is often associated with the pressure difference between these two areas. Statistically significant U200 anomaly correlations are found over the Indian and Pacific Oceans during weeks 5–3. At week 2, the signals appear above Africa and Indian Ocean, whereas at weeks 1–0 are located westward over the Atlantic Ocean. Negative correlations (in blue) correspond to an easterly wind anomaly, while positive correlations (in red) correspond to a westerly wind anomaly. These results in U200 correlations are related with the so-called Tropical Easterly Jet (TEJ) (Nicholson and Grist, 2003; Wu et al., 2009). The intensity and latitudinal position of this easterly jet influence precipitation in West Africa during the monsoon season (boreal summer). Specifically, a strengthening (intensification) of this easterly jet at 200 hPa is linked to increased rainfall over the Sahel and countries like Senegal. A more intense easterly jet promotes the ascent of humid air from the Gulf of Guinea, reinforcing convection and precipitation over the Sahel. The more northerly position of an intense jet (as in the case of weeks 2, 1, and 0) allows for a deeper penetration of the monsoon towards Sahelian latitudes, thus favoring precipitation. It appears that U850 correlations exhibit similar characteristics to U200 correlations in many regions, but with opposite signs, which is consistent with upper- and lower-level circulations associated with monsoon systems (tropical baroclinic circulations). Figure 3 depicts the correlation between the previous 10 to 60 days mean weekly H850, H500, and H200 signals and Senegal precipitation anomalies at different temporal lags. At weeks 5–4, significantly correlated H850 signals are primarily seen over Africa and progressively spread over the Atlantic. The signals are scattered across the tropics at lead times of 2 and 1 weeks. A pronounced signal emerges along the Senegalese coast at week 0. For H500 anomalies, the spatial distribution is broadly consistent with that of H850 correlations, with the strong coastal signal at week 0 also evident in H500 fields. In contrast, H200 anomaly correlations seem much more scattered compared to H850 and H500 throughout the weeks. At weeks 1–0, significantly correlated H200 signals are mainly seen over the tropical zone.
Figure 4 displays the correlation between the previous 10 to 60 days mean weekly SST signals and Senegal precipitation anomalies at different temporal lags. Positive signals are seen throughout the North Atlantic, the Mediterranean Sea, and the North Pacific during various weekly lead times. The positive signals over the North Atlantic become increasingly widespread and stronger as the lead time decreases. Accordingly, more robust and spatially extensive signals are observed over the North Atlantic from weeks 3 to 0. Significant negative correlations are also detected in several regions, including the South Atlantic, the eastern Indian Ocean, and parts of the South Pacific, depending on the lead time. These results are consistent with studies by Gaetani et al. (2010), Mohino et al. (2011), Fontaine et al. (2011), Diakhaté et al. (2019), and Thiam et al. (2024), showing that strong and significant positive SST anomalies in the Mediterranean precede a wetter than average summer in the Sahel. Jung (2006) identified a significant increase in precipitation in the Sahel following the Mediterranean heatwave of 2003. Strong associations are also found in the North Atlantic, with correlations exceeding 0.4 in the Gulf Stream region north of 30°N (Wang et al., 2012; Monerie et al., 2023; Liu et al., 2014) or in the northwest Pacific: extratropical warming of the northern hemisphere indeed induces a significant increase in precipitation in the Sahel through the modification of large-scale meridional heat distribution, according to Park et al. (2015) and Suárez-Moreno et al. (2018). Thus, all these results are in perfect agreement with our findings.
The spatiotemporal coupled covariance patterns are then constructed for a grid point where the correlation is statistically significant at the 5% level. The predictor is defined by calculating the sum of the products of the covariance patterns and the 10- to 60-day signals of atmospheric and oceanic fields for each preceding weekly interval following (Li et al., 2022):
denotes the weekly mean 10–60 day filtered signal of the atmospheric and oceanic field at grid point i, where the correlation between and Y is statistically significant at the 5% level. Here, k ranges from 1 to 7, representing 6 different atmospheric fields and the SST field. Y denotes the weekly mean precipitation anomalies. T is the total number of weeks, and N is the total number of grid points at which the correlation between and Y is statistically significant at the 5% level. Thus, for each atmospheric or oceanic field and each preceding week, there is only one predictor .
To ensure the robustness of the machine learning models, careful preprocessing of the predictors was implemented. The ISO signals from atmospheric and oceanic fields, as well as the precipitation anomalies, were normalized to improve numerical stability and model performance. We applied the Yeo-Johnson transformation (Yeo and Johnson, 2000), a method that generalizes the Box-Cox transformation to effectively handle zero and negative values. This approach helps make the variable distributions more symmetric and closer to normal, and brings them to a comparable scale, which is crucial for algorithms sensitive to data scaling, such as Support Vector Machine (SVM) and Ridge regression.
With the predictor set defined, we now turn to the machine learning models used to map these atmospheric and oceanic indices onto weekly precipitation anomalies over Senegal. We evaluate six algorithms that span a range of complexity and assumptions, from regularized linear models to ensemble tree methods.
2.2.3 Machine learning modeling
In the previous steps, predictors were defined by analyzing the relationship between global intraseasonal (weekly) signals, from atmospheric and oceanic fields, and precipitation anomalies in Senegal. The derived predictors can be used to forecast weekly precipitation anomalies in Senegal. Once the predictors were defined, ML models were constructed. Various modeling techniques were utilized, including Ridge regression, linear regression, random forests, multi-layer perceptron, and support vector machines.
Linear regression (LR) is a statistical method used to assess the linear relationship between a dependent variable y and one or more independent variables X (Draper, 1998). In general, linear regression is used when the relationship between the predictors X and outcome y can be reasonably approximated as linear (Weisberg, 2005). It works well when the predictors are not too highly correlated. The model can be used for prediction, inference and interpretation of the impact of each predictor on the outcome.
Linear regression is widely used in climate science for tasks like modeling temperature changes over time or predicting precipitation levels based on atmospheric variables. The interpretability of the model makes it a popular choice.
Ridge regression (also known as regularized regression) is a method that adds a regularization term to the cost function of linear regression in order to reduce the variance of the model (Hoerl and Kennard, 1970). The regularization term introduces bias into the estimates of the regression coefficients, but reduces variance by pulling them towards zero. In general, ridge regression outperforms simple linear regression in the case of multicollinearity between predictors, as it reduces the instability of coefficient estimates (Marquardt and Snee, 1975). Additionally, it is well-suited for high-dimensional problems by limiting overfitting risks (Zou and Hastie, 2005). This is why it is often used in climate science with a large number of variables (Faye et al., 2025). Ridge regression was used to model the relationship between CO2 concentrations and temperature (Gregory et al., 2004).
To capture potential nonlinearities that linear models may miss, we also employ support vector machines. Support vector machine (SVM) is a supervised ML algorithm and can be used for both classification and regression (Vapnik et al., 1998). SVM uses kernels function which can be linear or polynomial in order to obtain non-linear function (Gunn, 1998). SVM minimizes the error by adding the hyperplane and maximizing the margin between the prediction and the actual values (Karatzoglou et al., 2006). This model presents many advantages. It is very effective even with high dimensional data. It also works very well if the classes in the data are well separated points. SVM can also work with image data. Nevertheless, SVM model has some disadvantages. In fact, it is not easy to choose a good kernels function. Moreover, interpreting the final model remains challenging, particularly in understanding the contribution of individual variables and their relative importance, as well as in fine-tuning the model hyperparameters (Kirchner and Signorino, 2018). Complementing these methods, tree-based ensemble approaches, such as Random Forest, Gradient Boosting, and XGBoost, are used to capture non-linear interactions between variables. Random forests (RF) are an ensemble method based on decision trees. The principle is to construct a collection of decision trees, each trained on a bootstrap sample of the original data (resampling with replacement) (Breiman, 2001). In addition, for each split of a tree, only a random subset of the predictors is considered. Three hyperparameters need to be tuned in the RF algorithm, namely the number of trees in the “forest,” the number of features to consider when looking for the best split, and the maximum depth of the tree. In general, RF generates a higher quality global model than single decision tree models (Mutanga et al., 2012), because RF compensates for the bias introduced by the single decision tree due to its random character. Moreover, RF also shows efficiency in handling large dimensional datasets (Vincenzi et al., 2011), which helps to analyze the dataset of this study.
AdaBoost, or “Adaptive Boosting,” is an ensemble learning method that enhances the accuracy of prediction models by combining multiple weak models (often called “weak learners”) to create a strong model. This technique is particularly effective for classification tasks, although it can also be adapted for regression. AdaBoost uses weak learners, typically shallow decision trees (stumps), as base models. The models are trained sequentially, with each model attempting to correct the errors of previous models. The model predictions are combined with weighting to produce the final prediction.
Multilayer perceptron (MLP) is a type of artificial neural network capable of learning nonlinear relationships between variables (Rosenblatt, 1958). The MLP consists of an input layer, one or more hidden layers, and an output layer that are fully connected. Each neuron computes a linear combination of the inputs with associated weights, and applies a nonlinear activation function like the tanh or relu. The weights are adjusted through back propagation of the error gradient (Rumelhart et al., 1986). The MLP is capable of approximating any continuous function through its hidden-layer architecture (Cybenko, 1989). It is able to handle complex classification and regression problems. In climatology, the MLP is used for reconstructing missing data (Suzuki, 2011), precipitation forecasting (Hung et al., 2009), and solar radiation modeling (Li et al., 2020). Its ability to capture nonlinear relationships makes it a powerful tool.
The choice of these models was based on the performance they demonstrated in the studies conducted by Gerlitz et al. (2016), Cai et al. (2019), Toure et al. (2023), and Sarr and Sultan (2023). These references demonstrated the higher predictive effectiveness of these models compared to other alternatives. Therefore, we selected these particular ML models for our own analysis. In a recent study, PMM (Predictive Mean Matching), RF, and NORM (Bayesian Linear Regression) were used for multiple imputation to improve climate databases in Senegal (Toure et al., 2023). The results highlight the superior performance of the RF model in terms of accuracy and explained variance compared to other ML models. Sarr and Sultan (2023) used these ML techniques to predict crop yields in Senegal. Their results showed that combining climate and vegetation data with ML methods yields the best performance.
To determine optimal hyperparameters for selected ML models, we used a grid search method. This approach consists of exhaustively testing all possible combinations of hyperparameters from a predefined grid of values. Although computationally expensive, grid search has the advantage of being simple to implement. It allowed us to systematically identify the best hyperparameter configuration to maximize the performance of models. We used a one-year cross-validation (or k-fold cross-validation) to evaluate the machine learning models and assess the predictive performance of individual and combined predictors. Cross-validation is a procedure used to estimate the performance of a machine learning algorithm when making predictions on data not used during the training of the model. The cross-validation has a single hyperparameter “k” (here k = one-year) that controls the number of subsets that a dataset is split into. Once split, each subset is given the opportunity to be used as a test set while all other subsets together are used as a training dataset. This means that k-fold cross-validation involves fitting and evaluating k models. This, in turn, provides k estimates of a model’s performance on the dataset, which can be reported using summary statistics such as the mean and standard deviation. Then, to compare with the S2S models, we refined our approach by selecting the two best-performing models, Ridge and linear regression (LR). In addition to achieving the best results, these models require fewer computational resources, allowing us to apply the Leave-One-Year-Out (LOYO) validation without encountering limitations related to computing capacity. Although LOYO validation is a computationally demanding method, it provides a reliable and unbiased estimate of model performance. Leave-one-out cross-validation is a specific configuration of k-fold cross-validation, where k is equal to the number of examples in the dataset. LOYO represents an extreme version of this approach, involving the highest computational cost. Indeed, it requires training and evaluating a model for each example in the training set. The advantage of such a large number of evaluations is that it provides a more robust estimate of model performance, as each observation has the opportunity to represent the entire test dataset. Thus, with a dataset covering 38 years, thirty-eight training and validation processes were conducted using the LOYO method.
To ensure robust and generalizable performance estimates, we evaluate all models using a leave-one-year-out (LOYO) cross-validation strategy. The following section details the skill metrics used to quantify forecast accuracy and its decay with lead time.
2.2.4 Prediction skill-scores
We conducted leave-one-year-out cross-validation to assess the practicality of the models, i.e., using all years’ data from 1982 to 2019 except the target year to train the model and then make a prediction for the target year. This approach is an extensively used cross-validation method because of its simplicity, universality, and superiority in avoiding the issue of over-fitting. In this study, we utilize the mean absolute error (MAE) to provide an overall assessment of the forecast accuracy for weekly precipitation anomalies. The MAE is calculated as follows:
where is the total number of data (sample size), is the actual value of the -th data point, and is the corresponding predicted value. The mean absolute error, which quantifies the absolute difference between the values predicted by the model and the observations, is a positive metric. Thus, the closer the value of this average error tends towards zero, the more the model’s predictive performance is judged to be excellent. A low MAE therefore demonstrates high accuracy of the forecasts generated by the model compared to actual measurements.
Additionally, we used the Pearson correlation coefficient to evaluate the linear relationship between the predicted and observed values. The Pearson correlation coefficient is calculated as follows:
where and are the means of the actual and predicted values, respectively. The Pearson correlation coefficient, , ranges from −1 to 1, where values closer to 1 indicate a strong positive linear relationship, values closer to −1 indicate a strong negative linear relationship, and values around 0 indicate no linear relationship.
By computing both the MAE and the Pearson correlation coefficient (or anomaly correlation coefficient—ACC), we can provide a comprehensive evaluation of the model’s performance in predicting weekly precipitation anomalies.
3 Results
3.1 Assessment of S2S model performance for precipitation forecasting
In this section, we evaluate the quality of weekly precipitation hindcasts over West Africa using deterministic forecast verification metrics. The assessment covers the period from May to September during 1999–2010, focusing on three selected S2S models, namely, ECMWF, UKMO, and NCEP (for details on these models, see Section 2.1.4). Figure 5 illustrates the correlation between the hindcast ensemble mean and observed anomalies of accumulated precipitation for each S2S model at different weekly lead times. The highest associations are observed at week 1, with correlations decreasing as the lead time increases. Significant correlations are primarily found in countries such as Senegal, Mauritania, Niger, Nigeria, and Burkina Faso, notably during the first and second weeks, with positive correlations persisting across all lead times in Senegal. Weak associations are seen for lead times from 3 to 4 weeks ahead, particularly over the Sahel. A comparison between models reveals that UKMO outperforms ECMWF and NCEP in several Sahelian areas, including Senegal, Chad, and the coastal countries of the Gulf of Guinea. This finding aligns with the results of de Andrade et al. (2021), who demonstrated significant correlations up to week 4 over West Africa near the Gulf of Guinea for almost all models.
3.2 Comparative analysis of machine learning models for precipitation forecasting
We use six ML algorithms, detailed in the methods section: Ridge, LR, RF, SVM, Adaboost, and MLP (see section 2.2.3 for details). We then calculate the ensemble average of these models (hereinafter mean_ML). Figure 6 shows the MAE and ACC of weekly precipitation anomalies predicted by the different ML algorithms, as well as their ensemble average. The results show higher predictive ability for Ridge Regression in terms of ACC and MAE for all forecast intervals, followed by the LR model. This suggests that the Ridge method further enhances prediction skill across all intervals, yielding this technique for our subsequent analysis.
Building on the overall model ranking established above, we now investigate the physical sources of predictability by examining the contribution of each predictor individually. This decomposition helps clarify why atmospheric and oceanic variables dominate at different lead times.
3.3 Predictive performance of individual and combined predictors
To better understand the primary sources of predictability for intra-seasonal precipitation, we employ the Ridge regression model for each predictor (i.e., atmospheric and/or oceanic fields) individually. Figure 7 presents a comparison of MAE and ACC for forecasts of weekly precipitation anomalies obtained from different predictors. Overall, predictors such as OLR, U200, and U850 demonstrate high predictive ability (low MAE values and high ACC values) compared to predictors such as H200, H500, and H850 for nearly all lead times. This finding aligns with the ACC values shown in Figures 2, 3. The skillful predictors (OLR, U200, and U850) exhibit stronger linear relationships with precipitation, with OLR showing the strongest association. These results suggest that ISO signals from OLR, U200, and U850 contribute significantly to the prediction skill of sub-seasonal precipitation. The overall decline in forecast skill from week 0 to week 4 reflects the progressive loss of deterministic atmospheric predictability. At short lead times, atmospheric variables dominate: OLR captures convective activity, while U200 and U850 represent the baroclinic structure of the monsoon circulation. As these atmospheric signals decay with increasing lead time, the slowly evolving SST signal, linked to Mediterranean and North Atlantic teleconnections, contributes a growing share of the remaining predictability. This oceanic memory, arising from the ocean’s higher heat capacity and higher thermal inertia, does not raise absolute skill at short leads; rather, it partially offsets the continued atmospheric decay at longer leads. Consequently, the small recovery in ACC observed at week 5 is dominated by SST predictors, consistent with the highest SST-based ACC occurring at this lead time in Figure 7. This shift in the primary driver of predictability—from atmospheric processes at short leads to oceanic memory at longer leads—is physically consistent with the differing response timescales of the two components of the climate system. Comparing the Ridge regression model built with a single predictor to the one built with all atmospheric predictors (denoted as ), we find that the latter enhances forecasting capability. This improvement is even more evident when all atmospheric and oceanic predictors (denoted as all in Figure 7) are incorporated, thereby enhancing the predictive ability of the model. These results indicate enhanced forecast accuracy of intraseasonal precipitation with the simultaneous inclusion of different predictors, suggesting the relative relevance of oceanic and atmospheric variables under longer and shorter lead times, respectively.
While the previous analysis examined temporal variations in predictor importance, it does not reveal how forecast skill is distributed across Senegal. We therefore examine the spatial heterogeneity of model performance to identify regions where the Ridge regression model excels or struggles.
3.4 Spatial variability in forecast accuracy for weekly precipitation anomalies
In this section, we focus on the spatial aspects of prediction skill for weekly precipitation anomalies in Senegal, the core region of interest. Figure 8 illustrates the MAE obtained from the LOYO validation using a Ridge regression model for weekly precipitation anomaly forecasts. The analysis covers the period from June to September, 1982 to 2019, across Senegal under different weekly lead times. The 6-panel block on the left depicts the results using intraseasonal SST predictors. A northwest-southeast gradient of MAE values is observed, consistent across lead times. Relatively lower MAE values are consistent in the northwestern area of Senegal across the different lead times, indicating higher model performance. In contrast, the southeastern part of the country exhibits higher MAE values, showing moderate prediction skill. The apparent improvement in model performance as lead time increases is attributed to the longer-term predictive capability of ocean–atmosphere interactions compared to local atmospheric predictors. The block to the right in Figure 8 presents the MAE using solely intraseasonal atmospheric variables as predictors. The spatial pattern follows a similar northwest-southeast gradient as the previous case for the oceanic predictors. Model performance improves under shorter lead times and in the northwest region with a gradual decline towards the southeast. Figure 9 shows the results of evaluating the performance of the Ridge regression model using both intraseasonal atmospheric and oceanic variables as predictors. As in the previous cases, MAE scores reveal a persistent northwest-southeast gradient for all lead times. Decreased predictive ability (increased MAE scores) is observed as lead time increases. These spatial variations can be attributed to the differential contributions between atmospheric and oceanic variables according to lead times and impact regions. Given the similarity between Figure 9 and the atmospheric predictors-only block in Figure 8 (right panels), atmospheric signals appear to play a more prominent role in precipitation forecasting compared to SST variations when it comes to the subseasonal time scale (up to week 5 lead time). This is attributed to longer lead times in the significant response of precipitation to SST forcing, particularly in the southeastern area of Senegal. When assessing the weekly precipitation forecast in Senegal using the ACC metric with all predictors (Supplementary Figures S4, S5), we observe a significant decline in ACC as the forecast lead time increases, which further underscores the challenges in maintaining high prediction skill at longer lead times.
Supplementary Figure S1 further illustrates the impact of including oceanic predictors through the difference in MAE (MAE) between the two model configurations, calculated for each forecast lead time from week 0 to week 5. The values correspond to (). Negative values (in blue) indicate that adding oceanic variables reduces MAE, thereby improving model performance, while positive values (in red) reflect an increase in MAE, indicating reduced performance. Overall, the inclusion of oceanic predictors leads to a general improvement across most of Senegal, particularly in coastal and southern regions where MAE decreases significantly. This effect becomes more pronounced from week 2 onwards and persists until week 5, suggesting that oceanic variables (e.g., sea surface temperature – SST – or coupled ocean–atmosphere indices) provide useful information for medium-range forecasts. The slower evolution of oceanic signals compared to atmospheric parameters offers a form of climate system memory that enhances predictability. Conversely, some inland areas show slightly positive values, indicating that in these regions, oceanic forcing is less dominant and predictability relies more on local atmospheric dynamics. The same result is observed when considering other metrics such as ACC and RMSE (see Supplementary Figures S2–S4). To better contextualize the performance of the Ridge model trained using LOYO validation, we compared it against several reference methods (climatology, persistence, and damped persistence) with the aim of providing robust benchmarks to quantify the gains delivered by advanced forecasting systems. The calculation methods for these baselines are described in the Supplementary material. The results in the Table 2 show a consistent evolution of forecast performance across the forecast weeks.
| Model | Week-1 | Week-2 | Week-3 | Week-4 |
|---|---|---|---|---|
| Climatology | 1.20 | 1.21 | ||
| Persistence | 1.81 | 1.73 | 1.78 | 1.78 |
| Damped persistence | 1.72 | 1.57 | 1.56 | 1.51 |
| Ridge (OLR only) | 0.70 | 0.71 | 0.72 | 0.73 |
| Ridge (all predictors) | 0.69 | 0.70 | 0.70 | 0.71 |
| Obs. std. deviation | 1.97 | 1.90 | 1.75 | 1.74 |
| MAE as % of Std. Dev | 60.8% | 64.0% | 69.1% | 70.3% |
Model performance evaluation.
The reference methods (climatology, persistence, and damped persistence) exhibit higher errors, as expected. The climatology remains almost constant () across all weeks, reflecting its static nature. The simple persistence method shows larger errors (), highlighting the limitation of assuming that anomalies persist unchanged. The damped persistence slightly improves the results, especially from week 2 onward, by accounting for the gradual loss of atmospheric memory. The Ridge regression models clearly outperform the reference methods. In particular, the Ridge model using all predictors achieves the lowest errors (), indicating its superior ability to capture relevant predictive signals from multiple variables, compared to the OLR-only version. Finally, the MAE expressed as a percentage of the standard deviation indicates an overall satisfactory performance, with errors representing roughly 60–70% of the natural variability of the system. The gradual increase with lead time is expected, as forecast skill naturally decreases with time. Overall, these results confirm the robustness of the Ridge model especially when using multiple predictors and demonstrate its effectiveness in improving weekly forecasts compared to simple reference approaches. This is why we used the outputs of models trained with LOYO validation to compare them with our three models from the S2S database. Having established that the Ridge model outperforms simple reference baselines, we now place its performance in an operational context by comparing it directly against state-of-the-art dynamical S2S forecasting systems. This comparison addresses whether the ML approach offers genuine added value beyond current operational capabilities.
3.5 Comparative analysis of ML and S2S models for precipitation forecasting
This section presents a comprehensive comparison between the top-performing ML models and S2S models in forecasting precipitation in Senegal. Figure 10 illustrates the ACC between the ML approaches that relatively show the best predictive ability (Ridge, LR) and GCM-based models from the S2S database (ECMWF, UKMO, and NCEP). We noticed that UKMO seems more skillful than ECMWF and NCEP over Senegal for weeks 1 and 2. For all S2S models (GCM-based models), ACC scores show a consistent degradation as forecast lead time increases. Statistically significant ACC values are observed during weeks 1 and 2, spreading over much of Senegal, particularly for the UKMO and NCEP models, while ECMWF’s predictive skill is limited even under short lead times, being confined mainly to the coastal band. The remarkable improvement in skill is evident for ML-based models (Ridge, LR), which exhibit significant positive ACC values for all lead times, extending to the entire country. The improvement with respect to S2S models is outstanding as the prediction horizon increases, placing ML models as an efficient complement or even alternative to GCM-based subseasonal forecast systems. These performance differences are further explored in Table 3. It depicts the regionally averaged ACC scores, as a function of forecast lead times, for both S2S and ML models. The result unequivocally shows that ML models outperform S2S.
| Model | Week1 | Week2 | Week3 | Week4 |
|---|---|---|---|---|
| Ridge | 0.434 | 0.44 | 0.39 | 0.35 |
| LR | 0.433 | 0.43 | 0.38 | 0.34 |
| UKMO | 0.29 | 0.25 | 0.08 | 0.05 |
| ECMWF | 0.18 | 0.15 | 0.14 | 0.04 |
| NCEP | 0.28 | 0.18 | 0.07 | 0.16 |
Domain-averaged ACC for weekly precipitation forecasts over Senegal (week1–week4: forecast weeks 1 to 4).
Bold values in this Table indicate the highest domain-averaged ACC score for each forecast week (Week 1–Week 4).
In the southeastern region of Senegal, all models, particularly the S2S models, exhibit a notable skill degradation. This decrease in forecast accuracy could stem from the region’s unique geographical characteristics, notably the presence of vegetation and the land-atmosphere interactions that play a role in precipitation formation. This emphasizes the significance of considering these geographical factors when analyzing and predicting weather phenomena in Senegal, and highlights potential avenues for enhancing the accuracy of prediction models in the future. Supplementary Figure S5 presents the MAE-based Skill Score (), which evaluates the performance of the Ridge regression model relative to the three operational S2S models (ECMWF, UKMO, and NCEP) over the first four forecast weeks. SS values ranging from 0.96 to 0.97 indicate that Ridge significantly reduces the mean absolute error compared to the S2S models, demonstrating a robust improvement in predictive skill. Our comparative analysis underscores the remarkable potential of ML approaches in improving subseasonal precipitation forecasting in Senegal. The remarkable performance of these approaches, both spatially and for longer lead times, indicate the added value of ML to provide supplementary insights alongside existing dynamical forecasting systems. In the immediate future, while further work is needed to improve GCMs to achieve a better understanding of the physical mechanisms underlying climate predictability, ML techniques are highly valuable to improve the quality of subseasonal forecast, leading to more reliable predictions of precipitation, crucial for water resource management, agricultural planning, and risk assessment.
4 Discussion
This study demonstrates that machine learning models, particularly ridge regression, provide skillful subseasonal-to-seasonal precipitation forecasts in Senegal that surpass state-of-the-art dynamical prediction systems. Among the six models evaluated, ridge regression consistently achieved the lowest MAE and highest ACC across all lead times. The combination of atmospheric and oceanic predictors yielded the best overall performance, with atmospheric variables (OLR, U200, U850) dominating at shorter lead times and oceanic variables (SST) gaining relevance as the forecast horizon increases.
4.1 Comparison with prior research and novelty
The superior performance of the ML models aligns with the well-documented challenges GCMs face in representing West African monsoon dynamics, including the African Easterly Jet, mesoscale convective systems, and land-surface-atmosphere interactions (Rodrigues et al., 2014; Roehrig et al., 2013; Kniffka et al., 2020). Our findings extend recent advances in ML-based climate forecasting to the subseasonal time scale for West Africa, a region where few ML-based S2S studies exist.
The predictor hierarchy revealed by the ML models is consistent with known physical mechanisms. OLR emerges as the strongest predictor at short lead times, reflecting its role as a proxy for deep convection. The correlation maps (Figure 2) show convective signals propagating eastward from the Indian Ocean toward West Africa, consistent with the modulation of regional convective activity by the MJO during boreal summer. The complementary contributions of U850 and U200, with opposite-sign correlations in many regions, reflect the baroclinic structure of the monsoon: low-level southwesterly flow carries moisture from the Gulf of Guinea, while the upper-level Tropical Easterly Jet promotes ascent and convection over the Sahel. At longer lead times, SST gains relevance as the dominant predictor. This shift is not only consistent with the ocean’s higher heat capacity and thermal inertia providing a “memory” that extends predictability beyond the chaotic limits of atmospheric processes (Mariotti et al., 2018; Bach et al., 2019). Importantly, this oceanic memory does not improve absolute forecast skill at short leads; rather, it slows the decay of predictability as atmospheric signals lose coherence, producing the relative skill recovery observed at week 5. This also reflects specific teleconnection pathways identified in the correlation analysis (Figure 4): positive SST anomalies in the Mediterranean and North Atlantic (Gaetani et al., 2010; Mohino et al., 2011; Suárez-Moreno et al., 2018) and extratropical Northern Hemisphere warming influencing Sahel rainfall through meridional heat redistribution (Park et al., 2015).
A key finding is that the predictor hierarchy emerging from the ML models is physically interpretable, despite the models being purely statistical. Rather than operating as a black box, the data-driven approach recovers known physical teleconnections, which lends confidence to the forecasts and bridges the gap between statistical prediction and process-based understanding. Importantly, the non-filtering method employed for predictor extraction avoids the use of future information, making it suitable for real-time operational forecasting, unlike traditional bandpass filtering approaches. Beyond the physical interpretability of the predictor hierarchy, the reliability of these findings depends on the methodological framework employed. We therefore turn to the modeling choices, data limitations, and the spatial heterogeneity of forecast skill to assess the robustness and generalizability of our results.
4.2 Methodological considerations and limitations
A deliberate choice in this study was to employ interpretable and computationally efficient models rather than complex deep learning architectures. While such models are sometimes considered to have less predictive power for capturing complex nonlinear relationships, ridge regression consistently outperformed more complex alternatives within our evaluation, including Random Forest, AdaBoost, and MLP. This highlights the importance of aligning model complexity with the structure of the available data: after the predictor preprocessing described in Section 2, the underlying relationships are sufficiently linear for competitive forecasting. The simplicity of these models offers additional advantages in terms of transparency, reproducibility, and low computational cost. Moreover, the performance established here provides a rigorous benchmark against which future studies employing more complex architectures, such as convolutional neural networks for joint spatial forecasting, should demonstrate clear added value to justify their additional complexity.
Some limitations should be acknowledged though. CHIRPS and ERA5 carry observational uncertainties, though neither is a model forecast: CHIRPS blends satellite and station data, while ERA5 is a reanalysis constrained by observations. We mitigate these through (1) regional validation of CHIRPS (Tarnavsky et al., 2014; Funk et al., 2015; Faye et al., 2024), (2) intraseasonal anomaly computation that removes systematic bias, and (3) use of the best available long-record products. Residual uncertainty is unlikely to affect relative model rankings or our main conclusions. Although the modeling is performed at each grid point across Senegal, the observed northwest-southeast MAE gradient indicates that southeastern Senegal remains more challenging to predict. This spatial heterogeneity in predictive skill suggests that the dominant forcing mechanisms differ regionally, with oceanic teleconnections exerting stronger control in the northwest and local or land-surface processes playing a larger role in the southeast. Additionally, the set of predictors does not include land surface variables such as soil moisture, vegetation cover, or topographic features, which could further enhance model robustness, particularly in regions where oceanic forcing is less dominant.
This spatial gradient in forecast skill closely mirrors the climatological precipitation gradient across Senegal (Figure 1), where cumulative rainfall increases markedly from the arid northwest to the wetter southeast. In absolute terms, larger rainfall totals in the southeast naturally yield larger anomalies and thus higher MAE values. However, this pattern also reflects a physical limitation: the southeastern region, being closer to the core of the Intertropical Convergence Zone, is characterized by more intense mesoscale convective activity and stronger land-atmosphere coupling, processes that are not explicitly captured by the large-scale atmospheric and oceanic predictors used in this study. This suggests that incorporating land surface variables, such as soil moisture and vegetation indices, could be particularly beneficial for improving prediction skill in this region.
4.3 Implications for operational forecasting and climate adaptation
Despite these limitations, the results demonstrate that ML approaches can serve as a computationally efficient complement, or even alternative, to GCM-based subseasonal forecasting. The operational deployment of such models could significantly advance hydrometeorological risk assessment and water resource management in Senegal, with direct implications for the predominantly rain-fed agricultural sector. The low computational cost of these models makes them particularly suitable for operational deployment in resource-limited settings across West Africa.
4.4 Future directions
Future research should address the identified limitations through several avenues. Integrating ground observations from Senegal alongside additional explanatory variables (topography, vegetation cover, soil moisture) would enhance model robustness. Applying localized ML models based on homogeneous climatic zones could improve spatial prediction skill, particularly in southeastern Senegal. Deep neural network architectures, such as CNNs, should be explored for joint spatial forecasting of precipitation across the region. Finally, physics-informed machine learning models offer a promising path to combine the interpretability of physically based approaches with the flexibility and skill of data-driven methods. The methodology evaluated here is not limited to Senegal and can be extended to other climatic regions, broadening the applicability of ML-based subseasonal forecasting.
5 Summary and conclusion
Through this study, we advance on the field of sub-seasonal to seasonal precipitation forecasting in Senegal by leveraging state-of-the-art ML techniques, namely, Ridge regression, LR, RF, SVM, Adaboost, and MLP. As a starting point, our approach introduced a comprehensive analysis of the links between various intra-seasonal atmospheric and oceanic signals, and observed precipitation patterns in Senegal. In this framework, we conducted an in-depth examination of spatiotemporal correlations between atmospheric (U850, U200, OLR, H850, H500, and H200) and oceanic (SST) variables, and precipitation. With this analysis, we identified regions of statistically significant ACC scores, serving to construct spatiotemporal covariance models to define our set of predictors through the summation of the product of the covariance fields and intra-seasonal signals. To evaluate the performance of ML models, we used MAE and ACC metrics. The results were rigorously compared over the period 1982–2019 using a leave-one-year-out cross-validation method, showing that ML models provide skillful weekly precipitation forecasts for Senegal. Notably, the Ridge regression model consistently outperformed all other models and the ensemble mean across almost all lead times. Our analysis of individual sets of predictors, separating between atmospheric and oceanic variables, revealed that OLR, U200, and U850 contributed most significantly to improving the prediction skill of sub-seasonal precipitation under shorter lead times, particularly OLR. As the prediction horizon increases, oceanic variables improve their forecast ability, while atmospheric signals alone showed progressive skill degradation. As a third case study, the combination of both atmospheric and oceanic predictors enhanced forecasting skill-scores. This shift in the primary drivers of predictability from atmospheric to oceanic factors as forecast lead time increases occurs because oceans have a much greater heat capacity and thermal inertia than the atmosphere, allowing SST to maintain and gradually release energy over extended periods. As a result, oceanic conditions become increasingly important for longer lead times. On the side of the GCM-based models used in our study (NCEP, ECMWF and UKMO), we found a pronounced degradation in prediction skill as forecast lead time increases, with UKMO showing enhanced performance. Our comparative analysis against ML approaches revealed the potential of ML for subseasonal-to-seasonal forecasting, particularly for longer lead times, enhancing sub-seasonal forecasting performance compared to GCM-based systems.
Statements
Author contributions
DF: Investigation, Methodology, Software, Data curation, Writing – original draft, Writing – review & editing, Visualization, Conceptualization, Resources, Validation, Formal analysis. FA: Writing – original draft, Methodology, Validation, Writing – review & editing. RS-M: Writing – original draft, Supervision, Writing – review & editing, Validation. DW: Writing – review & editing, Writing – original draft. MH: Writing – review & editing. AD: Writing – review & editing. RL: Writing – review & editing. AG: Supervision, Writing – review & editing.
Funding
The author(s) declared that financial support was not received for this work and/or its publication.
Acknowledgments
I would like to express my sincere gratitude to the entire team who contributed to the success of this paper. Your hard work, dedication, and collaboration were invaluable throughout the research process. Thank you for your support and commitment to excellence. This work is based on S2S data. S2S is a joint initiative of the World Weather Research Programme (WWRP) and the World Climate Research Programme (WCRP). The original S2S database is hosted at ECMWF as an extension of the TIGGE database.
Conflict of interest
The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Generative AI statement
The author(s) declared that Generative AI was not used in the creation of this manuscript.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
References
-
AbbotJ.MarohasyJ. (2014). Input selection and optimisation for monthly rainfall forecasting in Queensland, Australia, using artificial neural networks. Atmos. Res.138, 166–178. doi: 10.1016/j.atmosres.2013.11.002
-
BachE.MotesharreiS.KalnayE.Ruiz-BarradasA. (2019). Local atmosphere–ocean predictability: dynamical origins, lead times, and seasonality. J. Clim.32, 7507–7519. doi: 10.1175/jcli-d-18-0817.1
-
BarnstonA. G.SmithT. M. (1996). Specification and prediction of global surface temperature and precipitation from global SST using CCA. J. Clim.9, 2660–2697. doi: 10.1175/1520-0442(1996)009<2660:SAPOGS>2.0.CO;2
-
BiasuttiM. (2019). Rainfall trends in the African Sahel: characteristics, processes, and causes. Wiley Interdiscip. Rev. Clim. Chang.10:e591. doi: 10.1002/wcc.591,
-
BreimanL. (2001). Random forests. Mach. Learn.45, 5–32. doi: 10.1023/a:1010933404324
-
BrunetG.ShapiroM.HoskinsB.MoncrieffM.DoleR.KiladisG. N.et al. (2010). Collaboration of the weather and climate communities to advance subseasonal-to-seasonal prediction. Bull. Am. Meteorol. Soc.91, 1397–1406. doi: 10.1175/2010bams3013.1
-
CaiY.GuanK.LobellD.PotgieterA. B.WangS.PengJ.et al. (2019). Integrating satellite and climate data to predict wheat yield in Australia using machine learning approaches. Agric. For. Meteorol.274, 144–159. doi: 10.1016/j.agrformet.2019.03.010
-
CybenkoG. (1989). Approximation by superpositions of a sigmoidal function. Math. Control Signals Syst.2, 303–314. doi: 10.1007/bf02551274
-
de AndradeF. M.CoelhoC. A. S.CavalcantiI. F. A. (2019). Global precipitation hindcast quality assessment of the subseasonal to seasonal (S2S) prediction project models. Clim. Dyn.52, 5451–5475. doi: 10.1007/s00382-018-4457-z
-
de AndradeF. M.YoungM. P.MacLeodD.HironsL. C.WoolnoughS. J.BlackE. (2021). Subseasonal precipitation prediction for Africa: forecast evaluation and sources of predictability. Weather Forecast.36, 265–284. doi: 10.1175/waf-d-20-0054.1
-
DiakhatéM.Rodriguez-FonsecaB.GómaraI.MohinoE.DiengA. L.GayeA. T. (2019). Oceanic forcing on interannual variability of Sahel heavy and moderate daily rainfall. J. Hydrometeorol.20, 397–410. doi: 10.1175/jhm-d-18-0035.1
-
DraperN. R. (1998). Applied regression analysis bibliography update 1994-97. Commun. Stat. Theory Methods27, 2581–2623. doi: 10.1080/03610929808832244
-
DuchonC. E. (1979). Lanczos filtering in one and two dimensions. J. Appl. Meteorol. Climatol.18, 1016–1022. doi: 10.1175/1520-0450(1979)018<1016:LFIOAT>2.0.CO;2
-
EdenJ. M.van OldenborghG. J.HawkinsE.SucklingE. B. (2015). A global empirical system for probabilistic seasonal climate prediction. Geosci. Model Dev.8, 3947–3973. doi: 10.5194/gmd-8-3947-2015
-
FallM.DiengA. L.SallSaϊdou M.SaneY.DiakhatéM. (2020). Synoptic analysis of extreme rainfall event in West Africa: the case of Linguère. Am. J. Environ. Prot.8, 1–9. doi: 10.12691/env-8-1-1
-
FayeD.KalyF.DiengA. L.WaneD.FallC. M. N.MignotJ.et al. (2024). Regionalization of the onset and offset of the rainy season in Senegal using Kohonen self-organizing maps. Atmos15:378. doi: 10.3390/atmos15030378
-
FayeD.LguensatR.KalyF.SudmantA.GayeA. T.KalisaE. (2025). Machine learning for air quality forecasting: insights from five Rwanda provinces. Sci. Afr.:e02959. doi: 10.1016/j.sciaf.2025.e02959
-
FontaineB.GaetaniM.UllmannA.RoucouP. (2011). Time evolution of observed July–September Sea surface temperature-Sahel climate teleconnection with removed quasi-global effect (1900–2008). J. Geophys. Res. Atmos.116. doi: 10.1029/2010JD014843
-
FunkC.PetersonP.LandsfeldM.PedrerosD.VerdinJ.ShuklaS.et al. (2015). The climate hazards infrared precipitation with stations—a new environmental record for monitoring extremes. Sci. Data2, 1–21. doi: 10.1038/sdata.2015.66,
-
GaetaniM.FontaineB.RoucouP.BaldiM. (2010). Influence of the Mediterranean Sea on the west African monsoon: intraseasonal variability in numerical simulations. J. Geophys. Res. Atmos.115. doi: 10.1029/2010jd014436
-
GerlitzL.VorogushynS.ApelH.GafurovA.Unger-ShayestehK.MerzB. (2016). A statistically based seasonal precipitation forecast model with automatic predictor selection and its application to central and South Asia. Hydrol. Earth Syst. Sci.20, 4605–4623. doi: 10.5194/hess-20-4605-2016
-
GregoryJ. M.IngramW. J.PalmerM. A.JonesG. S.StottP. A.ThorpeR. B.et al. (2004). A new method for diagnosing radiative forcing and climate sensitivity. Geophys. Res. Lett.31. doi: 10.1029/2003gl018747
-
GunnS. (1998). Support Vector Machines for Classification and Regression University of Southampton ISIS (Image Speech and Intelligent Systems Group) Technical Report, Southampton, UK: University of Southampton. 1–52.
-
HersbachH.BellB.BerrisfordP.DahlgrenP.HorányiA.Munoz-SebaterJ.et al. (2020). The ERA5 Global Reanalysis: Achieving a Detailed Record of the Climate and Weather for the Past 70 Years. European Geophysical Union General Assembly, 3–8. doi: 10.1002/qj.3803
-
HoerlA. E.KennardR. W. (1970). Ridge regression: biased estimation for nonorthogonal problems. Technometrics12, 55–67. doi: 10.1080/00401706.1970.10488634
-
HuangB.LiuC.BanzonV.FreemanE.GrahamG.HankinsB.et al. (2021). Improvements of the daily optimum interpolation sea surface temperature (DOISST) version 2.1. J. Clim.34, 2923–2939. doi: 10.1175/jcli-d-20-0166.1
-
HungN. Q.BabelM. S.WeesakulS.TripathiN. K. (2009). An artificial neural network model for rainfall forecasting in Bangkok, Thailand. Hydrol. Earth Syst. Sci.13, 1413–1425. doi: 10.5194/hess-13-1413-2009
-
HwangS.-O.SchemmJ.-K. E.BarnstonA. G.KwonW.-T. (2001). Long-lead seasonal forecast skill in far eastern Asia using canonical correlation analysis. J. Clim.14, 3005–3016. doi: 10.1175/1520-0442(2001)014<>2.0.co;2
-
JanicotS.ThorncroftC. D.AliA.AsencioN.BerryG.BockO.et al. (2008). Large-scale overview of the summer monsoon over West Africa during the AMMA field experiment in 2006. Ann. Geophys.26, 2569–2595. doi: 10.5194/angeo-26-2569-2008
-
JungG.. (2006). Regional Climate Change and the Impact on Hydrology in the Volta Basin of West Africa Publish/Report a Document. Germany: University of Augsburg, Augsburg.
-
KaratzoglouA.MeyerD.HornikK. (2006). Support vector machines in r. J. Stat. Softw. Augsburg, Germany: University of Augsburg. 15, 1–28. doi: 10.18637/jss.v015.i09
-
KirchnerA.SignorinoC. S. (2018). Using support vector machines for survey research. Surv. Pract.11, 1–14. doi: 10.29115/sp-2018-0001
-
KniffkaA.KnippertzP.FinkA. H.BenedettiA.BrooksM. E.HillP. G.et al. (2020). An evaluation of operational and research weather forecasts for southern West Africa using observations from the DACCIWA field campaign in June–July 2016. Q. J. R. Meteorol. Soc.146, 1121–1148. doi: 10.1002/qj.3729
-
LeeJ.-Y.WangB.WheelerM. C.FuX.WaliserD. E.KangI.-S. (2013). Real-time multivariate indices for the boreal summer intraseasonal oscillation over the Asian summer monsoon region. Clim. Dyn.40, 493–509. doi: 10.1007/s00382-012-1544-4
-
LeungJ. C.-H.QianW. (2017). Monitoring the madden–Julian oscillation with geopotential height. Clim. Dyn.49, 1981–2006. doi: 10.1007/s00382-016-3431-x
-
LiP.BessafiM.MorelB.ChabriatJ.-P.DelsautM.LiQ. (2020). Daily surface solar radiation prediction mapping using artificial neural network: the case study of Reunion Island. J. Sol. Energy Eng.142:021009. doi: 10.1115/1.4045274
-
LiY.WuZ.HeH.WangQ. J.XuH.LuG. (2021). Post-processing sub-seasonal precipitation forecasts at various spatiotemporal scales across China during boreal summer monsoon. J. Hydrol.598:125742. doi: 10.1016/j.jhydrol.2020.125742
-
LiY.WuZ.HeH.YinH. (2022). Probabilistic subseasonal precipitation forecasts using preceding atmospheric intraseasonal signals in a Bayesian perspective. Hydrol. Earth Syst. Sci.26, 4975–4994. doi: 10.5194/hess-26-4975-2022
-
LiebmannB.SmithC. A. (1996). Description of a complete (interpolated) outgoing longwave radiation dataset. Bull. Am. Meteorol. Soc.77, 1275–1277. Available online at: http://www.jstor.org/stable/26233278
-
LiuY.ChiangJ. C. H.ChouC.PatricolaC. M. (2014). Atmospheric teleconnection mechanisms of extratropical North Atlantic SST influence on Sahel rainfall. Clim. Dyn.43, 2797–2811. doi: 10.1007/s00382-014-2094-8
-
MariottiA.RutiP. M.RixenM. (2018). Progress in subseasonal to seasonal prediction through a joint weather and climate community effort. NPJ Clim. Atmos. Sci.1:4. doi: 10.1038/s41612-018-0014-z
-
MarquardtD. W.SneeR. D. (1975). Ridge regression in practice. Am. Stat.29, 3–20. doi: 10.1080/00031305.1975.10479105
-
Masson-DelmotteV.ZhaiP.PiraniA.ConnorsS. L.PéanC.BergerS.et al. (2021). Climate change 2021: the physical science basis. Contribution of working group I to the sixth assessment report of the intergovernmental panel on climate change. 2, 2391. doi: 10.1017/9781009157896,
-
MohinoE.JanicotS.BaderJ. (2011). Sahel rainfall and decadal to multi-decadal sea surface temperature variability. Clim. Dyn.37, 419–440. doi: 10.1007/s00382-010-0867-2
-
MonerieP.-a.BiasuttiM.MignotJ.MohinoE.PohlB.ZappaG. (2023). Storylines of Sahel precipitation change: roles of the North Atlantic and Euro-Mediterranean temperature. J. Geophys. Res. Atmos.128:e2023JD038712. doi: 10.1029/2023jd038712
-
MutangaO.AdamE.ChoM. A. (2012). High density biomass estimation for wetland vegetation using WorldView-2 imagery and random forest regression algorithm. Int. J. Appl. Earth Obs. Geoinf.18, 399–406. doi: 10.1016/j.jag.2012.03.012
-
NicholsonS. E.GristJ. P. (2003). The seasonal evolution of the atmospheric circulation over West Africa and equatorial Africa. J. Clim.16, 1013–1030. doi: 10.1175/1520-0442(2003)016<1013:TSEOTA>2.0.CO;2
-
ParkJ.-Y.BaderJ.MateiD. (2015). Northern-hemispheric differential warming is the key to understanding the discrepancies in the projected Sahel rainfall. Nat. Commun.6:5985. doi: 10.1038/ncomms6985,
-
PegionK.KirtmanB. P.BeckerE.CollinsD. C.LaJoieE.BurgmanR.et al. (2019). The subseasonal experiment (SubX): a multimodel subseasonal prediction experiment. Bull. Am. Meteorol. Soc.100, 2043–2060. doi: 10.1175/bams-d-18-0270.1
-
RobertsonA. W.VitartF.CamargoS. J. (2020). Subseasonal to seasonal prediction of weather to climate with application to tropical cyclones. J. Geophys. Res. Atmos.125:e2018JD029375. doi: 10.1029/2018jd029375
-
RodriguesL. R. L.Garcı́a-SerranoJ.Doblas-ReyesF. (2014). Seasonal forecast quality of the west African monsoon rainfall regimes by multiple forecast systems. J. Geophys. Res. Atmos.119, 7908–7930. doi: 10.1002/2013JD021316
-
RoehrigR.BouniolD.GuichardF.HourdinF.RedelspergerJ.-L. (2013). The present and future of the west African monsoon: a process-oriented assessment of CMIP5 simulations along the AMMA transect. J. Clim.26, 6471–6505. doi: 10.1175/jcli-d-12-00505.1
-
RosenblattF. (1958). The perceptron: a probabilistic model for information storage and Organization in the Brain. Psychol. Rev.65, 386–408. doi: 10.1037/h0042519,
-
RumelhartD. E.HintonG. E.WilliamsR. J. (1986). Learning representations by back-propagating errors. Nature323, 533–536. doi: 10.1038/323533a0
-
SaneY.PanthouG.BodianA.VischelT.LebelT.DacostaH.et al. (2018). Intensity–duration–frequency (IDF) rainfall curves in Senegal. Nat. Hazards Earth Syst. Sci.18, 1849–1866. doi: 10.5194/nhess-18-1849-2018
-
SarrA. B.SultanB. (2023). Predicting crop yields in Senegal using machine learning methods. Int. J. Climatol.43, 1817–1838. doi: 10.1002/joc.7947
-
SchepenA.WangQ. J.RobertsonD. (2012). Evidence for using lagged climate indices to forecast Australian seasonal rainfall. J. Clim.25, 1230–1246. doi: 10.1175/jcli-d-11-00156.1
-
SimpsonJ.AdlerR. F.NorthG. R. (1988). A proposed tropical rainfall measuring Mission (TRMM) satellite. Bull. Am. Meteorol. Soc.69, 278–295. doi: 10.1175/1520-0477(1988)069<0278:APTRMM>2.0.CO;2
-
Suárez-MorenoR.Rodrı́guez-FonsecaB. (2015). S 4 CAST V2. 0: sea surface temperature based statistical seasonal forecast model. Geosci. Model Dev.8, 3639–3658. doi: 10.5194/gmd-8-3639-2015
-
Suárez-MorenoR.Rodrı́guez-FonsecaB.BarrosoJ. A.FinkA. H. (2018). Interdecadal changes in the leading ocean forcing of Sahelian rainfall interannual variability: atmospheric dynamics and role of multidecadal SST background. J. Clim.31, 6687–6710. doi: 10.1175/JCLI-D-17-0367.1
-
SultanB.JanicotS.DiedhiouA. (2003). The West African monsoon dynamics. Part I: documentation of intraseasonal variability. J. Clim.16, 3389–3406. doi: 10.1175/1520-0442(2003)016<3389:TWAMDP>2.0.CO;2
-
SuzukiK. (2011). Artificial Neural Networks: Methodological Advances and Biomedical Applications. BoD–Books on Demand.
-
TarnavskyE.GrimesD.MaidmentR.BlackE.AllanR. P.StringerM.et al. (2014). Extension of the TAMSAT satellite-based rainfall monitoring over Africa and from 1983 to present. J. Appl. Meteorol. Climatol.53, 2805–2822. doi: 10.1175/jamc-d-14-0016.1
-
ThiamM.OrubaL.De CoetlogonG.WadeM.DiopB.FarotaA. K. (2024). Impact of the sea surface temperature in the North-eastern tropical Atlantic on precipitation over Senegal. J. Geophys. Res. Atmos.129:e2023JD040513. doi: 10.1029/2023jd040513
-
TotzS.TzipermanE.CoumouD.PfeifferK.CohenJ. (2017). Winter precipitation forecast in the European and Mediterranean regions using cluster analysis. Geophys. Res. Lett.44, 12–418. doi: 10.1002/2017GL075674,
-
ToureM.KlutseN. A. B.SarrM. A.KenneA. D.BhuiyanrM. A. E.NdiayeO.et al. (2023). A New Multiple Imputation Approach Using Machine Learning to Enhance Climate Databases in Senegal. Research Square Company. doi: 10.21203/rs.3.rs-3287168/v1
-
TrisosC. H.AdelekanI. O.TotinE.AyanladeA.EfitreJ.GemedaD. O.et al. (2022). “Africa,” in Climate Change 2022: Impacts, Adaptation and Vulnerability, eds. PörtnerH.-O.RobertsD. C.TignorM.PoloczanskaE. S.MintenbeckK.AlegríaA. (Cambridge, United Kingdom: Cambridge University Press), 1285–1455.
-
TuelA.EltahirE. A. B. (2018). Seasonal precipitation forecast over Morocco. Water Resour. Res.54, 9118–9130. doi: 10.1029/2018wr022984
-
VapnikV. N. (1998). Statistical Learning Theory. Wiley. doi: 10.1109/72.788640
-
VigaudN.GianniniA. (2019). West African convection regimes and their predictability from submonthly forecasts. Clim. Dyn.52, 7029–7048. doi: 10.1007/s00382-018-4563-y
-
VigaudN.TippettM. K.YuanJ.RobertsonA. W.AcharyaN. (2020). Spatial correction of multimodel ensemble subseasonal precipitation forecasts over North America using local Laplacian eigenfunctions. Mon. Weather Rev.148, 523–539. doi: 10.1175/mwr-d-19-0134.1
-
VincenziS.ZucchettaM.FranzoiP.PellizzatoM.PranoviF.De LeoG. A.et al. (2011). Application of a random forest algorithm to predict spatial distribution of the potential yield of Ruditapes philippinarum in the Venice lagoon, Italy. Ecol. Model.222, 1471–1478. doi: 10.1016/j.ecolmodel.2011.02.007
-
VitartF.ArdilouzeC.BonetA.BrookshawA.ChenM.CodoreanC.et al. (2017). The subseasonal to seasonal (S2S) prediction project database. Bull. Am. Meteorol. Soc.98, 163–173. doi: 10.1175/bams-d-16-0017.1
-
VitartF.RobertsonA. W. (2019). “Introduction: why sub-seasonal to seasonal prediction (S2S)?” in Sub-Seasonal to Seasonal Prediction, (Elsevier), 3–15. doi: 10.1016/B978-0-12-811714-9.00001-2
-
VitartF.RobertsonA.KumarA.HendonH.TakayaY.LinH.et al (2012). “Subseasonal to Seasonal Prediction: Research Implementation Plan.” WWRP/THORPEX-WCRP Report Geneva, Switzerland: WWRP/WCRP.
-
WangX. L.FengY.SwailV. R. (2012). North Atlantic wave height trends as reconstructed from the 20th century reanalysis. Geophys. Res. Lett.39. doi: 10.1029/2012gl053381
-
WeisbergS. (2005). Applied Linear Regression, vol. 528John Wiley & Sons. Minnesota: University of Minnesota School of Statistics Minneapolis.
-
WheelerM. C.HendonH. H. (2004). An all-season real-time multivariate MJO index: development of an index for monitoring and prediction. Mon. Weather Rev.132, 1917–1932. doi: 10.1175/1520-0493(2004)132<>2.0.co;2
-
WuM.-L. C.RealeO.SchubertS. D.SuarezM. J.KosterR. D.PegionP. J. (2009). African easterly jet: structure and maintenance. J. Clim.22, 4459–4480. doi: 10.1175/2009jcli2584.1
-
YeoI.-K.JohnsonR. A. (2000). A new family of power transformations to improve normality or symmetry. Biometrika87, 954–959. doi: 10.1093/biomet/87.4.954
-
ZhangC. (2005). Madden-Julian oscillation. Rev. Geophys.43. doi: 10.1029/2004rg000158
-
ZouH.HastieT. (2005). Regularization and variable selection via the elastic net. J. Royal Stat. Soc. Series B67, 301–320. doi: 10.1111/j.1467-9868.2005.00503.x
Summary
Keywords
subseasonal-to-seasonal prediction, West African monsoon, precipitation, machine learning, predictability, Senegal
Citation
Faye D, de Andrade FM, Suárez-Moreno R, Wane D, Hegglin MI, Dieng AL, Lguensat R and Gaye AT (2026) Data-driven approaches outperform dynamical models for subseasonal-to-seasonal precipitation forecasting in Senegal. Front. Clim. 8:1812437. doi: 10.3389/fclim.2026.1812437
Edited by
Qing Bao, Chinese Academy of Sciences (CAS), China
Updates
Check for updates
Copyright
© 2026 Faye, de Andrade, Suárez-Moreno, Wane, Hegglin, Dieng, Lguensat and Gaye.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Michaela I. Hegglin, m.i.hegglin@fz-juelich.de
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.
