Skip to content

historical_spatial_etl

Processes gridded climate data from global/regional sources (ERA5, CHIRPS, AgERA5) into spatial indicators and aggregated data.

Connectors

The connectors/ module downloads and transforms data from external sources:

Connector Source Variables Frequency
ERA5 Connector ECMWF ERA5 reanalysis Temperature, precipitation, solar radiation, humidity, wind Monthly downloads
CHIRPS Connector UCSB CHIRPS Precipitation only Monthly downloads
Local Data Connector Local/regional gridded datasets Configurable Configurable

Climate Processing

The climate_processing/ module handles data transformation:

Component Description
aggregate_daily_data.py Aggregates daily climate data to the required temporal resolution
aggregate_monthly_data.py Computes monthly aggregates from daily data
indicators_processor.py Orchestrates calculation of configured indicators per country

Tools

Tool Description
raster_clipper Clips raster data to country/region boundaries
raster_resampler Resamples rasters to match target resolution
raster_upload Uploads processed rasters to GeoServer for spatial services
download_pipeline Orchestrates download, processing, and upload workflow
config_manager Manages configuration files for data sources and parameters
file_namer Standardizes file naming according to conventions
validation_utils Validates data integrity and completeness
logging_manager Centralized logging with component tracking

Configuration

The config/ directory contains JSON configuration files:

  • chirps_config.json - CHIRPS download and processing parameters
  • clipping_config.json - Region/polygon clipping settings
  • copernicus_config.json - Copernicus/ERA5 API credentials and parameters
  • geoserver_config.json - GeoServer connection settings
  • local_data_config.json - Local data source configuration
  • naming_config.json - File naming conventions

Spatial Indicators

Calculated using a plugin-based architecture via CalculatorLoader. Currently supports 8 indicators:

Code Indicator Category Description
CDD Consecutive Dry Days Drought Maximum number of consecutive days with precipitation < 1mm
R95pTOT Very Wet Days Precipitation Total precipitation from days > 95th percentile
Rx1day Max 1-day Precipitation Precipitation Highest 1-day precipitation amount
SDII Simple Daily Intensity Index Precipitation Average precipitation on wet days (>=1mm)
TR20 Tropical Nights Temperature Days where minimum temperature > 20°C
TX10p Cool Days Temperature Days where maximum temperature < 10th percentile
TX90p Warm Days Temperature Days where maximum temperature > 90th percentile
TXX Maximum Temperature Temperature Monthly/seasonal maximum of daily maximum temperature

Each indicator uses a BaseCalculator class that reads climate data, applies the calculation formula, and stores results. The percentile_calculator.py computes the required percentiles for percentile-based indicators (R95pTOT, TX10p, TX90p).

Scheduled Execution (Crontab)

To keep spatial information up to date, the ETL is configured to run automatically each month. The scheduled execution uses a .sh script that activates the virtual or conda environment and executes the main ETL command with the corresponding dates calculated automatically.

Example of run_spatial_etl.sh:

#!/bin/bash
# Scheduled execution script for the spatial ETL
# Activates the environment and runs the pipeline with the current dates

source /path/to/env/bin/activate  # or: conda activate aclimate_v3

# Calculates the current month and the previous month
END_DATE=$(date +%Y-%m)
START_DATE=$(date -d "$(date +%Y-%m-01) -1 month" +%Y-%m)

python -m aclimate_v3_historical_spatial_etl.aclimate_run_etl \
  --country HONDURAS \
  --start_date ${START_DATE} \
  --end_date ${END_DATE} \
  --data_path /path/to/data \
  --climatology

Example crontab entry:

# Runs on the 22nd day of each month at 2:00 AM
0 2 22 * * /path/to/script/run_spatial_etl.sh

The script runs on the 22nd of each month. In addition to processing the most recent month's data, it is recommended to run it with the --climatology option to keep the climatologies up to date.

Manual Execution

To run the ETL manually with a specific date range:

python -m aclimate_v3_historical_spatial_etl.aclimate_run_etl \
  --country HONDURAS \
  --start_date 2020-01 \
  --end_date 2020-12 \
  --data_path /path/to/data \
  --climatology

Repository

For installation instructions, refer to the README in the repository: github.com/CIAT-DAPA/aclimate_v3_historical_spatial_etl