Skip to content

ETL de datos puntuales (Point Data ETL)

Container: Python 3.10

Processes climate data from weather stations and stores the results in the database. It is organized into the following modules:

Validation data

Verifies the quality of incoming station data. Checks that records are complete and consistent, and fills missing values using the station's climatology when available.

Responsibilities

  • Check that incoming station records are complete and consistent
  • Identify and flag invalid or out-of-range values
  • Fill missing values using the station's climatology

Key functionalities detailed:

  1. Data integrity checks — Validates completeness and consistency of each record before processing.
  2. Missing value imputation — Fills gaps using the station's climatological values when available.

CSV extract data (connector)

Reads station observations from CSV files. This is the data intake point of the pipeline, parsing the files provided by data sources and preparing the observations for processing.

Responsibilities

  • Read station observations from CSV files provided by data sources
  • Parse and standardize the data before processing
  • Support multiple data providers with a standardized format

Key functionalities detailed:

  1. File parsing — Reads CSV files and converts rows into observation records.
  2. Data preparation — Standardizes the observations for the next stages of the pipeline.

Calculate Monthly data

Aggregates daily observations into monthly values. This reduces the volume of data while preserving the patterns needed for analysis.

Responsibilities

  • Aggregate daily station observations into monthly values
  • Preserve the essential climate patterns in the aggregated data
  • Feed the aggregated data to the database

Key functionalities detailed:

  1. Temporal aggregation — Groups daily records by month and climate measure.
  2. Monthly output — Produces monthly values ready for storage and analysis.

Calculate Climatology

Computes long-term climatological averages per station and per month. These values serve as the reference baseline for the station.

Responsibilities

  • Compute long-term averages per station and per month
  • Provide the reference values used for anomaly detection
  • Store the climatological normals in the database

Key functionalities detailed:

  1. Monthly normals — Calculates the mean value for each month and measure based on historical records.
  2. Reference baseline — Serves as the baseline against which current conditions are compared.

Calculate Indicators

Generates point-based indicators from the processed station data. Applies the configured indicator formulas to produce metrics such as dry spells or accumulated precipitation.

Responsibilities

  • Apply the configured indicator formulas to the station data
  • Produce derived metrics that summarize climate patterns
  • Store the indicator results in the database

Key functionalities detailed:

  1. Indicator calculation — Computes indicators such as consecutive dry days, water stress, and accumulated precipitation.
  2. Configurable per country — Indicators are defined per country through the ORM configuration.

Main script

Coordinates the execution of the pipeline. Manages the sequence of operations from data extraction, validation, aggregation, and indicator calculation, and persists the results in the database.

Responsibilities

  • Orchestrate the execution order of all pipeline stages
  • Manage configuration and logging during execution
  • Ensure results are persisted correctly

Key functionalities detailed:

  1. Pipeline orchestration — Runs the stages in sequence: extraction, validation, aggregation, climatology, and indicators.
  2. Execution management — Configures parameters and coordinates the workflow.

For installation instructions, refer to the README in the repository: github.com/CIAT-DAPA/aclimate_v3_historical_location_etl