An end-to-end Data Engineering project for collecting, orchestrating, storing, processing and visualizing near-real-time and historical weather data from multiple sources.
The platform combines city-level weather observations, airport METAR reports, multi-day weather forecasts and historical climate datasets inside a reproducible local data infrastructure built with Python, PostgreSQL, Apache Airflow, Docker, Streamlit and Plotly.
This project is actively evolving. Apache Spark and additional data engineering components will be integrated in the next stage of the project.
The goal of this project is not only to display weather information, but to build a practical end-to-end data engineering environment around weather and climate data.
The platform currently supports:
- Near-real-time city weather ingestion
- Airport METAR observation ingestion
- Multi-day weather forecasts
- Historical climate data ingestion
- Multi-source ETL workflows
- Apache Airflow orchestration
- PostgreSQL relational storage
- Idempotent inserts and UPSERT operations
- Dockerized local infrastructure
- Automatic database schema initialization
- Persistent PostgreSQL and pgAdmin volumes
- Interactive Streamlit & Plotly analytics
- Logging
- Retry mechanisms
- API request caching
- Environment-based configuration
The project currently tracks 27 cities and 27 airports through its reference tables.
flowchart LR
OM["Open-Meteo API<br/>Current Weather & Forecast"]
METAR["AviationWeather<br/>METAR Observations"]
HIST["Historical Weather Sources<br/>Meteostat / Open-Meteo"]
AF["Apache Airflow<br/>Orchestration"]
PG[("PostgreSQL 15<br/>Weather Database")]
ST["Streamlit + Plotly<br/>Analytics Dashboard"]
PA["pgAdmin<br/>Database Management"]
OM --> AF
METAR --> AF
AF --> PG
HIST --> PG
PG --> ST
PG --> PA
The architecture intentionally separates the main responsibilities of the system:
Apache Airflow β Orchestration
PostgreSQL β Data Storage
Streamlit β Analytics & Presentation
pgAdmin β Database Administration
Apache Airflow orchestrates the automated near-real-time ingestion workflow.
The main DAG is:
weather_etl_pipeline
The pipeline is scheduled to run every 30 minutes.
flowchart LR
CENTER["Extract & Load<br/>City Weather"]
AIRPORT["Extract & Load<br/>Airport METAR"]
FORECAST["Extract & Load<br/>Forecast Data"]
CENTER --> FORECAST
AIRPORT --> FORECAST
The city-center and airport ingestion tasks execute independently.
After both upstream tasks finish successfully, the forecast task begins.
The DAG includes:
- Automatic retries
- Retry delays
- Structured logging
- Task dependency management
catchup=False- Production-oriented DAG tags
- Separate extraction and loading functions
The local environment contains two logically separate PostgreSQL databases with different responsibilities.
| Database | Purpose |
|---|---|
| PostgreSQL 15 | Stores weather observations, forecasts, city metadata, airport metadata and historical climate records |
| Airflow Metadata PostgreSQL | Stores DAG runs, task instances, scheduler state and internal Airflow metadata |
Airflow does not store the weather datasets in its metadata database.
Instead, Airflow DAG tasks execute Python ingestion functions which connect to the main Weather PostgreSQL database.
Current local communication flow:
Airflow Container
β
β host.docker.internal:5432
βΌ
Weather PostgreSQL 15
This allows the Astro-managed Airflow environment and the main Docker Compose application stack to remain independent while still communicating with each other.
Open-Meteo is used for city-level weather observations and forecast data.
Collected information includes:
- Temperature
- Apparent temperature
- Relative humidity
- Precipitation
- Rain
- Snowfall
- Cloud cover
- Sea-level pressure
- Wind speed
- Wind direction
- Weather codes
- Forecast data
API requests use caching and retry mechanisms to reduce unnecessary requests and improve pipeline reliability.
Airport observations are retrieved using ICAO station codes from METAR-compatible aviation weather services.
Collected information includes:
- Temperature
- Dew point
- Relative humidity
- Wind speed
- Wind direction
- Atmospheric pressure
- Precipitation
- Airport weather observations
This allows the platform to compare city-center weather observations with real airport station measurements.
Historical weather ingestion uses historical weather providers such as:
- Meteostat
- Open-Meteo historical datasets
depending on location and source availability.
Historical data is available across long time ranges for supported locations, with some datasets extending back to 1940.
The historical ETL pipeline performs extraction, normalization, basic validation and cleaning before loading records into PostgreSQL.
The main application database currently contains six core tables:
cities
airports
center_hourly_weather_data
airport_hourly_weather_data
daily_forecast_data
historical_weather
Main logical relationships:
flowchart TD
CITIES["cities"]
AIRPORTS["airports"]
CENTER["center_hourly_weather_data"]
FORECAST["daily_forecast_data"]
HIST["historical_weather"]
AIRPORT_DATA["airport_hourly_weather_data"]
CITIES --> CENTER
CITIES --> FORECAST
CITIES --> HIST
AIRPORTS --> AIRPORT_DATA
Reference metadata is stored separately from observation data.
This makes the schema easier to maintain and extend as new cities, airports and weather datasets are introduced.
The near-real-time ingestion layer uses relational constraints and conflict handling to prevent duplicate records.
City observations use:
UNIQUE (city_id, record_time)Airport observations use:
UNIQUE (airport_id, record_time)Forecast records use:
UNIQUE (city_id, forecast_date)Current observation ingestion uses:
ON CONFLICT DO NOTHINGwhile forecast ingestion uses UPSERT-style behavior:
ON CONFLICT (...)
DO UPDATEThis allows Airflow tasks to be retried without unnecessarily creating duplicate records.
Historical ingestion idempotency is planned as an upcoming improvement.
Historical weather processing is currently separated from the near-real-time Airflow DAG.
The historical pipeline follows the general flow:
Historical Weather Source
β
Chunked Extraction
β
Pandas DataFrames
β
Type Normalization
β
Sanity Checks
β
Missing Value Handling
β
Duplicate Date Removal
β
PostgreSQL
Historical data is requested in smaller time windows rather than through a single extremely large request.
This provides several advantages:
- Better API reliability
- Easier rate-limit management
- Improved error recovery
- Lower memory pressure
- Easier debugging
- Better control over partial failures
Basic sanity checks are also applied before records are persisted.
The main application infrastructure is managed using Docker Compose.
flowchart TD
COMPOSE["Docker Compose"]
DB["PostgreSQL 15<br/>weather_postgres"]
WEB["Streamlit<br/>weather_streamlit"]
ADMIN["pgAdmin<br/>weather_pgadmin"]
SCHEMA["database/schema.sql"]
SEED["database/seed.sql"]
PGVOL["pg_data"]
PAVOL["pgadmin_data"]
COMPOSE --> DB
COMPOSE --> WEB
COMPOSE --> ADMIN
SCHEMA --> DB
SEED --> DB
DB --> PGVOL
ADMIN --> PAVOL
DB -->|Healthcheck| WEB
DB -->|Healthcheck| ADMIN
The Streamlit application is built directly from the repository's Dockerfile.
PostgreSQL includes a healthcheck using:
pg_isready
and dependent services wait until the database becomes healthy before starting.
When PostgreSQL starts with a fresh Docker volume, the initialization scripts are automatically executed in order:
database/schema.sql
β
database/seed.sql
schema.sql contains the relational database structure including:
- Tables
- Sequences
- Primary keys
- Foreign keys
- Unique constraints
seed.sql initializes reference metadata for:
27 cities
27 airports
Existing PostgreSQL volumes are not reinitialized, so accumulated weather data remains intact during normal container recreation or restart operations.
Two Docker volumes are currently used:
pg_data
βββ PostgreSQL application data
pgadmin_data
βββ pgAdmin configuration
Container recreation therefore does not automatically remove stored database data.
Avoid using
docker compose down -vunless you intentionally want to delete the Docker volumes and rebuild the database from scratch.
The Streamlit dashboard reads data directly from PostgreSQL and provides interactive analytics using Plotly.
Current dashboard capabilities include:
- Global city map
- Current city weather
- Current airport METAR observations
- Multi-day forecasts
- Temperature analysis
- Precipitation analysis
- Wind analysis
- Historical climate exploration
- Monthly aggregations
- Yearly aggregations
- Long-term precipitation trends
- Airport weather analysis
- Interactive visualizations
| Category | Technologies |
|---|---|
| Language | Python |
| Orchestration | Apache Airflow |
| Airflow Runtime | Astronomer / Astro CLI |
| Database | PostgreSQL |
| Database Administration | pgAdmin |
| Containerization | Docker, Docker Compose |
| Data Processing | Pandas |
| Database Connectivity | psycopg2, SQLAlchemy |
| Dashboard | Streamlit |
| Visualization | Plotly |
| Weather APIs | Open-Meteo, AviationWeather, Meteostat |
| API Reliability | requests-cache, retry-requests |
| Configuration | python-dotenv |
| Version Control | Git, GitHub |
Weather-Data-Engineering-Platform/
β
βββ airflow/
β βββ .astro/
β β βββ config.yaml
β β
β βββ dags/
β β βββ src/
β β β βββ database.py
β β β βββ extract.py
β β β βββ load.py
β β β βββ logger.py
β β β
β β βββ weather_pipeline_dag.py
β β
β βββ .env.example
β βββ Dockerfile
β βββ requirements.txt
β βββ packages.txt
β βββ README.md
β
βββ database/
β βββ schema.sql
β βββ seed.sql
β
βββ docs/
β βββ images/
β βββ dashboard-overview.png
β βββ historical-analysis.png
β βββ airflow-dag.png
β
βββ history/
β βββ config.py
β βββ historical_etl.py
β βββ openm_rescue.py
β
βββ scripts/
β βββ check_station_inventory.py
β
βββ .dockerignore
βββ .env.example
βββ .gitignore
βββ app.py
βββ docker-compose.yml
βββ Dockerfile
βββ main.py
βββ requirements.txt
βββ README.md
Install:
- Git
- Docker Desktop
- Docker Compose
- Astronomer Astro CLI if you want to run Apache Airflow locally
git clone https://github.com/Avdatek5003/Weather_Data_Pipeline.git
cd Weather_Data_PipelineCopy:
.env.example
to:
.env
Linux / macOS:
cp .env.example .envWindows PowerShell:
Copy-Item .env.example .envThen configure your local credentials.
Example:
DB_HOST=localhost
DB_PORT=5432
DB_NAME=Weather_Data_Pipeline
DB_USER=postgres
DB_PASSWORD=your_password
RAPIDAPI_KEY=your_rapidapi_key
PGADMIN_DEFAULT_EMAIL=admin@example.com
PGADMIN_DEFAULT_PASSWORD=your_pgadmin_passwordNever commit the real .env file.
Build and start the application:
docker compose up -d --buildThis starts:
PostgreSQL β localhost:5432
pgAdmin β localhost:5050
Streamlit β localhost:8501
Open the Streamlit dashboard:
http://localhost:8501
Open pgAdmin:
http://localhost:5050
Check container status:
docker compose psThe PostgreSQL service should report:
healthy
Apache Airflow is currently managed as a separate Astro project.
Move into the Airflow directory:
cd airflowCopy the Airflow environment template.
Linux / macOS:
cp .env.example .envWindows PowerShell:
Copy-Item .env.example .envThe Airflow environment uses:
DB_HOST=host.docker.internal
DB_PORT=5432
DB_NAME=Weather_Data_Pipeline
DB_USER=postgres
DB_PASSWORD=your_passwordStart Airflow:
astro dev startOpen the Airflow UI:
http://localhost:8080
The Airflow metadata database and the Weather application database are separate.
Airflow DAG containers reach the application PostgreSQL database through:
host.docker.internal:5432
Stop the Astro environment with:
astro dev stopHistorical ingestion can be executed separately from the near-real-time Airflow workflow.
From the project root:
python history/historical_etl.pyRequired API credentials must be configured in .env.
Sensitive credentials are not intended to be committed to the repository.
The project uses environment variables for:
Database username
Database password
Database host
Database port
Database name
RapidAPI key
pgAdmin credentials
Real credentials belong in .env.
The repository contains .env.example files only to document the required configuration.
This project currently demonstrates practical implementations of:
- ETL pipeline design
- Multi-source data ingestion
- Workflow orchestration
- DAG design
- Task dependencies
- Retry strategies
- API caching
- Environment-based configuration
- Relational data modeling
- Primary keys
- Foreign keys
- Unique constraints
- SQL UPSERT operations
- Idempotent ingestion
- Docker containerization
- Docker Compose
- Service healthchecks
- Persistent Docker volumes
- Database initialization scripts
- Historical batch ingestion
- Time-series storage
- Structured logging
- Data cleaning
- Data validation
- Interactive analytics
The project is still under active development.
- Multi-source weather ingestion
- Open-Meteo current weather ingestion
- Airport METAR ingestion
- Weather forecast ingestion
- PostgreSQL relational model
- Historical climate pipeline
- Apache Airflow orchestration
- Dockerized development environment
- Automatic database initialization
- Streamlit analytics dashboard
- Persistent database infrastructure
- Historical pipeline idempotency improvements
- Apache Spark processing layer
- Automated data quality validation
- Unit testing
- Integration testing
- CI/CD with GitHub Actions
- Cloud deployment
Apache Spark is intentionally listed as a future component rather than a currently implemented technology.
The next stage of the project will introduce Spark after an appropriate processing use case has been designed for the growing historical weather datasets.
The goal is to use Spark for meaningful distributed data processing rather than adding it only as a portfolio technology.
Possible Spark responsibilities include:
Historical Weather Data
β
Apache Spark
β
Distributed Transformations
β
Aggregated / Curated Layer
β
Analytics Storage
β
Streamlit
This will allow the architecture to evolve from a traditional Python/PostgreSQL ETL pipeline toward a more scalable data-processing platform.
The core ingestion, storage, orchestration, containerization and visualization layers are operational.
Current development path:
Apache Spark Fundamentals
β
Spark Integration
β
Historical Pipeline Improvements
β
Data Quality
β
Testing
β
CI/CD
β
Project Completion
Ahmet Avdatek
Computer Engineering student focused on Data Engineering, distributed data systems and cloud data platforms.



