diff --git a/docs/CICs/PriogridCountryMapper.md b/docs/CICs/PriogridCountryMapper.md deleted file mode 100644 index e706249..0000000 --- a/docs/CICs/PriogridCountryMapper.md +++ /dev/null @@ -1,156 +0,0 @@ - -# Class Intent Contract: PriogridCountryMapper - -**Status:** Active -**Owner:** PRIO MD&D Team -**Last reviewed:** 2026-06-02 -**Related ADRs:** ADR-001, ADR-003, ADR-009 - ---- - -## 1. Purpose - -> **What is this class for?** - -`PriogridCountryMapper` is the spatial mapping engine that maps PRIO-GRID cell IDs to administrative boundary metadata (country, admin level 1, admin level 2) using largest-overlap spatial intersection against bundled shapefiles. - -It is the single authoritative source of geographic assignment in this repository. - ---- - -## 2. Non-Goals (Explicit Exclusions) - -- This class does **not** perform data pipeline orchestration or delivery -- This class does **not** fetch data from external services (Appwrite, ViewsER) -- This class does **not** manage pipeline lifecycle (read/transform/validate/save) -- This class does **not** modify or persist prediction data -- This class does **not** perform geodesic area calculations — it uses planar area on EPSG:4326 geometries -- This class does **not** handle partner-specific output formatting - ---- - -## 3. Responsibilities and Guarantees - -- Guarantees deterministic assignment: the same GID always maps to the same country/admin region given the same input shapefiles and equivalent cache state (caveat: stale disk cache from before a shapefile update violates this — see C-14) -- Guarantees that assignment uses the **largest overlap** rule: the administrative region with the highest area overlap ratio with the grid cell wins -- Guarantees that all loaded shapefiles have invalid geometries fixed via `make_valid()` and are validated for required columns and CRS at initialization -- Guarantees that spatial indices are built for all loaded GeoDataFrames to enable efficient queries -- Guarantees that cache mode (disk vs memory) does not affect mapping results — only performance -- Provides batch processing capabilities for enriching DataFrames with geographic metadata -- Provides both forward (GID → country) and reverse (country → GIDs) lookup - ---- - -## 4. Inputs and Assumptions - -- Requires 4 shapefiles at initialization: Natural Earth 10m countries, PRIO-GRID cells, GAUL 2024 L1, GAUL 2024 L2 -- Assumes all shapefiles use EPSG:4326 (WGS84) coordinate reference system -- Assumes Natural Earth data contains columns: `ISO_A3`, `NAME_EN`, `geometry` -- Assumes PRIO-GRID data contains column: `gid`, `geometry` -- Assumes GAUL data contains columns: `gaul1_code`/`gaul2_code`, `gaul1_name`/`gaul2_name`, `iso3_code`, `geometry` -- Assumes `joblib` is available for disk caching and `cachetools` for memory caching - -Assumptions that are not met **must cause failure**, not fallback behavior. - ---- - -## 5. Outputs and Side Effects - -**Outputs:** -- `find_country_for_gid(gid)` → dict with `iso_a3`, `country_name`, `overlap_ratio`, `method`, or `None` -- `find_admin1_for_gid(gid)` → dict with `gaul1_code`, `gaul1_name`, `iso3_code`, `method`, or `None` -- `find_admin2_for_gid(gid)` → dict with `gaul2_code`, `gaul2_name`, `iso3_code`, `method`, or `None` -- `find_gids_for_country(iso_a3)` → list of integer GIDs -- `enrich_dataframe_with_pg_info(df)` → DataFrame with added geographic columns - -**Side Effects:** -- Creates disk cache directory (`~/.priogrid_mapper_cache/`) if disk caching enabled -- Writes joblib cache files to disk (persistent across sessions) -- Allocates a multiprocessing Pool (cleaned up in `__del__`) -- Logs progress at INFO level during batch operations - ---- - -## 6. Failure Modes and Loudness - -- **Missing shapefiles:** `FileNotFoundError` at initialization — the mapper cannot be instantiated -- **Invalid geometries in shapefiles:** Fixed automatically via `make_valid()` at load time. If intersection still fails at runtime, logged as WARNING and the country is skipped in the overlap calculation. -- **GID not found:** Returns `None` — this is not an error (ocean cells legitimately have no country) -- **Zero-area grid geometry:** Returns `None` with WARNING log — a degenerate cell (collapsed polygon) cannot have meaningful overlap -- **Invalid point geometry:** Raises `ValueError` -- **Required columns missing from shapefiles:** Raises `ValueError` at initialization - -The following **must never** fail silently: -- Shapefile loading failures -- CRS validation failures -- Required column absence - ---- - -## 7. Boundaries and Interactions - -**Allowed interactions:** -- Reads shapefiles from the `views_postprocessing/shapefiles/` directory (Geographic Data Assets layer) -- Reads/writes to disk cache directory -- Uses `geopandas`, `shapely`, `joblib`, `cachetools` for spatial operations and caching - -**Must not depend on:** -- Pipeline Managers (`UNFAOPostProcessorManager` or any orchestration code) -- External services (Appwrite, ViewsER) -- Environment variables or runtime configuration beyond cache directory path - -This anchors the class within ADR-002 (topology): it sits at the Spatial Mapping Engine layer, above Geographic Data Assets, below Pipeline Managers. - ---- - -## 8. Examples of Correct Usage - -```python -from views_postprocessing.unfao.mapping.mapping import PriogridCountryMapper - -mapper = PriogridCountryMapper(use_disk_cache=True) - -# Single lookup -result = mapper.find_country_for_gid(148345) -# Returns: {"gid": 148345, "iso_a3": "TZA", "country_name": "Tanzania", "overlap_ratio": 0.95, "method": "largest overlap", ...} - -# DataFrame enrichment -enriched = mapper.enrich_dataframe_with_pg_info(df, pg_id_col="priogrid_gid", batch_size=1000) -``` - ---- - -## 9. Examples of Incorrect Usage - -- **Using the mapper to infer country from coordinates without going through the shapefile intersection** — this bypasses the declared assignment algorithm -- **Calling `batch_country_mapping()` with disk cache enabled** — this method only works with in-memory cache due to a known bug (C-05 in risk register) -- **Modifying `countries_gdf` or `priogrid_gdf` attributes directly** — these are loaded at initialization and treated as immutable reference data -- **Assuming `find_gids_for_country()` returns the exact inverse of `find_country_for_gid()`** — they use different threshold rules (known inconsistency, C-04 in risk register) - ---- - -## 10. Test Alignment - -- **Green tests:** Known-good GID → country mappings for unambiguous cells; schema compliance of enriched DataFrames; determinism across repeated calls -- **Beige tests:** Border cells (cells spanning 2+ countries); batch processing with mixed valid/invalid GIDs; both cache modes producing identical results -- **Red tests:** Non-existent GIDs; ocean-only cells; extreme latitude cells where planar distortion is maximal; corrupted geometries - -Currently: **no tests exist** (C-03 in risk register). This contract defines what tests must verify when written. - ---- - -## 11. Evolution Notes - -- The assignment algorithm (largest overlap) is considered **stable** — changes require a new ADR -- Cache implementation details are **evolving** — may be refactored to eliminate duplication (C-06) -- The forward/reverse lookup inconsistency (C-04) is a known defect to be resolved -- `batch_country_mapping` crash with disk cache (C-05) is a known defect to be resolved - ---- - -## End of Contract - -This document defines the **intended meaning** of `PriogridCountryMapper`. - -Changes to behavior that violate this intent are bugs. -Changes to intent must update this contract. diff --git a/docs/CICs/README.md b/docs/CICs/README.md index 1b88b25..50c2d6f 100644 --- a/docs/CICs/README.md +++ b/docs/CICs/README.md @@ -50,8 +50,9 @@ Contracts must be clear enough that: ## Active Contracts -- `PriogridCountryMapper.md` — Core spatial mapping engine (GID → country/admin boundary assignment) - `UNFAOPostProcessorManager.md` — Pipeline orchestration manager (read → transform → validate → save) +- `GaulLookupEnricher.md` — Precomputed GAUL lookup enrichment (ADR-011; replaced the runtime mapper) +- `ReconciliationModule.md` — Reconcile pgm forecasts to cm country totals (numpy-native) --- diff --git a/pyproject.toml b/pyproject.toml index a162a83..841ff2b 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -10,7 +10,6 @@ readme = "README.md" [tool.poetry.dependencies] python = ">=3.11,<3.15" views-pipeline-core = ">=2.1.3,<3.0.0" -cachetools = "==6.2.1" views-frames = ">=1.0,<2" [build-system] diff --git a/reports/technical_risk_register.md b/reports/technical_risk_register.md index 473e2b8..142b71e 100644 --- a/reports/technical_risk_register.md +++ b/reports/technical_risk_register.md @@ -5,9 +5,9 @@ | Project | views-postprocessing | | Owner | Dylan Pinheiro / PRIO MD&D Team | | Last Updated | 2026-06-24 | -| Total Concerns | 38 | -| Open Concerns | 35 | -| Resolved Concerns | 3 | +| Total Concerns | 40 | +| Open Concerns | 24 | +| Resolved Concerns | 16 | --- @@ -30,7 +30,7 @@ **Highest tier:** 1 (C-16) **Fix strategy:** Extract `CacheStrategy` interface with disk/memory implementations. Thread-lock in memory impl. Shapefile hash in disk cache keys. **Resolution scope:** Full -**⚠ CONTINGENT ON ADR-011:** If precomputed lookup table replaces mapping.py, this entire cluster is eliminated. Defer until ADR-011 is executed. +**✅ RESOLVED 2026-06-24:** ADR-011 is executed and the runtime mapper (`mapping.py`) was deleted (C-39, PR #42). The cache machinery this cluster describes no longer exists — C-05, C-06, C-14, C-16, D-01, D-02 are all resolved. ### Cluster B: Silent error hiding architecture **Root cause:** The codebase suppresses problem signals at three levels — global warning filter, DEBUG-level exception logging with `continue`, and raises without preceding logs. The impact propagates through a delivery chain with no correction mechanism. @@ -45,7 +45,7 @@ **Highest tier:** 2 (C-02) **Fix strategy:** Lazy initialization or removal of module-level call. Add `mapper` constructor parameter to manager. **Resolution scope:** Full -**⚠ CONTINGENT ON ADR-011:** If precomputed lookup table replaces mapping.py, this entire cluster is eliminated. Defer until ADR-011 is executed. +**✅ RESOLVED 2026-06-24:** ADR-011 is executed and the runtime mapper (`mapping.py`) was deleted (C-39, PR #42). The module-level `set_default_mapper()` side effect this cluster describes no longer exists — C-02 and C-10 are both resolved. ### Cluster D: Mapper-manager boundary contract **Root cause:** No explicit contract declares what columns the mapper produces and the manager consumes. @@ -53,7 +53,7 @@ **Highest tier:** 2 (C-04) **Fix strategy:** Define `ENRICHMENT_SCHEMA` constant. Harmonize forward/reverse thresholds. Add end-to-end integration test. **Resolution scope:** Full -**⚠ CONTINGENT ON ADR-011:** If precomputed lookup table replaces mapping.py, the boundary simplifies to a Parquet schema. C-17 is eliminated; C-04 is eliminated (no runtime forward/reverse divergence). +**✅ RESOLVED 2026-06-24:** ADR-011 is executed and the runtime mapper (`mapping.py`) was deleted (C-39, PR #42). The mapper/manager column boundary is now a precomputed Parquet schema (`gaul_schema.py` + `GaulLookupEnricher`); there is no runtime forward/reverse divergence — C-04 and C-17 are both resolved. ### Cluster F: CIC-code drift (documentation describes aspirational, not actual behavior) **Root cause:** CICs were written as design contracts and never validated against the code. Multiple guarantees are false. @@ -61,6 +61,7 @@ **Highest tier:** Not a code risk — documentation accuracy risk **Fix strategy:** Update CICs to describe actual code behavior. Specifically: (1) cache guarantee needs C-05 caveat, (2) return types need full key listing, (3) §6 log level should say DEBUG not WARNING, (4) ADR-008 compliance claim needs qualifying, (5) env var boundary validation claim needs qualifying. Pure documentation, no code changes. **Resolution scope:** Full +**✅ PARTIALLY RESOLVED 2026-06-24:** the `PriogridCountryMapper` CIC was deleted with the mapper (C-39, PR #42), so its drift findings (1.1, 1.2) are moot. The `UNFAOPostProcessorManager` CIC remains and is kept current (it now describes `GaulLookupEnricher`). ### Cluster E: Replace runtime mapper with precomputed lookup table **Root cause:** The area-majority algorithm is a confirmed FAO requirement (D-05 resolved), but it doesn't need 3,100 lines of geopandas runtime code — a one-time precomputation produces a ~65K-row Parquet lookup table that replaces the entire mapper with a dictionary join. @@ -68,86 +69,27 @@ **Highest tier:** 2 (C-23) **Fix strategy (revised 2026-06-12):** (1) Build the lookup by joining views-datafactory's 7 area-majority GAUL parquets (regenerated June 11, 259,200 rows each) plus the GID→lat/lon formula — the original "run the current mapper with LFS" precomputation is obsolete. (2) Replace `mapping.py` with a simple Parquet-join enricher. (3) Remove 774 MB shapefile bundle, geopandas dependency, and all cache machinery. See `docs/cross_repo_integration_report.md` and ADR-011 assessment §10. **Resolution scope:** Full — resolves Clusters A and C entirely. Eliminates C-07, C-08, C-11. Reduces Cluster B to manager-side concerns only (C-19 unfao.py raises, C-21 batch tracking, C-22 correction process). +**✅ EXECUTED 2026-06-24:** the lookup enricher shipped (`GaulLookupEnricher`, ADR-011) and the runtime mapper + shapefiles + geopandas were deleted (C-39, PR #42). The remaining open entries here are the datafactory-side area-math (C-08/C-31, tracked in views-datafactory) and the manager-side Cluster B residue (C-19/C-22) — not mapper code. --- ## Open Concerns -### C-02: Module-level side effect blocks package import on shapefile failure - -| Field | Value | -|-------|-------| -| ID | C-02 | -| Tier | 2 | -| Source | `repo-assimilation` (2026-06-02) | -| Trigger | When updating, relocating, or removing any shapefile in `views_postprocessing/shapefiles/`, verify that import of the package still succeeds | -| Location | `views_postprocessing/unfao/mapping/mapping.py:3122` | - -Line 3122 calls `set_default_mapper()` at module scope, which instantiates `PriogridCountryMapper` and loads 4 large shapefiles (Natural Earth 10m, PRIO-GRID, GAUL L1, GAUL L2). If any shapefile is missing or malformed, the import of `mapping.py` raises an exception. Because `unfao.py` imports `get_default_mapper` from this module, and the `managers/__init__.py` re-exports `UNFAOPostProcessorManager`, any consumer importing from this package will fail — even code paths that do not need the mapper. This makes the package entirely unusable if a single shapefile is corrupted. - ---- - -### C-03: Test coverage gaps across mapper-manager boundary and manager code +### C-03: Test coverage gaps in manager validation and the enrich→validate path | Field | Value | |-------|-------| | ID | C-03 | | Tier | 3 | | Source | `repo-assimilation` (2026-06-02), `test-review` (2026-06-02) | -| Trigger | When modifying the manager's `_validate()` or the mapper's `find_*` methods, verify that the test suite covers the changed behavior — integration-level coverage across the mapper-manager boundary is still absent | -| Location | `tests/test_mapping.py`, `tests/test_validation.py`, `views_postprocessing/unfao/managers/unfao.py` | - -Initial state was zero test coverage. A 73-test suite was written (2026-06-02) covering the mapper's core guarantees (determinism, largest-overlap, admin assignment, cache equivalence, GID-not-found, missing shapefile, C-05 disk-cache bug) and the validation logic (missing columns, null rejection, error messages). Remaining gaps: (1) the validation tests replicate `_validate()` logic in a standalone function because `views-pipeline-core` is unavailable in test environments — if the real `_validate()` diverges, tests pass while production fails; (2) no end-to-end test enriches through the mapper then validates through the manager; (3) the `ThreadPoolExecutor` code path is never exercised in tests; (4) no tests run against real shapefiles (Git LFS not installed). CI pytest step has been uncommented. - -Tier recalibrated from 2 to 3 during review-rr (2026-06-02): 73 tests now exist covering core mapper guarantees. The gap is maintainability (test-code divergence, missing integration path), not structural fragility. - ---- - -### C-04: Inconsistent forward/reverse mapping breaks expected bijection - -| Field | Value | -|-------|-------| -| ID | C-04 | -| Tier | 2 | -| Source | `repo-assimilation` (2026-06-02) | -| Trigger | When using `find_gids_for_country()` to enumerate cells for aggregation, verify that the result set is consistent with what `find_country_for_gid()` would assign to that country | -| Location | `views_postprocessing/unfao/mapping/mapping.py:920-922,662-668` | - -`find_country_for_gid()` assigns a GID to the country with the largest overlap regardless of magnitude (a 30% overlap wins if it's the largest). `_find_dominant_country_gids()` (used by `find_gids_for_country()`) requires `overlap_ratio > 0.5` to include a GID (line 922). A border cell with 40% overlap in country A and 35% in country B will be assigned to A by the forward lookup but excluded from `find_gids_for_country("A")`. Any downstream aggregation using the reverse lookup will silently miss cells that the enrichment step assigned to that country. - ---- +| Trigger | When modifying the manager's `_validate()` or the enricher, verify that the test suite covers the changed behavior — end-to-end coverage across the enrich→validate path is still absent | +| Location | `tests/test_validation.py`, `views_postprocessing/unfao/managers/unfao.py` | -### C-05: Multiple methods crash when disk caching is active +Initial state was zero test coverage. A 73-test suite was written (2026-06-02) covering the (now-deleted) mapper's core guarantees and the validation logic (missing columns, null rejection, error messages). Remaining gaps after the mapper removal: (1) the validation tests replicate `_validate()` logic in a standalone function because `views-pipeline-core` is unavailable in test environments — if the real `_validate()` diverges, tests pass while production fails; (2) no end-to-end test enriches through `GaulLookupEnricher` then validates through the manager. -| Field | Value | -|-------|-------| -| ID | C-05 | -| Tier | 2 | -| Source | `repo-assimilation` (2026-06-02), `graphify` (2026-06-02) | -| Trigger | When calling `batch_country_mapping()`, `batch_admin_mapping()`, `find_multiple_countries_by_iso_a3()`, `get_cache_stats()`, or `clear_cache()` on a mapper initialized with `use_disk_cache=True`, the call will raise `AttributeError` | -| Location | `views_postprocessing/unfao/mapping/mapping.py:786-787,1665-1668,2331-2332,286-294,332-344` | +Tier recalibrated from 2 to 3 during review-rr (2026-06-02): the gap is maintainability (test-code divergence, missing integration path), not structural fragility. -Multiple methods directly access `self._country_cache`, `self._admin1_cache`, and `self._admin2_cache`, which are in-memory LRU/TTL cache attributes only created when `use_disk_cache=False`. When `use_disk_cache=True`, the `__init__` method only creates `self._disk_*_cache` variants and never initializes the in-memory attributes. Affected methods: `batch_country_mapping()` (line 787), `batch_admin_mapping()` (lines 1665-1668), `find_multiple_countries_by_iso_a3()` (line 2331), `clear_cache()` (lines 286-294), and `get_cache_stats()` (lines 332-344). The default mapper is initialized with `use_disk_cache=True` (line 3102), making this the default failure mode for all these methods. - -Graphify graph traversal (2026-06-02) identified the broader scope: the same `self._*_cache` pattern appears in 6 methods beyond the originally identified `batch_country_mapping`. - -See also C-03 (no tests to catch this), C-06 (duplication is the root cause — each cache branch has its own API surface). - ---- - -### C-06: Massive code duplication across disk/memory cache branches - -| Field | Value | -|-------|-------| -| ID | C-06 | -| Tier | 3 | -| Source | `repo-assimilation` (2026-06-02) | -| Trigger | When fixing a bug in the spatial overlap logic within any `find_*` method, verify that the fix is applied to BOTH the disk-cache and memory-cache branches of that method | -| Location | `views_postprocessing/unfao/mapping/mapping.py:609-775,807-900,1130-1362,1364-1612` | - -Every method supporting both cache modes (`find_country_for_gid`, `find_admin1_for_gid`, `find_admin2_for_gid`, `find_gids_for_country`) contains the full spatial lookup implementation duplicated in both `if self.use_disk_cache` branches. The `find_country_for_gid` method alone has ~165 lines of identical logic in each branch. Total duplication accounts for approximately 1500 of the file's 3122 lines (~48%). A bug fixed in one branch may be missed in the other, creating divergent behavior depending on cache mode. - -See also C-05 (an example of this risk materializing). +**Update 2026-06-24:** narrowed with the mapper deletion (C-39, PR #42). The mapper-coverage dimension is gone with the mapper (`tests/test_mapping.py` deleted; the determinism/cache/shapefile/`ThreadPoolExecutor` gaps no longer exist). Two manager-side gaps remain: the standalone `_validate()` replica and the missing enrich→validate end-to-end test. --- @@ -195,64 +137,6 @@ The "Validate Version" step fetches the latest version from `https://pypi.org/py --- -### C-10: Manager-to-Mapper coupling via module-level global state - -| Field | Value | -|-------|-------| -| ID | C-10 | -| Tier | 3 | -| Source | `graphify` (2026-06-02) | -| Trigger | When writing unit tests for `UNFAOPostProcessorManager` that need to mock the mapper, verify that the dependency is injectable — currently it is hard-wired through the `_DEFAULT_MAPPER` global | -| Location | `views_postprocessing/unfao/mapping/mapping.py:3088-3122`, `views_postprocessing/unfao/managers/unfao.py:43` | - -Graphify graph traversal revealed that `UNFAOPostProcessorManager` (component 4) and `PriogridCountryMapper` (component 0) are in completely disconnected graph components — they share no structural edge despite being tightly coupled at runtime. The connection goes through `get_default_mapper()` which returns the module-level `_DEFAULT_MAPPER` global (line 3116). In `unfao.py` line 43, the manager calls `self._mapper = get_default_mapper()`, coupling it to a singleton created at import time rather than accepting the mapper as an injected dependency. This makes the manager untestable in isolation (our test suite had to intercept `gpd.read_file` at conftest module level to work around this), and prevents using different mapper configurations in different contexts. - -See also C-02 (the same module-level side effect that creates the global mapper). - ---- - -### C-11: PriogridCountryMapper is a god class bridging 3 architectural concerns - -| Field | Value | -|-------|-------| -| ID | C-11 | -| Tier | 3 | -| Source | `graphify` (2026-06-02) | -| Trigger | When adding a new spatial lookup method or output format, verify whether it belongs in the mapper or should be extracted into a separate class with a focused responsibility | -| Location | `views_postprocessing/unfao/mapping/mapping.py:133-3085` | - -Graphify community detection identified `PriogridCountryMapper` as the highest-betweenness node in the repository graph (0.111), with 50 edges bridging 3 distinct functional communities: Spatial Mapping Engine (core overlap logic), Country Lookup Methods (`find_country_by_iso_a3`, `search_countries_by_name`, `get_country_summary`), and Utilities & Visualization (`visualize_grid_and_country`, `calculate_capital_distance`, `cached_haversine`). The class cohesion score is 0.04, indicating its internal methods are weakly interconnected. This single 3000-line class mixes spatial intersection logic, country metadata retrieval, DataFrame enrichment, batch processing, cache management, multiprocessing orchestration, and matplotlib visualization — violating the single-responsibility principle. The god class pattern increases the cost of change: modifying visualization code risks breaking spatial assignment, and vice versa. - -Deep graph traversal (2026-06-02) quantified the structural fragility: 33.2% of all nodes are articulation points (removing any one disconnects subgraphs) and 37.7% of edges are bridges. Both metrics exceed the 30% threshold for tree-like fragility, confirming that the codebase has a star topology centered on this single class rather than a resilient mesh. - -See also C-06 (code duplication within the same class amplifies the SRP violation). - ---- - -### C-12: Silent wrong-country assignment when geometry intersection fails - -| Field | Value | -|-------|-------| -| ID | C-12 | -| Tier | 2 | -| Source | `expert-review` (2026-06-02) | -| Trigger | When Natural Earth or GAUL shapefiles contain an invalid polygon for a country that is the correct assignment for a PRIO-GRID cell, verify that the geometry error is surfaced — currently the cell is silently assigned to the next-best country | -| Location | `views_postprocessing/unfao/mapping/mapping.py:658-660` (also duplicated at ~line 1192, 1428 in admin1/admin2 branches) | - -In `find_country_for_gid()`, the overlap calculation loop (line 644-660) wraps each `country["geometry"].intersection(grid_geometry)` call in a bare `except Exception` that logs at DEBUG and skips the country with `continue`. The identical pattern exists in 7 locations across both cache branches: lines 659, 744 (country), 925 (reverse lookup), 1202, 1320 (admin1), 1436, 1562 (admin2). All 7 use `logger.debug` — invisible in production where DEBUG is disabled. The exception type is `Exception` (widest possible), catching not just `GEOSException` but also `TypeError`, `KeyError`, and any code bug. - -If the correct country's polygon has an invalid geometry, its intersection fails, it's skipped, and the cell is assigned to the next-best country. The output contains no error signal — `method` still reads `"largest overlap"`. The wrong result is cached persistently by `joblib.Memory`. Critically, this is a *correlated* failure: if country X's polygon is invalid, ALL cells overlapping country X are misassigned — not just one. The UN FAO would receive data showing country X has zero conflict predictions while neighbors show inflated numbers. The failure is both systematic and undetectable from the output. - -The mapper's own `_validate_naturalearth_data` (line 473-477) detects invalid geometries at load time and logs a WARNING, but does not call `make_valid()` or reject the data — a validate-then-proceed-anyway pattern. D-04 resolved (2026-06-02) to apply `make_valid()` at load time in all three `_load_*` methods. `make_valid()` has been applied to all three methods. - -Critically, the manager's `_validate()` method — the only safety gate before upload — is structurally incapable of catching Cluster B's primary failure mode. `_validate()` checks column presence and null counts. A misassigned cell has a valid ISO code (just the wrong one), non-null values, and all required columns present. Validation passes. The safety net has a hole shaped exactly like the failure mode it should catch: wrong-but-valid geographic assignments are invisible to every automated check in the pipeline. - -Tier recalibrated from 1 to 2 during post-campaign review-rr (2026-06-02): `make_valid()` now applied at load time mitigates the root cause. With valid geometries, intersection failures are limited to precision edge cases. The 7 DEBUG handlers remain but are defense-in-depth, not the primary failure path. Campaign Claim 2.1-2.4 confirmed correct algorithm behavior. - -See also C-08 (a different mechanism for wrong-country at high latitudes), D-05 (if mapper eliminated, this concern is moot). Contingent on ADR-011. - ---- - ### C-13: No timeout on Appwrite operations — pipeline can hang indefinitely | Field | Value | @@ -267,28 +151,6 @@ See also C-08 (a different mechanism for wrong-country at high latitudes), D-05 --- -### C-14: Stale disk cache returns outdated mappings after shapefile update - -| Field | Value | -|-------|-------| -| ID | C-14 | -| Tier | 2 | -| Source | `expert-review` (2026-06-02) | -| Trigger | When updating GAUL or Natural Earth shapefiles to a new version, verify that the disk cache at `~/.priogrid_mapper_cache/` is manually cleared — there is no automatic invalidation | -| Location | `views_postprocessing/unfao/mapping/mapping.py:55-69,183-192` | - -The `joblib.Memory` disk cache at `~/.priogrid_mapper_cache/` stores spatial mapping results keyed by function arguments (GID values), not by shapefile content or version. When shapefiles are updated (e.g., GAUL boundary revision, Natural Earth ISO code correction), the cache continues returning results computed against the old shapefiles. No staleness signal is emitted. The `clear_cache()` method exists but must be called manually — nothing in the initialization or shapefile loading path checks whether the cached results match the current shapefiles. A shapefile hash or modification timestamp in the cache key would provide automatic invalidation. - -Wrong assignments persist in three layers beyond the cache: (1) joblib disk cache, (2) local timestamped parquet files in `data_generated/`, (3) Appwrite UN FAO bucket. Correcting a discovered error requires clearing the cache, re-running the pipeline, re-uploading to Appwrite, and notifying the partner — but no documented procedure exists for identifying which uploaded files contain affected GIDs or triggering a partner-side data retraction. - -Additionally, stale cache violates the PriogridCountryMapper CIC's determinism guarantee ("the same GID always maps to the same country given the same input shapefiles"). Two environments with different cache states produce different output. - -FAO Release Note 01, Topic C confirms: "FAO will release updated GAUL boundaries annually. Only boundaries that change will be updated." This establishes an annual regeneration cadence for the precomputed lookup table (ADR-011). Partial boundary updates mean only affected cells need re-mapping, but the current cache has no mechanism to identify which cells are affected by a partial shapefile update. - -See also C-05 (cache API bugs), C-22 (no post-delivery correction process), ADR-011 (precomputed lookup replaces runtime cache). - ---- - ### C-15: Upload metadata lacks enrichment provenance and carries test description | Field | Value | @@ -307,36 +169,6 @@ See also C-14 (stale cache without version tracking), C-22 (no post-delivery cor --- -### C-16: Thread-unsafe caches under concurrent enrichment - -| Field | Value | -|-------|-------| -| ID | C-16 | -| Tier | 1 | -| Source | `expert-review` (2026-06-02) | -| Trigger | When `enrich_dataframe_with_pg_info` processes >1000 unique GIDs (triggering the `ThreadPoolExecutor` path at line 2546), concurrent threads read/write the shared `self._country_cache` LRUCache — verify thread safety of cache access | -| Location | `views_postprocessing/unfao/mapping/mapping.py:2557,697-774,1250-1362,1490-1612` | - -`enrich_dataframe_with_pg_info` dispatches concurrent threads via `ThreadPoolExecutor` (line 2557) when `use_multiprocessing=True` (the default) and `total_ids > batch_size`. Each thread calls `find_country_for_gid`, `find_admin1_for_gid`, and `find_admin2_for_gid`, all of which read and write the shared `self._country_cache`, `self._admin1_cache`, and `self._admin2_cache` instances. The `cachetools` documentation explicitly states: "all these classes are not thread-safe. Access to a shared cache from multiple threads must be properly synchronized." Concurrent cache mutations can corrupt the internal `OrderedDict`, causing `RuntimeError: dictionary changed size during iteration`, `KeyError` on entries that should exist, or silently returning wrong cached values. The race window is a classic TOCTOU: thread A checks `cache_key in cache` (True), thread B evicts that key (LRU full), thread A reads `cache[cache_key]` → `KeyError` or wrong value. - -See also C-05 (cache attribute bugs), D-02 (disagreement on fix approach). - ---- - -### C-17: Implicit column naming contract between mapper and manager - -| Field | Value | -|-------|-------| -| ID | C-17 | -| Tier | 3 | -| Source | `expert-review` (2026-06-02) | -| Trigger | When Natural Earth updates their shapefile and renames or adds columns (e.g., `ISO_A3` → `ISO_A3_EH`), verify that the mapper's 3-step column derivation still produces the names the manager expects in `filter_cols` | -| Location | `views_postprocessing/unfao/mapping/mapping.py:675,2709-2714`, `views_postprocessing/unfao/managers/unfao.py:146-158` | - -The manager expects columns like `country_iso_a3` (unfao.py:151). The mapper produces this via a 3-step chain: Natural Earth column `ISO_A3` → `find_country_for_gid` dict key `iso_a3` (mapping.py:675) → `_process_pg_batch` prepends `country_` prefix (mapping.py:2709) → `country_iso_a3`. No explicit contract declares this mapping. Furthermore, when `country_cols=None` (the default), `_process_pg_batch` (lines 2710-2714) dumps ALL shapefile columns with the `country_` prefix, making the output schema dependent on external shapefile content. A Natural Earth column rename would silently break the chain, causing `_validate()` to reject the enriched data with a confusing "missing required metadata column" error that doesn't mention the shapefile as the root cause. - ---- - ### C-19: Systematic ADR-008 non-compliance — 23 of 24 raises lack preceding log | Field | Value | @@ -351,39 +183,7 @@ ADR-008 requires structural failures to be both logged persistently AND raised e Part of Cluster B (expanded scope). ---- - -### C-20: ZeroDivisionError on degenerate zero-area cells - -| Field | Value | -|-------|-------| -| ID | C-20 | -| Tier | 3 | -| Source | `falsification-audit` (2026-06-02) | -| Trigger | When a new shapefile version introduces a near-zero-but-not-exactly-zero area cell geometry, verify the `== 0.0` guard catches it — very small positive areas bypass the guard | -| Location | `views_postprocessing/unfao/mapping/mapping.py` (7 guard sites) | - -Zero-area guards (`if grid_geometry.area == 0.0: logger.warning; return None`) were added at all 7 overlap calculation sites including `_find_dominant_country_gids`. A degenerate cell now returns `None` with a WARNING instead of silently vanishing. Remaining theoretical risk: near-zero-but-not-exactly-zero areas (e.g., 5e-31 square degrees) bypass the `== 0.0` check. This is physically impossible for real PRIO-GRID cells (~0.25 sq degrees) but could be strengthened with an epsilon-based guard. - -Tier recalibrated from 1 to 3 during review-rr (2026-06-02): all 7 sites now have guards. The concern is mitigated from "silent vanishing" to "theoretical precision edge case." - -See also C-12 (geometry intersection failures in the same catch block). - ---- - -### C-21: Batch enrichment failures produce incomplete DataFrames (observability-only fix) - -| Field | Value | -|-------|-------| -| ID | C-21 | -| Tier | 2 | -| Source | `falsification-audit` (2026-06-02) | -| Trigger | When `enrich_dataframe_with_pg_info` or `enrich_dataframe_with_country_info` encounters a batch-level processing failure, verify whether the failure is surfaced to the caller — currently failed batches are logged but not raised | -| Location | `views_postprocessing/unfao/mapping/mapping.py:2597-2600,2636-2638,2871-2874,2902-2904` | - -Four `except Exception: logger.error(...); continue` handlers drop failed batches during DataFrame enrichment in both `enrich_dataframe_with_pg_info` and `enrich_dataframe_with_country_info`. `failed_batches` counters and summary ERROR logs were added to both methods. However, the fix is observability-only: `failed_batches` and `failed_gids` are tracked and logged but NOT included in the return value or raised as an exception. The caller receives the partial DataFrame with no programmatic signal. This does not satisfy ADR-003's fail-loud requirement. - -Part of Cluster B (expanded scope). See also C-06 (duplication root cause). +**Update 2026-06-24:** the mapper portion (20 of the 23 raises, in `mapping.py`) is gone with the deleted runtime mapper (C-39); the **3 raises in `unfao.py`** remain (tracked by issue #13). Narrowed to the manager. --- @@ -653,30 +453,24 @@ Tier 2: structural fragility under the realistic change of wiring to global, wit --- -## Disagreements - -### D-01: Cache strategy refactoring — extract now vs. characterize first +### C-40: FAO delivery logic fused to pipeline-core via double inheritance + interleaved infrastructure | Field | Value | |-------|-------| -| ID | D-01 | -| Source | `expert-review` (2026-06-02) | -| Perspectives | Martin/GoF (extract Strategy pattern now), Feathers/Beck (characterize disk-cache branch with tests first) | -| Resolution | **⚠ CONTINGENT ON ADR-011.** If precomputed lookup table replaces mapping.py, no cache to refactor — this disagreement is moot. Only resolve if mapper is kept. | - ---- +| ID | C-40 | +| Tier | 2 | +| Source | `expert-code-review` (2026-06-24) | +| Trigger | When pipeline-core changes `PGMDataset` / the data loader / the postprocessor base (it is mid-migration: their #186/#188/#161), or when wanting to unit-test or numpy-ify the FAO enrich/validate without standing up the whole framework | +| Location | `views_postprocessing/unfao/managers/unfao.py:26` (double inheritance); `:78-98` and `:185-199` (inline env/AppwriteConfig/DatastoreModule); `:112-122` (`_append_metadata`), `:147-172` (`_validate`) | -### D-02: Thread safety fix — lock caches vs. remove threading +`UNFAOPostProcessorManager` subclasses **two concrete** pipeline-core base classes (`PostprocessorManager`, `ForecastingModelManager`) and **interleaves infrastructure** (env reading, `AppwriteConfig` construction, `DatastoreModule`, path resolution) with the FAO **business logic** (GAUL enrichment, the 9-column null gate) inside the lifecycle hooks. Consequences: (a) the FAO logic cannot be instantiated or unit-tested without the full framework + Appwrite env + viewser; (b) **pandas cannot leave the delivery path** because the inherited data loader and `PGMDataset` are pandas — gated on pipeline-core's own DataFrame retirement; (c) **SDP exposure** — heavy *inheritance* coupling to a pipeline-core that is itself unstable (mid-migration), so upstream changes break far from their cause (cf. C-27, C-29); (d) it's the repo's only composition-over-inheritance violation. The dependency itself is correct (`unfao.py` genuinely *is* a pipeline-core postprocessor) — the issue is its **blast radius**. Mitigation (does **not** fight the Template-Method framework): keep the subclass as a **thin shell** but extract `enrich` + `validate` + the 9-column contract into a pipeline-core-free core object the manager *calls*, and wrap the Appwrite I/O behind a small delivery-sink adapter (DIP). This makes the FAO logic testable standalone and insulates it from pipeline-core churn. -| Field | Value | -|-------|-------| -| ID | D-02 | -| Source | `expert-review` (2026-06-02) | -| Perspectives | Nygard (threading.Lock), Beck (remove ThreadPoolExecutor), Hickey (sequential is simpler), Ousterhout (keep threading for speed) | -| Resolution | **⚠ CONTINGENT ON ADR-011.** If precomputed lookup table replaces mapping.py, no threading — this disagreement is moot. Only resolve if mapper is kept. | +See also C-07/C-27/C-29 (pipeline-core coupling symptoms), C-39 (the dead-mapper cleanup that precedes any unfao restructuring). --- +## Disagreements + ### D-05: Strategic direction — eliminate runtime mapper vs keep area-based algorithm | Field | Value | @@ -741,6 +535,136 @@ Tier 2: structural fragility under the realistic change of wiring to global, wit ## Resolved Concerns +### C-10: Manager-to-Mapper coupling via module-level global state — RESOLVED + +| Field | Value | +|-------|-------| +| ID | C-10 | +| Resolved | 2026-06-24 | +| Resolution | The `PriogridCountryMapper` runtime mapper was deleted (C-39, PR #42); the module-level `_DEFAULT_MAPPER` global and `get_default_mapper()` coupling this concern describes no longer exist (the manager now uses `GaulLookupEnricher`). | + +--- + +### C-04: Inconsistent forward/reverse mapping breaks expected bijection — RESOLVED + +| Field | Value | +|-------|-------| +| ID | C-04 | +| Resolved | 2026-06-24 | +| Resolution | The `PriogridCountryMapper` runtime mapper was deleted (C-39, PR #42); the code this concern describes no longer exists. | + +--- + +### C-02: Module-level side effect blocks package import on shapefile failure — RESOLVED + +| Field | Value | +|-------|-------| +| ID | C-02 | +| Resolved | 2026-06-24 | +| Resolution | The `PriogridCountryMapper` runtime mapper was deleted (C-39, PR #42); the code this concern describes no longer exists. | + +--- + +### C-05: Multiple methods crash when disk caching is active — RESOLVED + +| Field | Value | +|-------|-------| +| ID | C-05 | +| Resolved | 2026-06-24 | +| Resolution | The `PriogridCountryMapper` runtime mapper was deleted (C-39, PR #42); the code this concern describes no longer exists. | + +--- + +### C-06: Massive code duplication across disk/memory cache branches — RESOLVED + +| Field | Value | +|-------|-------| +| ID | C-06 | +| Resolved | 2026-06-24 | +| Resolution | The `PriogridCountryMapper` runtime mapper was deleted (C-39, PR #42); the code this concern describes no longer exists. | + +--- + +### C-11: PriogridCountryMapper is a god class bridging 3 architectural concerns — RESOLVED + +| Field | Value | +|-------|-------| +| ID | C-11 | +| Resolved | 2026-06-24 | +| Resolution | The `PriogridCountryMapper` runtime mapper was deleted (C-39, PR #42); the code this concern describes no longer exists. | + +--- + +### C-12: Silent wrong-country assignment when geometry intersection fails — RESOLVED + +| Field | Value | +|-------|-------| +| ID | C-12 | +| Resolved | 2026-06-24 | +| Resolution | The `PriogridCountryMapper` runtime mapper was deleted (C-39, PR #42); the code this concern describes no longer exists. | + +--- + +### C-14: Stale disk cache returns outdated mappings after shapefile update — RESOLVED + +| Field | Value | +|-------|-------| +| ID | C-14 | +| Resolved | 2026-06-24 | +| Resolution | The `PriogridCountryMapper` runtime mapper was deleted (C-39, PR #42); the code this concern describes no longer exists. | + +--- + +### C-16: Thread-unsafe caches under concurrent enrichment — RESOLVED + +| Field | Value | +|-------|-------| +| ID | C-16 | +| Resolved | 2026-06-24 | +| Resolution | The `PriogridCountryMapper` runtime mapper was deleted (C-39, PR #42); the code this concern describes no longer exists. | + +--- + +### C-17: Implicit column naming contract between mapper and manager — RESOLVED + +| Field | Value | +|-------|-------| +| ID | C-17 | +| Resolved | 2026-06-24 | +| Resolution | The `PriogridCountryMapper` runtime mapper was deleted (C-39, PR #42); the code this concern describes no longer exists. | + +--- + +### C-20: ZeroDivisionError on degenerate zero-area cells — RESOLVED + +| Field | Value | +|-------|-------| +| ID | C-20 | +| Resolved | 2026-06-24 | +| Resolution | The `PriogridCountryMapper` runtime mapper was deleted (C-39, PR #42); the code this concern describes no longer exists. | + +--- + +### C-21: Batch enrichment failures produce incomplete DataFrames (observability-only fix) — RESOLVED + +| Field | Value | +|-------|-------| +| ID | C-21 | +| Resolved | 2026-06-24 | +| Resolution | The `PriogridCountryMapper` runtime mapper was deleted (C-39, PR #42); the code this concern describes no longer exists. | + +--- + +### C-39: Dead geopandas runtime mapper + 1.3 GB shapefiles — RESOLVED + +| Field | Value | +|-------|-------| +| ID | C-39 | +| Resolved | 2026-06-24 | +| Resolution | Deleted the dead mapper cluster on branch `chore/remove-dead-geopandas-mapper`: `unfao/mapping/` (3,171 lines), the 1.3 GB `shapefiles/` bundle, the mapper tests + `conftest.py`, the two ADR-011 diff scripts, and the CIC — **45 files / ~5,069 deletions**. **geopandas + shapely are now gone from the codebase** (zero references); `cachetools` dropped from pyproject (mapper-only). Verified: deletion broke nothing (keep-tests green). **Dissolves the old-mapper concern cluster** — C-02, C-05, C-06, C-11, C-12, C-14, C-16, C-17, C-19, C-20, C-21 and D-01, D-02 describe code that no longer exists; they are superseded by this deletion and should be relocated to Resolved in a register-curation pass. (C-08's high-latitude-area note also dies on the mapper side; its datafactory dimension stays under C-31.) | + +--- + ### C-01: Silent upload of incomplete geographic metadata to UN FAO — RESOLVED | Field | Value | @@ -773,6 +697,26 @@ Tier 2: structural fragility under the realistic change of wiring to global, wit ## Resolved Disagreements +### D-01: Cache strategy refactoring — extract now vs. characterize first — RESOLVED + +| Field | Value | +|-------|-------| +| ID | D-01 | +| Resolved | 2026-06-24 | +| Resolution | The `PriogridCountryMapper` runtime mapper was deleted (C-39, PR #42); the code this concern describes no longer exists. | + +--- + +### D-02: Thread safety fix — lock caches vs. remove threading — RESOLVED + +| Field | Value | +|-------|-------| +| ID | D-02 | +| Resolved | 2026-06-24 | +| Resolution | The `PriogridCountryMapper` runtime mapper was deleted (C-39, PR #42); the code this concern describes no longer exists. | + +--- + ### D-03: Warning suppression — RESOLVED | Field | Value | diff --git a/scripts/diff_enrichment.py b/scripts/diff_enrichment.py deleted file mode 100644 index 7763c6e..0000000 --- a/scripts/diff_enrichment.py +++ /dev/null @@ -1,164 +0,0 @@ -"""Shadow diff: current runtime mapper vs new lookup enricher (ADR-011, Stage 2 prep). - -Runs BOTH geographic-enrichment engines on the africa_me_legacy cell set and -compares all 9 contract columns per cell. Every difference is classified so we -can account, precisely, for what changes when the engine is swapped — and -confirm there are no UNEXPLAINED differences before going live. - -Requires the real shapefiles (git-lfs pull) so the current mapper can run, and -the committed lookup parquet for the new enricher. Writes: - reports/enrichment_diff/diff_cells.csv one row per africa_me cell - reports/enrichment_diff/summary.json counts per class (for plots/report) - -The classification deliberately separates "country code-system difference" -(same GAUL country, Natural Earth ISO vs GAUL ISO string) from genuine -"country reassignment" (different GAUL country), so the headline change count -is not inflated by a pure naming convention. -""" - -from __future__ import annotations - -import json -import logging -import os -import warnings -from pathlib import Path - -import numpy as np -import pandas as pd - -from views_postprocessing.unfao.enrichment import GaulLookupEnricher -from views_postprocessing.unfao.gaul_schema import METADATA_COLS - -# Quiet the noisy mapper run; the imports above are clean and stay at top. -warnings.filterwarnings("ignore") -logging.getLogger().setLevel(logging.WARNING) - - -def _resolve_datafactory() -> Path: - env = os.environ.get("VIEWS_DATAFACTORY") - if env: - return Path(env) - return Path(__file__).resolve().parents[2] / "views-datafactory" - - -_DATAFACTORY = _resolve_datafactory() -_AFRICA_ME = _DATAFACTORY / "src" / "datafactory_query" / "africa_me_legacy_pgids.json" -_AZORES_GIDS = {182470, 183190, 183909, 183910, 186058, 186778} -_OUT = Path("reports/enrichment_diff") - - -def _run_new(gids: list[int]) -> pd.DataFrame: - enr = GaulLookupEnricher() - df = pd.DataFrame({"priogrid_gid": gids, "month_id": 0}) - out = enr.enrich_dataframe_with_pg_info( - df, pg_id_col="priogrid_gid", time_id_col="month_id", - ).set_index("priogrid_gid")[METADATA_COLS] - return out.add_suffix("__new") - - -def _run_old(gids: list[int]) -> pd.DataFrame: - from views_postprocessing.unfao.mapping.mapping import get_default_mapper - mapper = get_default_mapper() - df = pd.DataFrame({"priogrid_gid": gids, "month_id": 0}) - raw = mapper.enrich_dataframe_with_pg_info( - df, pg_id_col="priogrid_gid", time_id_col="month_id", - only_metadata=True, batch_size=1000, use_multiprocessing=True, - show_progress=False, - ) - # The mapper may omit columns for unmapped cells; reindex to the contract. - raw = raw.set_index("priogrid_gid") - for c in METADATA_COLS: - if c not in raw.columns: - raw[c] = np.nan - return raw[METADATA_COLS].add_suffix("__old") - - -def _classify(row: pd.Series) -> tuple[str, str]: - """Return (presence, change_class) for one cell.""" - has_old = pd.notna(row["country_iso_a3__old"]) - has_new = pd.notna(row["country_iso_a3__new"]) - - if has_new and not has_old: - if int(row.name) in _AZORES_GIDS: - return "new_only", "azores_supplement" - return "new_only", "coastal_recovery" - if has_old and not has_new: - return "old_only", "dropped_by_land_gaul" - if not has_old and not has_new: - return "neither", "both_unmapped" - - # Present in both — compare the 9 columns. - diffs = [c for c in METADATA_COLS - if not _eq(row[f"{c}__old"], row[f"{c}__new"])] - if not diffs: - return "both", "identical" - - gaul0_changed = "admin1_gaul0_code" in diffs - iso_changed = "country_iso_a3" in diffs - admin_changed = any(c in diffs for c in - ["admin1_gaul1_code", "admin2_gaul2_code"]) - - if gaul0_changed: - return "both", "country_reassignment" - if iso_changed and not gaul0_changed: - return "both", "country_code_system" - if admin_changed: - return "both", "admin_reallocation" - # Only names/coords differ, codes identical. - if any(c in diffs for c in ["pg_xcoord", "pg_ycoord"]): - return "both", "coordinate_difference" - return "both", "name_only_difference" - - -def _eq(a, b) -> bool: - if pd.isna(a) and pd.isna(b): - return True - if pd.isna(a) or pd.isna(b): - return False - if isinstance(a, float) or isinstance(b, float): - try: - return abs(float(a) - float(b)) < 1e-6 - except (TypeError, ValueError): - return str(a) == str(b) - return str(a) == str(b) - - -def main() -> None: - gids = sorted(json.loads(_AFRICA_ME.read_text())) - print(f"africa_me_legacy cells: {len(gids):,}") - - print("running NEW enricher (lookup merge)...") - new = _run_new(gids) - print("running OLD mapper (real shapefiles, this takes a few minutes)...") - old = _run_old(gids) - - both = old.join(new, how="outer") - both.index.name = "priogrid_gid" - - classified = both.apply(_classify, axis=1, result_type="expand") - both["presence"] = classified[0] - both["change_class"] = classified[1] - - _OUT.mkdir(parents=True, exist_ok=True) - both.to_csv(_OUT / "diff_cells.csv") - - counts = both["change_class"].value_counts().to_dict() - n_changed = int((both["change_class"] != "identical").sum() - - counts.get("both_unmapped", 0)) - summary = { - "n_cells": len(both), - "counts": counts, - "n_changed_excluding_identical_and_unmapped": n_changed, - "unexplained": int(both["change_class"].isin(["UNEXPLAINED"]).sum()), - } - (_OUT / "summary.json").write_text(json.dumps(summary, indent=2)) - - print("\n=== difference classes ===") - for k, v in sorted(counts.items(), key=lambda kv: -kv[1]): - print(f" {k:28s} {v:>6,}") - print(f"\nwrote {_OUT/'diff_cells.csv'} and summary.json") - - -if __name__ == "__main__": - main() diff --git a/scripts/plot_enrichment_diff.py b/scripts/plot_enrichment_diff.py deleted file mode 100644 index 0102a9a..0000000 --- a/scripts/plot_enrichment_diff.py +++ /dev/null @@ -1,186 +0,0 @@ -"""Rich verification maps for the enrichment diff (ADR-011, Stage 2). - -Renders the old-mapper-vs-new-lookup differences over africa_me as maps, so the -change can be shown and accounted for visually. Reads -reports/enrichment_diff/diff_cells.csv and writes PNGs to -reports/enrichment_diff/maps/. - -Cells are PRIO-GRID 0.5-degree squares drawn with pcolormesh on a regular grid; -country outlines come from Natural Earth 110m (fast context layer). -""" - -from __future__ import annotations - -from pathlib import Path - -import matplotlib -matplotlib.use("Agg") -import matplotlib.colors as mcolors -import matplotlib.patches as mpatches -import matplotlib.pyplot as plt -import numpy as np -import pandas as pd - -from views_postprocessing.unfao.gaul_schema import CELL_SIZE, colrow - -_DIFF = Path("reports/enrichment_diff/diff_cells.csv") -_MAPS = Path("reports/enrichment_diff/maps") -_NE110 = Path("views_postprocessing/shapefiles/ne_110m_admin_0_countries/" - "ne_110m_admin_0_countries.shp") -_CELL = CELL_SIZE - -CLASS_COLORS = { - "identical": "#e8e8e8", - "country_reassignment": "#d62728", - "admin_reallocation": "#ff7f0e", - "coastal_recovery": "#2ca02c", - "dropped_by_land_gaul": "#9467bd", - "both_unmapped": "#7f7f7f", -} - - -def _gid_to_colrow(gid): - return colrow(gid) - - -def _borders(ax, bbox): - try: - import geopandas as gpd - ne = gpd.read_file(_NE110) - ne.boundary.plot(ax=ax, color="black", linewidth=0.4, zorder=3) - except Exception: - pass - ax.set_xlim(bbox[0], bbox[1]) - ax.set_ylim(bbox[2], bbox[3]) - ax.set_xlabel("longitude") - ax.set_ylabel("latitude") - - -def _grid(values: pd.Series, bbox): - """Build a masked 2D array over bbox at 0.5deg from a per-gid value Series.""" - col0 = int((bbox[0] + 180) / _CELL) - row0 = int((bbox[2] + 90) / _CELL) - ncol = int((bbox[1] - bbox[0]) / _CELL) - nrow = int((bbox[3] - bbox[2]) / _CELL) - grid = np.full((nrow, ncol), np.nan) - for gid, v in values.items(): - c, r = _gid_to_colrow(gid) - cc, rr = c - col0, r - row0 - if 0 <= cc < ncol and 0 <= rr < nrow: - grid[rr, cc] = v - extent = [bbox[0], bbox[1], bbox[2], bbox[3]] - return grid, extent - - -def _bbox(df, pad=2.0): - xs = [(_gid_to_colrow(g)[0]) * _CELL - 180 for g in df.index] - ys = [(_gid_to_colrow(g)[1]) * _CELL - 90 for g in df.index] - return [min(xs) - pad, max(xs) + _CELL + pad, - min(ys) - pad, max(ys) + _CELL + pad] - - -def map_agree_disagree(df, bbox): - fig, ax = plt.subplots(figsize=(11, 11)) - code = (df["change_class"] != "identical").astype(int) - grid, extent = _grid( - pd.Series(code.values, index=df.index), bbox) - cmap = mcolors.ListedColormap(["#e8e8e8", "#d62728"]) - ax.imshow(np.ma.masked_invalid(grid), origin="lower", extent=extent, - cmap=cmap, vmin=0, vmax=1, interpolation="nearest", zorder=2) - _borders(ax, bbox) - n_changed = int(code.sum()) - ax.set_title(f"africa_me: enrichment agreement\n" - f"{len(df)-n_changed:,} identical (grey) | " - f"{n_changed:,} changed (red)") - ax.legend(handles=[mpatches.Patch(color="#e8e8e8", label="identical"), - mpatches.Patch(color="#d62728", label="changed")], - loc="lower left") - fig.savefig(_MAPS / "01_agree_disagree.png", dpi=130, bbox_inches="tight") - plt.close(fig) - - -def map_by_class(df, bbox): - classes = [c for c in CLASS_COLORS if c in set(df["change_class"])] - code_of = {c: i for i, c in enumerate(classes)} - codes = df["change_class"].map(code_of) - grid, extent = _grid(pd.Series(codes.values, index=df.index), bbox) - cmap = mcolors.ListedColormap([CLASS_COLORS[c] for c in classes]) - fig, ax = plt.subplots(figsize=(11, 11)) - ax.imshow(np.ma.masked_invalid(grid), origin="lower", extent=extent, - cmap=cmap, vmin=-0.5, vmax=len(classes) - 0.5, - interpolation="nearest", zorder=2) - _borders(ax, bbox) - ax.set_title("africa_me: difference class per cell") - ax.legend(handles=[mpatches.Patch(color=CLASS_COLORS[c], - label=f"{c} ({int((df['change_class']==c).sum())})") - for c in classes], loc="lower left", fontsize=9) - fig.savefig(_MAPS / "02_by_class.png", dpi=130, bbox_inches="tight") - plt.close(fig) - - -def map_country_before_after(df, bbox): - present = df[df["country_iso_a3__old"].notna() - | df["country_iso_a3__new"].notna()] - countries = sorted(set(present["country_iso_a3__old"].dropna()) - | set(present["country_iso_a3__new"].dropna())) - cof = {c: i for i, c in enumerate(countries)} - cmap = plt.get_cmap("tab20", max(len(countries), 1)) - fig, axes = plt.subplots(1, 2, figsize=(20, 10)) - for ax, col, title in [ - (axes[0], "country_iso_a3__old", "OLD mapper (Natural Earth country)"), - (axes[1], "country_iso_a3__new", "NEW lookup (GAUL country)"), - ]: - codes = df[col].map(cof) - grid, extent = _grid(pd.Series(codes.values, index=df.index), bbox) - ax.imshow(np.ma.masked_invalid(grid), origin="lower", extent=extent, - cmap=cmap, vmin=0, vmax=len(countries), interpolation="nearest", - zorder=2) - _borders(ax, bbox) - ax.set_title(title) - fig.suptitle("Country assignment: before vs after (color = country)") - fig.savefig(_MAPS / "03_country_before_after.png", dpi=120, - bbox_inches="tight") - plt.close(fig) - - -def map_zoom(df, bbox, name, title): - sub = df.copy() - classes = [c for c in CLASS_COLORS if c in set(sub["change_class"])] - code_of = {c: i for i, c in enumerate(classes)} - grid, extent = _grid( - pd.Series(sub["change_class"].map(code_of).values, index=sub.index), bbox) - cmap = mcolors.ListedColormap([CLASS_COLORS[c] for c in classes]) - fig, ax = plt.subplots(figsize=(10, 10)) - ax.imshow(np.ma.masked_invalid(grid), origin="lower", extent=extent, - cmap=cmap, vmin=-0.5, vmax=len(classes) - 0.5, - interpolation="nearest", zorder=2, alpha=0.85) - _borders(ax, bbox) - ax.set_title(title) - ax.legend(handles=[mpatches.Patch(color=CLASS_COLORS[c], label=c) - for c in classes if c != "identical"], - loc="lower left", fontsize=9) - fig.savefig(_MAPS / name, dpi=140, bbox_inches="tight") - plt.close(fig) - - -def main(): - _MAPS.mkdir(parents=True, exist_ok=True) - df = pd.read_csv(_DIFF).set_index("priogrid_gid") - bbox = _bbox(df) - - map_agree_disagree(df, bbox) - map_by_class(df, bbox) - map_country_before_after(df, bbox) - # Zoom 1: Lesotho / South Africa enclave (biggest reassignment cluster). - map_zoom(df, [24, 36, -32, -24], "04_zoom_lesotho_sa.png", - "Zoom: Lesotho / Eswatini / South Africa border reassignments") - # Zoom 2: Namibia / Botswana / South Africa tri-border. - map_zoom(df, [13, 30, -30, -17], "05_zoom_namibia_botswana.png", - "Zoom: Namibia / Botswana / South Africa / Mozambique borders") - print(f"wrote maps to {_MAPS}/") - for p in sorted(_MAPS.glob("*.png")): - print(f" {p.name} ({p.stat().st_size//1024} KB)") - - -if __name__ == "__main__": - main() diff --git a/tests/conftest.py b/tests/conftest.py deleted file mode 100644 index 0fcb97b..0000000 --- a/tests/conftest.py +++ /dev/null @@ -1,143 +0,0 @@ -"""Test fixtures for views-postprocessing. - -Handles the module-level set_default_mapper() side effect (C-02) by -intercepting shapefile loading during import and substituting synthetic -geodata. This lets tests run without Git LFS or production shapefiles. -""" - -import pytest -import geopandas as gpd -from shapely.geometry import box -import tempfile -import shutil -import atexit -from pathlib import Path - - -# ── Synthetic shapefile data ────────────────────────────────────────────── - -_COUNTRIES = gpd.GeoDataFrame( - { - "ISO_A3": ["AAA", "BBB"], - "NAME_EN": ["CountryA", "CountryB"], - "geometry": [box(0.0, 0.0, 1.0, 2.0), box(1.0, 0.0, 2.0, 2.0)], - }, - crs="EPSG:4326", -) - -_PRIOGRID = gpd.GeoDataFrame( - { - "gid": [1, 2, 3, 4], - "xcoord": [0.25, 1.75, 0.85, 0.25], - "ycoord": [0.25, 0.25, 0.25, 1.25], - "geometry": [ - box(0.0, 0.0, 0.5, 0.5), # cell 1: fully in AAA - box(1.5, 0.0, 2.0, 0.5), # cell 2: fully in BBB - box(0.7, 0.0, 1.2, 0.5), # cell 3: 60% AAA / 40% BBB - box(0.0, 1.0, 0.5, 1.5), # cell 4: fully in AAA (upper) - ], - }, - crs="EPSG:4326", -) - -_ADMIN1 = gpd.GeoDataFrame( - { - "gaul1_code": [101, 102, 201, 202], - "gaul1_name": ["A-North", "A-South", "B-North", "B-South"], - "gaul0_code": [1, 1, 2, 2], - "gaul0_name": ["CountryA", "CountryA", "CountryB", "CountryB"], - "iso3_code": ["AAA", "AAA", "BBB", "BBB"], - "geometry": [ - box(0.0, 1.0, 1.0, 2.0), box(0.0, 0.0, 1.0, 1.0), - box(1.0, 1.0, 2.0, 2.0), box(1.0, 0.0, 2.0, 1.0), - ], - }, - crs="EPSG:4326", -) - -_ADMIN2 = gpd.GeoDataFrame( - { - "gaul2_code": [1011, 1012, 1021, 1022, 2011, 2012, 2021, 2022], - "gaul2_name": [ - "A-North-W", "A-North-E", "A-South-W", "A-South-E", - "B-North-W", "B-North-E", "B-South-W", "B-South-E", - ], - "gaul1_code": [101, 101, 102, 102, 201, 201, 202, 202], - "gaul1_name": [ - "A-North", "A-North", "A-South", "A-South", - "B-North", "B-North", "B-South", "B-South", - ], - "gaul0_code": [1, 1, 1, 1, 2, 2, 2, 2], - "gaul0_name": [ - "CountryA", "CountryA", "CountryA", "CountryA", - "CountryB", "CountryB", "CountryB", "CountryB", - ], - "iso3_code": ["AAA", "AAA", "AAA", "AAA", "BBB", "BBB", "BBB", "BBB"], - "geometry": [ - box(0.0, 1.0, 0.5, 2.0), box(0.5, 1.0, 1.0, 2.0), - box(0.0, 0.0, 0.5, 1.0), box(0.5, 0.0, 1.0, 1.0), - box(1.0, 1.0, 1.5, 2.0), box(1.5, 1.0, 2.0, 2.0), - box(1.0, 0.0, 1.5, 1.0), box(1.5, 0.0, 2.0, 1.0), - ], - }, - crs="EPSG:4326", -) - - -# ── Write synthetics to disk ───────────────────────────────────────────── - -_tmpdir = Path(tempfile.mkdtemp(prefix="views_test_")) -atexit.register(lambda: shutil.rmtree(_tmpdir, ignore_errors=True)) - -_synth_paths = {} -for _name, _gdf in [("countries", _COUNTRIES), ("priogrid", _PRIOGRID), - ("admin1", _ADMIN1), ("admin2", _ADMIN2)]: - _d = _tmpdir / _name - _d.mkdir() - _p = _d / f"{_name}.shp" - _gdf.to_file(_p) - _synth_paths[_name] = str(_p) - -_PATH_MAP = { - "ne_10m_admin_0_countries": _synth_paths["countries"], - "ne_110m_admin_0_countries": _synth_paths["countries"], - "priogrid_cellshp": _synth_paths["priogrid"], - "GAUL_2024_L1": _synth_paths["admin1"], - "GAUL_2024_L2": _synth_paths["admin2"], -} - - -# ── Permanent gpd.read_file intercept ──────────────────────────────────── -# Keep the patch active for the entire test session so that constructing -# new PriogridCountryMapper instances in fixtures also uses synthetics. - -_orig_read_file = gpd.read_file - - -def _intercepting_read_file(path, *args, **kwargs): - path_str = str(path) - for substr, synth in _PATH_MAP.items(): - if substr in path_str: - return _orig_read_file(synth, *args, **kwargs) - return _orig_read_file(path, *args, **kwargs) - - -gpd.read_file = _intercepting_read_file - -# Import the mapping module — set_default_mapper() now uses synthetics -import views_postprocessing.unfao.mapping.mapping as _mapping_mod # noqa: E402 - - -# ── Session-scoped fixtures ────────────────────────────────────────────── - -@pytest.fixture(scope="session") -def mapper(): - """PriogridCountryMapper with in-memory cache, backed by synthetic geodata.""" - return _mapping_mod.PriogridCountryMapper(use_disk_cache=False) - - -@pytest.fixture(scope="session") -def mapper_disk_cache(tmp_path_factory): - """PriogridCountryMapper with disk cache, backed by synthetic geodata.""" - cache_dir = tmp_path_factory.mktemp("disk_cache") - return _mapping_mod.PriogridCountryMapper(use_disk_cache=True, cache_dir=str(cache_dir)) diff --git a/tests/test_datafactory_deploy_readiness.py b/tests/test_datafactory_deploy_readiness.py index 5d95147..b1792b4 100644 --- a/tests/test_datafactory_deploy_readiness.py +++ b/tests/test_datafactory_deploy_readiness.py @@ -36,9 +36,17 @@ class TestReleaseGate: views-models#127's REGION='land_gaul' flip would hit a package without it. """ - # views-datafactory#224 RESOLVED 2026-06-24 (datafactory 6f7f4ec bumped to - # v1.4.0, untagged) — xfail(strict) flipped this to a pass, so it's promoted - # to a live guard: it now fails if a future version collides with a tag again. + # #224 resolved (v1.4.0), then datafactory TAGGED v1.4.0 (2026-06-24) — so + # pyproject version == latest tag, which this over-strict check reads as a + # "stale version". Re-xfail(strict): the gate flaps with datafactory's release + # cadence; it flips back to a failure (demanding promotion) only when + # datafactory bumps past the tag to the next -dev. + @pytest.mark.xfail( + reason="cross-repo deploy gate flaps with datafactory's release cadence: " + "v1.4.0 is now tagged and == pyproject version; xfail(strict) " + "re-promotes when datafactory bumps past the tag.", + strict=True, + ) def test_version_bumped_past_latest_tag(self): version_line = (_DF / "pyproject.toml").read_text() current = next( diff --git a/tests/test_integration.py b/tests/test_integration.py deleted file mode 100644 index ff8ac73..0000000 --- a/tests/test_integration.py +++ /dev/null @@ -1,140 +0,0 @@ -"""Integration tests for the mapper-to-manager boundary contract. - -Verifies that PriogridCountryMapper.enrich_dataframe_with_pg_info() produces -output compatible with UNFAOPostProcessorManager._append_metadata()'s filter_cols -and _validate()'s null checks. - -Closes: https://github.com/views-platform/views-postprocessing/issues/4 -Campaign: Claim 3.4 (FALSIFIED — no integration test existed) -""" - -import pandas as pd - -FILTER_COLS_METADATA = [ - "pg_xcoord", - "pg_ycoord", - "country_iso_a3", - "admin1_gaul1_code", - "admin1_gaul1_name", - "admin1_gaul0_code", - "admin1_gaul0_name", - "admin2_gaul2_code", - "admin2_gaul2_name", -] - -REQUIRED_METADATA_COLS = FILTER_COLS_METADATA - - -class TestMapperToManagerBoundary: - """Verify the enrichment output matches what the manager consumes.""" - - def test_all_filter_cols_present_in_enrichment(self, mapper): - df = pd.DataFrame({ - "priogrid_gid": [1, 2, 3, 4], - "month_id": [100, 100, 100, 100], - }) - enriched = mapper.enrich_dataframe_with_pg_info( - df, pg_id_col="priogrid_gid", time_id_col="month_id", - only_metadata=True, batch_size=10, - use_multiprocessing=False, show_progress=False, - ) - for col in FILTER_COLS_METADATA: - assert col in enriched.columns, ( - f"filter_cols column '{col}' missing from enrichment output. " - f"Available: {sorted(enriched.columns.tolist())}" - ) - - def test_no_nulls_in_metadata_columns_for_valid_gids(self, mapper): - df = pd.DataFrame({ - "priogrid_gid": [1, 2, 4], - "month_id": [100, 100, 100], - }) - enriched = mapper.enrich_dataframe_with_pg_info( - df, pg_id_col="priogrid_gid", time_id_col="month_id", - only_metadata=True, batch_size=10, - use_multiprocessing=False, show_progress=False, - ) - for col in FILTER_COLS_METADATA: - null_count = enriched[col].isnull().sum() - assert null_count == 0, ( - f"Column '{col}' has {null_count} nulls in enrichment output for valid GIDs" - ) - - def test_country_iso_a3_contains_valid_codes(self, mapper): - df = pd.DataFrame({ - "priogrid_gid": [1, 2, 3, 4], - "month_id": [100, 100, 100, 100], - }) - enriched = mapper.enrich_dataframe_with_pg_info( - df, pg_id_col="priogrid_gid", time_id_col="month_id", - only_metadata=True, batch_size=10, - use_multiprocessing=False, show_progress=False, - ) - valid_codes = {"AAA", "BBB"} - actual_codes = set(enriched["country_iso_a3"].dropna().unique()) - assert actual_codes.issubset(valid_codes), ( - f"Unexpected ISO codes: {actual_codes - valid_codes}" - ) - - def test_admin_codes_are_numeric(self, mapper): - df = pd.DataFrame({ - "priogrid_gid": [1, 2, 4], - "month_id": [100, 100, 100], - }) - enriched = mapper.enrich_dataframe_with_pg_info( - df, pg_id_col="priogrid_gid", time_id_col="month_id", - only_metadata=True, batch_size=10, - use_multiprocessing=False, show_progress=False, - ) - for col in ["admin1_gaul1_code", "admin1_gaul0_code", "admin2_gaul2_code"]: - assert pd.api.types.is_numeric_dtype(enriched[col]), ( - f"Column '{col}' should be numeric, got {enriched[col].dtype}" - ) - - def test_coordinates_are_float(self, mapper): - df = pd.DataFrame({ - "priogrid_gid": [1, 2], - "month_id": [100, 100], - }) - enriched = mapper.enrich_dataframe_with_pg_info( - df, pg_id_col="priogrid_gid", time_id_col="month_id", - only_metadata=True, batch_size=10, - use_multiprocessing=False, show_progress=False, - ) - for col in ["pg_xcoord", "pg_ycoord"]: - assert pd.api.types.is_float_dtype(enriched[col]), ( - f"Column '{col}' should be float, got {enriched[col].dtype}" - ) - - def test_row_count_preserved(self, mapper): - df = pd.DataFrame({ - "priogrid_gid": [1, 2, 3, 4], - "month_id": [100, 100, 100, 100], - }) - enriched = mapper.enrich_dataframe_with_pg_info( - df, pg_id_col="priogrid_gid", time_id_col="month_id", - only_metadata=True, batch_size=10, - use_multiprocessing=False, show_progress=False, - ) - assert len(enriched) == 4 - - -class TestEnrichmentThenValidation: - """Verify that mapper output passes the manager's validation logic.""" - - def test_enriched_output_passes_validation(self, mapper): - df = pd.DataFrame({ - "priogrid_gid": [1, 2, 4], - "month_id": [100, 100, 100], - }) - enriched = mapper.enrich_dataframe_with_pg_info( - df, pg_id_col="priogrid_gid", time_id_col="month_id", - only_metadata=True, batch_size=10, - use_multiprocessing=False, show_progress=False, - ) - for col in REQUIRED_METADATA_COLS: - assert col in enriched.columns, f"Validation would fail: '{col}' missing" - null_count = enriched[col].isnull().sum() - assert null_count == 0, ( - f"Validation would fail: '{col}' has {null_count} nulls" - ) diff --git a/tests/test_mapping.py b/tests/test_mapping.py deleted file mode 100644 index ff6d7ec..0000000 --- a/tests/test_mapping.py +++ /dev/null @@ -1,274 +0,0 @@ -"""Tests for PriogridCountryMapper. - -Covers CIC guarantees and failure modes using synthetic shapefiles. -See docs/CICs/PriogridCountryMapper.md for the contract under test. -""" - -import pytest -import pandas as pd -from shapely.geometry import Point - - -# --------------------------------------------------------------------------- -# GREEN: Determinism (CIC §3 — "same GID always maps to same country") -# --------------------------------------------------------------------------- - -class TestDeterminism: - def test_same_gid_returns_identical_result(self, mapper): - result1 = mapper.find_country_for_gid(1) - result2 = mapper.find_country_for_gid(1) - assert result1 == result2 - - def test_determinism_across_all_cells(self, mapper): - for gid in [1, 2, 3, 4]: - assert mapper.find_country_for_gid(gid) == mapper.find_country_for_gid(gid) - - def test_admin1_determinism(self, mapper): - for gid in [1, 2, 3, 4]: - assert mapper.find_admin1_for_gid(gid) == mapper.find_admin1_for_gid(gid) - - def test_admin2_determinism(self, mapper): - for gid in [1, 2, 3, 4]: - assert mapper.find_admin2_for_gid(gid) == mapper.find_admin2_for_gid(gid) - - -# --------------------------------------------------------------------------- -# GREEN: Largest-overlap assignment (CIC §3 — core algorithm) -# --------------------------------------------------------------------------- - -class TestLargestOverlap: - def test_cell_fully_in_country_a(self, mapper): - result = mapper.find_country_for_gid(1) - assert result is not None - assert result["iso_a3"] == "AAA" - assert result["method"] == "largest overlap" - - def test_cell_fully_in_country_b(self, mapper): - result = mapper.find_country_for_gid(2) - assert result is not None - assert result["iso_a3"] == "BBB" - - def test_border_cell_assigned_to_larger_overlap(self, mapper): - """Cell 3 spans x=[0.7, 1.2]: 0.3 in AAA, 0.2 in BBB → AAA wins.""" - result = mapper.find_country_for_gid(3) - assert result is not None - assert result["iso_a3"] == "AAA" - assert result["overlap_ratio"] > 0.5 - - def test_overlap_ratio_is_float(self, mapper): - result = mapper.find_country_for_gid(1) - assert isinstance(result["overlap_ratio"], float) - assert 0.0 < result["overlap_ratio"] <= 1.0 - - -# --------------------------------------------------------------------------- -# GREEN: Admin-level assignment (CIC §3) -# --------------------------------------------------------------------------- - -class TestAdminAssignment: - def test_admin1_for_interior_cell(self, mapper): - result = mapper.find_admin1_for_gid(1) - assert result is not None - assert result["gaul1_name"] == "A-South" - assert result["iso3_code"] == "AAA" - - def test_admin1_for_upper_cell(self, mapper): - result = mapper.find_admin1_for_gid(4) - assert result is not None - assert result["gaul1_name"] == "A-North" - - def test_admin2_for_interior_cell(self, mapper): - result = mapper.find_admin2_for_gid(1) - assert result is not None - assert result["gaul2_name"] == "A-South-W" - - def test_admin1_consistent_with_country(self, mapper): - for gid in [1, 2, 4]: - country = mapper.find_country_for_gid(gid) - admin1 = mapper.find_admin1_for_gid(gid) - if country and admin1: - assert admin1["iso3_code"] == country["iso_a3"] - - -# --------------------------------------------------------------------------- -# GREEN: Output schema for DataFrame enrichment (CIC §3, §5) -# --------------------------------------------------------------------------- - -class TestEnrichDataframe: - def test_enrichment_adds_expected_columns(self, mapper): - df = pd.DataFrame({ - "priogrid_id": [1, 2, 4], - "month_id": [100, 100, 100], - "pred_value": [0.5, 0.3, 0.1], - }) - enriched = mapper.enrich_dataframe_with_pg_info( - df, pg_id_col="priogrid_id", time_id_col="month_id", - batch_size=10, use_multiprocessing=False, show_progress=False - ) - expected_cols = [ - "pg_xcoord", "pg_ycoord", "country_iso_a3", - "admin1_gaul1_code", "admin1_gaul1_name", - "admin1_gaul0_code", "admin1_gaul0_name", - "admin2_gaul2_code", "admin2_gaul2_name", - ] - for col in expected_cols: - assert col in enriched.columns, f"Missing column: {col}" - - def test_enrichment_preserves_row_identity(self, mapper): - df = pd.DataFrame({ - "priogrid_id": [1, 2], - "month_id": [100, 100], - }) - enriched = mapper.enrich_dataframe_with_pg_info( - df, pg_id_col="priogrid_id", time_id_col="month_id", - batch_size=10, use_multiprocessing=False, show_progress=False - ) - assert list(enriched["priogrid_id"]) == [1, 2] - - def test_enrichment_row_count_preserved(self, mapper): - df = pd.DataFrame({ - "priogrid_id": [1, 2, 3, 4], - "month_id": [100, 100, 100, 100], - }) - enriched = mapper.enrich_dataframe_with_pg_info( - df, pg_id_col="priogrid_id", time_id_col="month_id", - batch_size=10, use_multiprocessing=False, show_progress=False - ) - assert len(enriched) == 4 - - -# --------------------------------------------------------------------------- -# GREEN: Cache equivalence (CIC §3 — "cache mode doesn't affect results") -# --------------------------------------------------------------------------- - -class TestCacheEquivalence: - def test_memory_and_disk_cache_produce_same_country(self, mapper, mapper_disk_cache): - for gid in [1, 2, 3, 4]: - mem_result = mapper.find_country_for_gid(gid) - disk_result = mapper_disk_cache.find_country_for_gid(gid) - if mem_result is None: - assert disk_result is None - else: - assert mem_result["iso_a3"] == disk_result["iso_a3"] - assert abs(mem_result["overlap_ratio"] - disk_result["overlap_ratio"]) < 1e-9 - - def test_memory_and_disk_cache_produce_same_admin1(self, mapper, mapper_disk_cache): - for gid in [1, 2, 4]: - mem = mapper.find_admin1_for_gid(gid) - disk = mapper_disk_cache.find_admin1_for_gid(gid) - if mem is None: - assert disk is None - else: - assert mem["gaul1_code"] == disk["gaul1_code"] - - -# --------------------------------------------------------------------------- -# RED: GID not found returns None (CIC §6) -# --------------------------------------------------------------------------- - -class TestGidNotFound: - def test_nonexistent_gid_returns_none(self, mapper): - assert mapper.find_country_for_gid(999999) is None - - def test_nonexistent_gid_admin1_returns_none(self, mapper): - assert mapper.find_admin1_for_gid(999999) is None - - def test_nonexistent_gid_admin2_returns_none(self, mapper): - assert mapper.find_admin2_for_gid(999999) is None - - -# --------------------------------------------------------------------------- -# RED: Invalid point geometry raises ValueError (CIC §6) -# --------------------------------------------------------------------------- - -class TestInvalidGeometry: - def test_invalid_point_raises(self, mapper): - from shapely import wkt - invalid_point = wkt.loads("POINT (0 0)") - invalid_point = Point(float("nan"), float("nan")) - with pytest.raises((ValueError, Exception)): - mapper.find_gid_for_point(invalid_point) - - -# --------------------------------------------------------------------------- -# RED: Missing shapefile raises at init (CIC §6) -# --------------------------------------------------------------------------- - -class TestMissingShapefile: - def test_missing_country_shapefile_raises(self): - import views_postprocessing.unfao.mapping.mapping as mod - import tempfile - import os - - tmpdir = tempfile.mkdtemp() - missing_path = os.path.join(tmpdir, "nonexistent", "missing.shp") - orig = mod.NATURAL_EARTH_COUNTRY_PATH - mod.NATURAL_EARTH_COUNTRY_PATH = missing_path - try: - with pytest.raises(Exception): - mod.PriogridCountryMapper(use_disk_cache=False) - finally: - mod.NATURAL_EARTH_COUNTRY_PATH = orig - os.rmdir(tmpdir) - - -# --------------------------------------------------------------------------- -# RED: batch_country_mapping crashes with disk cache (C-05) -# --------------------------------------------------------------------------- - -class TestBatchCountryMappingDiskCacheBug: - def test_batch_country_mapping_with_disk_cache_raises(self, mapper_disk_cache): - """Documents known bug C-05: batch_country_mapping accesses - self._country_cache which doesn't exist in disk-cache mode.""" - with pytest.raises(AttributeError): - mapper_disk_cache.batch_country_mapping([1, 2]) - - -# --------------------------------------------------------------------------- -# GREEN: Reverse lookup (find_gids_for_country) -# --------------------------------------------------------------------------- - -class TestReverseLookup: - def test_finds_gids_for_known_country(self, mapper): - gids = mapper.find_gids_for_country("AAA") - assert isinstance(gids, list) - assert len(gids) > 0 - assert 1 in gids # cell 1 is fully in AAA - - def test_unknown_country_returns_empty(self, mapper): - gids = mapper.find_gids_for_country("ZZZ") - assert gids == [] - - def test_interior_cells_in_reverse_lookup(self, mapper): - """Cells fully contained by a country should appear in its reverse lookup.""" - gids_a = mapper.find_gids_for_country("AAA") - assert 1 in gids_a # fully in AAA - assert 4 in gids_a # fully in AAA - - gids_b = mapper.find_gids_for_country("BBB") - assert 2 in gids_b # fully in BBB - - -# --------------------------------------------------------------------------- -# BEIGE: Utility methods -# --------------------------------------------------------------------------- - -class TestUtilityMethods: - def test_get_all_iso_a3_codes(self, mapper): - codes = mapper.get_all_iso_a3_codes() - assert "AAA" in codes - assert "BBB" in codes - - def test_find_country_by_iso_a3(self, mapper): - result = mapper.find_country_by_iso_a3("AAA") - assert result is not None - assert result["country_name"] == "CountryA" - - def test_find_country_by_invalid_code(self, mapper): - assert mapper.find_country_by_iso_a3("ZZ") is None # not 3 chars - - def test_find_all_admin_for_gid(self, mapper): - result = mapper.find_all_admin_for_gid(1) - assert "country" in result - assert "admin1" in result - assert "admin2" in result diff --git a/tests/test_regression.py b/tests/test_regression.py deleted file mode 100644 index 05a429e..0000000 --- a/tests/test_regression.py +++ /dev/null @@ -1,179 +0,0 @@ -"""Regression tests for code fixes applied during the audit session. - -Each test protects a specific fix. Reverting the fix causes the test to fail. - -Closes: https://github.com/views-platform/views-postprocessing/issues/6 -Closes: https://github.com/views-platform/views-postprocessing/issues/7 -Closes: https://github.com/views-platform/views-postprocessing/issues/8 -Campaign: Claims 3.1, 3.3 (test gaps for fixes) -""" - -import logging -import tempfile -from pathlib import Path - -import geopandas as gpd -import pandas as pd -from shapely.geometry import Polygon - - -# --------------------------------------------------------------------------- -# Issue #6: make_valid() preprocessing -# --------------------------------------------------------------------------- - -class TestMakeValidPreprocessing: - """Verify that invalid geometries are corrected at load time.""" - - def test_invalid_country_geometry_fixed_at_load(self): - """A self-intersecting country polygon is fixed by make_valid().""" - import views_postprocessing.unfao.mapping.mapping as mod - - bowtie = Polygon([(0, 0), (1, 1), (1, 0), (0, 1)]) - assert not bowtie.is_valid - - gdf = gpd.GeoDataFrame( - {"ISO_A3": ["TST"], "NAME_EN": ["TestCountry"], "geometry": [bowtie]}, - crs="EPSG:4326", - ) - with tempfile.TemporaryDirectory() as tmpdir: - path = Path(tmpdir) / "countries.shp" - gdf.to_file(path) - - mapper = mod.PriogridCountryMapper.__new__(mod.PriogridCountryMapper) - loaded = mapper._load_and_preprocess_naturalearth(str(path)) - - assert loaded["geometry"].is_valid.all(), "make_valid() should fix invalid geometries" - assert loaded["geometry"].iloc[0].area > 0 - - def test_invalid_admin_geometry_fixed_at_load(self): - """An invalid admin polygon is fixed by make_valid().""" - import views_postprocessing.unfao.mapping.mapping as mod - - bowtie = Polygon([(0, 0), (1, 1), (1, 0), (0, 1)]) - gdf = gpd.GeoDataFrame( - { - "gaul1_code": [1], "gaul1_name": ["Test"], - "gaul0_code": [1], "gaul0_name": ["TestC"], - "iso3_code": ["TST"], "geometry": [bowtie], - }, - crs="EPSG:4326", - ) - with tempfile.TemporaryDirectory() as tmpdir: - path = Path(tmpdir) / "admin.shp" - gdf.to_file(path) - - mapper = mod.PriogridCountryMapper.__new__(mod.PriogridCountryMapper) - loaded = mapper._load_admin_data(str(path), "admin1") - - assert loaded["geometry"].is_valid.all() - - def test_raw_invalid_geometry_is_genuinely_invalid(self): - """Proves the fixture IS invalid before make_valid — the guard is real.""" - bowtie = Polygon([(0, 0), (1, 1), (1, 0), (0, 1)]) - gdf = gpd.GeoDataFrame( - {"ISO_A3": ["TST"], "NAME_EN": ["Test"], "geometry": [bowtie]}, - crs="EPSG:4326", - ) - with tempfile.TemporaryDirectory() as tmpdir: - path = Path(tmpdir) / "raw.shp" - gdf.to_file(path) - raw = gpd.read_file(path) - - assert not raw["geometry"].is_valid.all(), ( - "Test fixture should contain invalid geometry before make_valid" - ) - - -# --------------------------------------------------------------------------- -# Issue #7: Global warning suppression must not return -# --------------------------------------------------------------------------- - -class TestWarningSuppressionRegression: - """Verify global warning suppression is absent.""" - - def test_no_global_warning_suppression(self): - """warnings.filterwarnings('ignore') must not be active globally.""" - import warnings - # Force import of the mapping module (conftest already did this) - import views_postprocessing.unfao.mapping.mapping # noqa: F401 - - global_ignores = [ - f for f in warnings.filters - if f[0] == "ignore" and f[2] is Warning - ] - assert len(global_ignores) == 0, ( - f"Global warning suppression detected: {global_ignores}. " - "Use targeted warnings.catch_warnings() instead." - ) - - def test_non_centroid_warnings_are_visible(self): - """A UserWarning should be visible — global suppression is not active.""" - import warnings - with warnings.catch_warnings(record=True) as w: - warnings.simplefilter("always") - warnings.warn("regression test canary", UserWarning) - assert len(w) == 1, ( - "UserWarning should be visible — global suppression must not be active" - ) - - -# --------------------------------------------------------------------------- -# Issue #8: Zero-area grid geometry guard -# --------------------------------------------------------------------------- - -class TestZeroAreaGuardRegression: - """Verify that degenerate zero-area cells are handled with WARNING, not silently.""" - - def test_zero_area_country_returns_none_with_warning(self, mapper, caplog): - """A zero-area cell should return None with a WARNING log.""" - degenerate = Polygon([(0.25, 0.25), (0.5, 0.25), (0.25, 0.25)]) - assert degenerate.area == 0.0 - - mapper.priogrid_gdf = pd.concat([ - mapper.priogrid_gdf, - gpd.GeoDataFrame( - {"gid": [99], "xcoord": [0.3], "ycoord": [0.25], "geometry": [degenerate]}, - crs="EPSG:4326", - ), - ], ignore_index=True) - mapper.priogrid_gdf["centroid"] = mapper.priogrid_gdf["geometry"].centroid - - with caplog.at_level(logging.WARNING): - result = mapper.find_country_for_gid(99) - - assert result is None, "Zero-area cell should return None" - assert any("Zero-area" in r.message and "99" in r.message for r in caplog.records), ( - "Zero-area cell should produce a WARNING log mentioning the GID" - ) - - def test_zero_area_admin1_returns_none(self, mapper): - """Zero-area cell returns None for admin1 lookup too.""" - degenerate = Polygon([(0.25, 0.25), (0.5, 0.25), (0.25, 0.25)]) - if 99 not in mapper.priogrid_gdf["gid"].values: - mapper.priogrid_gdf = pd.concat([ - mapper.priogrid_gdf, - gpd.GeoDataFrame( - {"gid": [99], "xcoord": [0.3], "ycoord": [0.25], "geometry": [degenerate]}, - crs="EPSG:4326", - ), - ], ignore_index=True) - mapper.priogrid_gdf["centroid"] = mapper.priogrid_gdf["geometry"].centroid - - result = mapper.find_admin1_for_gid(99) - assert result is None - - def test_zero_area_admin2_returns_none(self, mapper): - """Zero-area cell returns None for admin2 lookup too.""" - degenerate = Polygon([(0.25, 0.25), (0.5, 0.25), (0.25, 0.25)]) - if 99 not in mapper.priogrid_gdf["gid"].values: - mapper.priogrid_gdf = pd.concat([ - mapper.priogrid_gdf, - gpd.GeoDataFrame( - {"gid": [99], "xcoord": [0.3], "ycoord": [0.25], "geometry": [degenerate]}, - crs="EPSG:4326", - ), - ], ignore_index=True) - mapper.priogrid_gdf["centroid"] = mapper.priogrid_gdf["geometry"].centroid - - result = mapper.find_admin2_for_gid(99) - assert result is None diff --git a/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.cpg b/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.cpg deleted file mode 100644 index 105bf61..0000000 --- a/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.cpg +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:3ad3031f5503a4404af825262ee8232cc04d4ea6683d42c5dd0a2f2a27ac9824 -size 5 diff --git a/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.dbf b/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.dbf deleted file mode 100644 index c7b2d2f..0000000 --- a/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.dbf +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:e036df5fb3a5ee64ec2ea17a640fac406ecf0cb908f79f41c7208b431ba6ca44 -size 1533520 diff --git a/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.prj b/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.prj deleted file mode 100644 index 2098c4d..0000000 --- a/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.prj +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:a02a27b1d1982c8516d83398e85a3c8b1aef1713c13ef4d84d7bde17430c07c4 -size 145 diff --git a/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.sbn b/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.sbn deleted file mode 100644 index 3a09541..0000000 Binary files a/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.sbn and /dev/null differ diff --git a/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.sbx b/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.sbx deleted file mode 100644 index ae3c383..0000000 Binary files a/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.sbx and /dev/null differ diff --git a/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.shp b/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.shp deleted file mode 100644 index cba612e..0000000 --- a/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.shp +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:dc7ec18e5f64c6a6591d88ac60be15a60e3a0bcae57cb489a91108b674c33392 -size 451373368 diff --git a/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.shp.xml b/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.shp.xml deleted file mode 100644 index c7d69d9..0000000 --- a/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.shp.xml +++ /dev/null @@ -1,2 +0,0 @@ - -20250620145811001.0FALSEUpdateSchema "CIMDATA=<CIMStandardDataConnection xsi:type='typens:CIMStandardDataConnection' xmlns:xsi='http://www.w3.org/2001/XMLSchema-instance' xmlns:xs='http://www.w3.org/2001/XMLSchema' xmlns:typens='http://www.esri.com/schemas/ArcGIS/3.4.0'><WorkspaceConnectionString>DATABASE=W:\Gaul_2024_Shapefiles\Admin0_Error\Revised_June25\OUTput\New Folder</WorkspaceConnectionString><WorkspaceFactory>Shapefile</WorkspaceFactory><Dataset>GAUL_2024_L1.shp</Dataset><DatasetType>esriDTFeatureClass</DatasetType></CIMStandardDataConnection>" <operationSequence><workflow><DeleteField><field_name>gaul2_code</field_name><field_name>gaul2_name</field_name></DeleteField></workflow></operationSequence>CalculateField GAUL_2024_L1 disp_en !gaul1_name!+", "+!gaul0_name! PYTHON3 # TEXT NO_ENFORCE_DOMAINSExportFeatures GAUL_2024_L1 W:\Gaul_2024_Shapefiles\Admin0_Error\Output\GAUL_2024_L1.shp # NOT_USE_ALIAS "iso3_code "iso3_code" true true false 3 Text 0 0,First,#,GAUL_2024_L1,iso3_code,0,2;map_code "map_code" true true false 3 Text 0 0,First,#,GAUL_2024_L1,map_code,0,2;gaul0_code "gaul0_code" true true false 10 Long 0 10,First,#,GAUL_2024_L1,gaul0_code,-1,-1;gaul0_name "gaul0_name" true true false 100 Text 0 0,First,#,GAUL_2024_L1,gaul0_name,0,99;gaul1_code "gaul1_code" true true false 10 Long 0 10,First,#,GAUL_2024_L1,gaul1_code,-1,-1;gaul1_name "gaul1_name" true true false 100 Text 0 0,First,#,GAUL_2024_L1,gaul1_name,0,99;continent "continent" true true false 12 Text 0 0,First,#,GAUL_2024_L1,continent,0,11;disp_en "disp_en" true true false 254 Text 0 0,First,#,GAUL_2024_L1,disp_en,0,253" #ExportFeatures GAUL_2024_L1 "W:\GAUL_2025\GAUL_2024\July25\GAUL_2024_L1\New Folder\GAUL_2024_L1.shp" # NOT_USE_ALIAS "iso3_code "iso3_code" true true false 3 Text 0 0,First,#,GAUL_2024_L1,iso3_code,0,2;map_code "map_code" true true false 3 Text 0 0,First,#,GAUL_2024_L1,map_code,0,2;gaul0_code "gaul0_code" true true false 10 Long 0 10,First,#,GAUL_2024_L1,gaul0_code,-1,-1;gaul0_name "gaul0_name" true true false 100 Text 0 0,First,#,GAUL_2024_L1,gaul0_name,0,99;gaul1_code "gaul1_code" true true false 10 Long 0 10,First,#,GAUL_2024_L1,gaul1_code,-1,-1;gaul1_name "gaul1_name" true true false 100 Text 0 0,First,#,GAUL_2024_L1,gaul1_name,0,99;continent "continent" true true false 12 Text 0 0,First,#,GAUL_2024_L1,continent,0,11;disp_en "disp_en" true true false 254 Text 0 0,First,#,GAUL_2024_L1,disp_en,0,253" #GAUL_2024_L10020.000file://\\LAPTOP-7171PVOF\W$\GAUL_2025\GAUL_2024\July25\GAUL_2024_L1\New Folder\GAUL_2024_L1.shpLocal Area NetworkGeographicGCS_WGS_1984Angular Unit: Degree (0.017453)<GeographicCoordinateSystem xsi:type='typens:GeographicCoordinateSystem' xmlns:xsi='http://www.w3.org/2001/XMLSchema-instance' xmlns:xs='http://www.w3.org/2001/XMLSchema' xmlns:typens='http://www.esri.com/schemas/ArcGIS/3.4.0'><WKT>GEOGCS[&quot;GCS_WGS_1984&quot;,DATUM[&quot;D_WGS_1984&quot;,SPHEROID[&quot;WGS_1984&quot;,6378137.0,298.257223563]],PRIMEM[&quot;Greenwich&quot;,0.0],UNIT[&quot;Degree&quot;,0.0174532925199433],AUTHORITY[&quot;EPSG&quot;,4326]]</WKT><XOrigin>-400</XOrigin><YOrigin>-400</YOrigin><XYScale>11258999068426.238</XYScale><ZOrigin>-100000</ZOrigin><ZScale>10000</ZScale><MOrigin>-100000</MOrigin><MScale>10000</MScale><XYTolerance>8.983152841195215e-09</XYTolerance><ZTolerance>0.001</ZTolerance><MTolerance>0.001</MTolerance><HighPrecision>true</HighPrecision><LeftLongitude>-180</LeftLongitude><WKID>4326</WKID><LatestWKID>4326</LatestWKID></GeographicCoordinateSystem>20250701164720002025070116472000Microsoft Windows 10 Version 10.0 (Build 26100) ; Esri ArcGIS 13.4.0.55405GAUL_2024_L1Shapefile0.000datasetEPSG6.2(3.0.1)0SimpleFALSE0FALSEFALSEGAUL_2024_L1Feature Class0FIDFIDOID400Internal feature number.EsriSequential unique whole numbers that are automatically generated.ShapeShapeGeometry000Feature geometry.EsriCoordinates defining the features.iso3_codeiso3_codeString300map_codemap_codeString300gaul0_codegaul0_codeInteger10100gaul0_namegaul0_nameString10000gaul1_codegaul1_codeInteger10100gaul1_namegaul1_nameString10000continentcontinentString1200disp_endisp_enString2540020250701 diff --git a/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.shx b/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.shx deleted file mode 100644 index a61bf85..0000000 --- a/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1.shx +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:9ca5f533c9888e7140f409b0659ef06f24d5443a09365fb99542fc95f7936558 -size 24980 diff --git a/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1_ChangeLog.docx b/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1_ChangeLog.docx deleted file mode 100644 index bf46d4a..0000000 Binary files a/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1_ChangeLog.docx and /dev/null differ diff --git a/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1_Metadata.xlsx b/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1_Metadata.xlsx deleted file mode 100644 index 2937626..0000000 Binary files a/views_postprocessing/shapefiles/GAUL_2024_L1/GAUL_2024_L1_Metadata.xlsx and /dev/null differ diff --git a/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.cpg b/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.cpg deleted file mode 100644 index 105bf61..0000000 --- a/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.cpg +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:3ad3031f5503a4404af825262ee8232cc04d4ea6683d42c5dd0a2f2a27ac9824 -size 5 diff --git a/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.dbf b/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.dbf deleted file mode 100644 index 0a43ebd..0000000 --- a/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.dbf +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:4a69f4e7d30439006c27d718037db7dd3889e49227f7ca42fda0c85cbd0e18de -size 27451326 diff --git a/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.prj b/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.prj deleted file mode 100644 index 2098c4d..0000000 --- a/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.prj +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:a02a27b1d1982c8516d83398e85a3c8b1aef1713c13ef4d84d7bde17430c07c4 -size 145 diff --git a/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.sbn b/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.sbn deleted file mode 100644 index bfac991..0000000 Binary files a/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.sbn and /dev/null differ diff --git a/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.sbx b/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.sbx deleted file mode 100644 index a7c6eca..0000000 Binary files a/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.sbx and /dev/null differ diff --git a/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.shp b/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.shp deleted file mode 100644 index f50aaac..0000000 --- a/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.shp +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:c150365c48c26648f2b07c7322db589594c465aefcf1d44e843e46d46fd44872 -size 774793140 diff --git a/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.shp.xml b/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.shp.xml deleted file mode 100644 index 89916ff..0000000 --- a/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.shp.xml +++ /dev/null @@ -1,2 +0,0 @@ - -20250701164247001.0FALSEGAUL_2024_L20020.000file://\\LAPTOP-7171PVOF\W$\GAUL_2025\GAUL_2024\July25\GAUL_2024_L2\New Folder\GAUL_2024_L2.shpLocal Area NetworkGeographicGCS_WGS_1984Angular Unit: Degree (0.017453)<GeographicCoordinateSystem xsi:type='typens:GeographicCoordinateSystem' xmlns:xsi='http://www.w3.org/2001/XMLSchema-instance' xmlns:xs='http://www.w3.org/2001/XMLSchema' xmlns:typens='http://www.esri.com/schemas/ArcGIS/3.4.0'><WKT>GEOGCS[&quot;GCS_WGS_1984&quot;,DATUM[&quot;D_WGS_1984&quot;,SPHEROID[&quot;WGS_1984&quot;,6378137.0,298.257223563]],PRIMEM[&quot;Greenwich&quot;,0.0],UNIT[&quot;Degree&quot;,0.0174532925199433],AUTHORITY[&quot;EPSG&quot;,4326]]</WKT><XOrigin>-400</XOrigin><YOrigin>-400</YOrigin><XYScale>11258999068426.238</XYScale><ZOrigin>-100000</ZOrigin><ZScale>10000</ZScale><MOrigin>-100000</MOrigin><MScale>10000</MScale><XYTolerance>8.983152841195215e-09</XYTolerance><ZTolerance>0.001</ZTolerance><MTolerance>0.001</MTolerance><HighPrecision>true</HighPrecision><LeftLongitude>-180</LeftLongitude><WKID>4326</WKID><LatestWKID>4326</LatestWKID></GeographicCoordinateSystem>ExportFeatures GAUL_2024_L2 "W:\GAUL_2025\GAUL_2024\July25\GAUL_2024_L2\New Folder\GAUL_2024_L2.shp" # NOT_USE_ALIAS "iso3_code "iso3_code" true true false 3 Text 0 0,First,#,GAUL_2024_L2,iso3_code,0,2;map_code "map_code" true true false 3 Text 0 0,First,#,GAUL_2024_L2,map_code,0,2;gaul0_code "gaul0_code" true true false 10 Long 0 10,First,#,GAUL_2024_L2,gaul0_code,-1,-1;gaul0_name "gaul0_name" true true false 100 Text 0 0,First,#,GAUL_2024_L2,gaul0_name,0,99;gaul1_code "gaul1_code" true true false 10 Long 0 10,First,#,GAUL_2024_L2,gaul1_code,-1,-1;gaul1_name "gaul1_name" true true false 100 Text 0 0,First,#,GAUL_2024_L2,gaul1_name,0,99;gaul2_code "gaul2_code" true true false 10 Long 0 10,First,#,GAUL_2024_L2,gaul2_code,-1,-1;gaul2_name "gaul2_name" true true false 100 Text 0 0,First,#,GAUL_2024_L2,gaul2_name,0,99;continent "continent" true true false 12 Text 0 0,First,#,GAUL_2024_L2,continent,0,11;disp_en "disp_en" true true false 254 Text 0 0,First,#,GAUL_2024_L2,disp_en,0,253" #20250701164247002025070116424700Microsoft Windows 10 Version 10.0 (Build 26100) ; Esri ArcGIS 13.4.0.55405GAUL_2024_L2Shapefile0.000datasetEPSG6.2(3.0.1)0SimpleFALSE0FALSEFALSEGAUL_2024_L2Feature Class0FIDFIDOID400Internal feature number.EsriSequential unique whole numbers that are automatically generated.ShapeShapeGeometry000Feature geometry.EsriCoordinates defining the features.iso3_codeiso3_codeString300map_codemap_codeString300gaul0_codegaul0_codeInteger10100gaul0_namegaul0_nameString10000gaul1_codegaul1_codeInteger10100gaul1_namegaul1_nameString10000gaul2_codegaul2_codeInteger10100gaul2_namegaul2_nameString10000continentcontinentString1200disp_endisp_enString2540020250701 diff --git a/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.shx b/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.shx deleted file mode 100644 index cd01932..0000000 --- a/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2.shx +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:f31ef2e9380774399a9b7b168837225290eeadd60e1d184da13b468672c6bb63 -size 364292 diff --git a/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2_ChangeLog.docx b/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2_ChangeLog.docx deleted file mode 100644 index 57fae4f..0000000 Binary files a/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2_ChangeLog.docx and /dev/null differ diff --git a/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2_Metadata.xlsx b/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2_Metadata.xlsx deleted file mode 100644 index 10d9288..0000000 Binary files a/views_postprocessing/shapefiles/GAUL_2024_L2/GAUL_2024_L2_Metadata.xlsx and /dev/null differ diff --git a/views_postprocessing/shapefiles/ne_10m_admin_0_countries/ne_10m_admin_0_countries.cpg b/views_postprocessing/shapefiles/ne_10m_admin_0_countries/ne_10m_admin_0_countries.cpg deleted file mode 100644 index 105bf61..0000000 --- a/views_postprocessing/shapefiles/ne_10m_admin_0_countries/ne_10m_admin_0_countries.cpg +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:3ad3031f5503a4404af825262ee8232cc04d4ea6683d42c5dd0a2f2a27ac9824 -size 5 diff --git a/views_postprocessing/shapefiles/ne_10m_admin_0_countries/ne_10m_admin_0_countries.dbf b/views_postprocessing/shapefiles/ne_10m_admin_0_countries/ne_10m_admin_0_countries.dbf deleted file mode 100644 index b36e8ca..0000000 --- a/views_postprocessing/shapefiles/ne_10m_admin_0_countries/ne_10m_admin_0_countries.dbf +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:c5dbd3dd5fd7e2ef49051fc88562c03819e8ea63a382642df6eadd1243bf4b49 -size 878482 diff --git a/views_postprocessing/shapefiles/ne_10m_admin_0_countries/ne_10m_admin_0_countries.prj b/views_postprocessing/shapefiles/ne_10m_admin_0_countries/ne_10m_admin_0_countries.prj deleted file mode 100644 index 2098c4d..0000000 --- a/views_postprocessing/shapefiles/ne_10m_admin_0_countries/ne_10m_admin_0_countries.prj +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:a02a27b1d1982c8516d83398e85a3c8b1aef1713c13ef4d84d7bde17430c07c4 -size 145 diff --git a/views_postprocessing/shapefiles/ne_10m_admin_0_countries/ne_10m_admin_0_countries.shp b/views_postprocessing/shapefiles/ne_10m_admin_0_countries/ne_10m_admin_0_countries.shp deleted file mode 100644 index 29356e6..0000000 --- a/views_postprocessing/shapefiles/ne_10m_admin_0_countries/ne_10m_admin_0_countries.shp +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:7ce119ef6342e43cff7c0c3004e0911ab7ec1988a14734372031d2012180e7bc -size 8806224 diff --git a/views_postprocessing/shapefiles/ne_10m_admin_0_countries/ne_10m_admin_0_countries.shx b/views_postprocessing/shapefiles/ne_10m_admin_0_countries/ne_10m_admin_0_countries.shx deleted file mode 100644 index 1fef3c4..0000000 --- a/views_postprocessing/shapefiles/ne_10m_admin_0_countries/ne_10m_admin_0_countries.shx +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:ca19ec112d054c77bc8f7ac00e3b110d5dff32cc9bcf4cd1b8b66bdd0f611d32 -size 2164 diff --git a/views_postprocessing/shapefiles/ne_110m_admin_0_countries/ne_110m_admin_0_countries.cpg b/views_postprocessing/shapefiles/ne_110m_admin_0_countries/ne_110m_admin_0_countries.cpg deleted file mode 100644 index 105bf61..0000000 --- a/views_postprocessing/shapefiles/ne_110m_admin_0_countries/ne_110m_admin_0_countries.cpg +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:3ad3031f5503a4404af825262ee8232cc04d4ea6683d42c5dd0a2f2a27ac9824 -size 5 diff --git a/views_postprocessing/shapefiles/ne_110m_admin_0_countries/ne_110m_admin_0_countries.dbf b/views_postprocessing/shapefiles/ne_110m_admin_0_countries/ne_110m_admin_0_countries.dbf deleted file mode 100755 index 85e0a7d..0000000 --- a/views_postprocessing/shapefiles/ne_110m_admin_0_countries/ne_110m_admin_0_countries.dbf +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:1fee677cd4e03b367876e03861eb10197e4022a846bf92060e0313432863785b -size 531808 diff --git a/views_postprocessing/shapefiles/ne_110m_admin_0_countries/ne_110m_admin_0_countries.prj b/views_postprocessing/shapefiles/ne_110m_admin_0_countries/ne_110m_admin_0_countries.prj deleted file mode 100755 index 92bc4cc..0000000 --- a/views_postprocessing/shapefiles/ne_110m_admin_0_countries/ne_110m_admin_0_countries.prj +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:3259f0e55290a82b1350646f604e8a7ee1e2136c0320a40fad838ab40819fff8 -size 147 diff --git a/views_postprocessing/shapefiles/ne_110m_admin_0_countries/ne_110m_admin_0_countries.shp b/views_postprocessing/shapefiles/ne_110m_admin_0_countries/ne_110m_admin_0_countries.shp deleted file mode 100755 index 6fded2d..0000000 --- a/views_postprocessing/shapefiles/ne_110m_admin_0_countries/ne_110m_admin_0_countries.shp +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:08e341606e8391e458c3f08deb312de664b56bfae376064c5aa0aee6681a5f55 -size 180924 diff --git a/views_postprocessing/shapefiles/ne_110m_admin_0_countries/ne_110m_admin_0_countries.shx b/views_postprocessing/shapefiles/ne_110m_admin_0_countries/ne_110m_admin_0_countries.shx deleted file mode 100755 index a96b56d..0000000 --- a/views_postprocessing/shapefiles/ne_110m_admin_0_countries/ne_110m_admin_0_countries.shx +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:8b0be2ad97dd484aee5c2ebc98697d5372e832b8ae58a35a661aeef6b985668d -size 1516 diff --git a/views_postprocessing/shapefiles/priogrid_cellshp/priogrid_cell.dbf b/views_postprocessing/shapefiles/priogrid_cellshp/priogrid_cell.dbf deleted file mode 100644 index 5fb0751..0000000 --- a/views_postprocessing/shapefiles/priogrid_cellshp/priogrid_cell.dbf +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:af378cd89a2784a6b6d250e0d264f64d3402b58f9e68d3e1f87e79b1ce0d5902 -size 20476994 diff --git a/views_postprocessing/shapefiles/priogrid_cellshp/priogrid_cell.prj b/views_postprocessing/shapefiles/priogrid_cellshp/priogrid_cell.prj deleted file mode 100644 index 195a501..0000000 --- a/views_postprocessing/shapefiles/priogrid_cellshp/priogrid_cell.prj +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:98aaf3d1c0ecadf1a424a4536de261c3daf4e373697cb86c40c43b989daf52eb -size 143 diff --git a/views_postprocessing/shapefiles/priogrid_cellshp/priogrid_cell.qpj b/views_postprocessing/shapefiles/priogrid_cellshp/priogrid_cell.qpj deleted file mode 100644 index edb0142..0000000 --- a/views_postprocessing/shapefiles/priogrid_cellshp/priogrid_cell.qpj +++ /dev/null @@ -1 +0,0 @@ -GEOGCS["WGS 84",DATUM["WGS_1984",SPHEROID["WGS 84",6378137,298.257223563,AUTHORITY["EPSG","7030"]],TOWGS84[0,0,0,0,0,0,0],AUTHORITY["EPSG","6326"]],PRIMEM["Greenwich",0,AUTHORITY["EPSG","8901"]],UNIT["degree",0.0174532925199433,AUTHORITY["EPSG","9108"]],AUTHORITY["EPSG","4326"]] diff --git a/views_postprocessing/shapefiles/priogrid_cellshp/priogrid_cell.shp b/views_postprocessing/shapefiles/priogrid_cellshp/priogrid_cell.shp deleted file mode 100644 index 58fca9c..0000000 --- a/views_postprocessing/shapefiles/priogrid_cellshp/priogrid_cell.shp +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:c61a5bce9d35b8d1568a601be7367204a36c2d84ccd9caf3f87b9542c1305bab -size 35251300 diff --git a/views_postprocessing/shapefiles/priogrid_cellshp/priogrid_cell.shx b/views_postprocessing/shapefiles/priogrid_cellshp/priogrid_cell.shx deleted file mode 100644 index 642c7cc..0000000 --- a/views_postprocessing/shapefiles/priogrid_cellshp/priogrid_cell.shx +++ /dev/null @@ -1,3 +0,0 @@ -version https://git-lfs.github.com/spec/v1 -oid sha256:b7e37604cc28f71e20369d4e27b411289575afc44012a1417fd4025cb56659b3 -size 2073700 diff --git a/views_postprocessing/unfao/managers/README.md b/views_postprocessing/unfao/managers/README.md index 60a9948..59156ad 100644 --- a/views_postprocessing/unfao/managers/README.md +++ b/views_postprocessing/unfao/managers/README.md @@ -384,5 +384,4 @@ except Exception as e: ## See Also -- [GaulLookupEnricher](../../../docs/CICs/GaulLookupEnricher.md) - Precomputed-lookup enrichment (ADR-011) -- [PriogridCountryMapper](../mapping/README.md) - Legacy runtime mapper (retained, no longer used by the manager) \ No newline at end of file +- [GaulLookupEnricher](../../../docs/CICs/GaulLookupEnricher.md) - Precomputed-lookup enrichment (ADR-011) \ No newline at end of file diff --git a/views_postprocessing/unfao/mapping/README.md b/views_postprocessing/unfao/mapping/README.md deleted file mode 100644 index a754de8..0000000 --- a/views_postprocessing/unfao/mapping/README.md +++ /dev/null @@ -1,579 +0,0 @@ -# PRIO-GRID Spatial Mapping Module - -## Overview - -The `mapping.py` module provides a comprehensive spatial mapping solution for the VIEWS platform. It enables bidirectional mapping between **PRIO-GRID cells** (a standardized global grid system with ~50×50 km cells) and administrative boundaries at multiple levels: - -- **Country level** (using Natural Earth data) -- **Admin Level 1** (GAUL 2024 - provinces/states) -- **Admin Level 2** (GAUL 2024 - districts/counties) - -This module is critical for the UN FAO postprocessing pipeline, where conflict predictions made at the PRIO-GRID level need to be aggregated and reported at various administrative boundary levels. - ---- - -## Table of Contents - -1. [Key Concepts](#key-concepts) -2. [Data Sources](#data-sources) -3. [Classes](#classes) -4. [Mapping Decision Logic](#mapping-decision-logic) -5. [Caching Architecture](#caching-architecture) -6. [Key Methods](#key-methods) -7. [Usage Examples](#usage-examples) -8. [Performance Considerations](#performance-considerations) -9. [Configuration](#configuration) - ---- - -## Key Concepts - -### PRIO-GRID -The **Peace Research Institute Oslo GRID** (PRIO-GRID) is a standardized global grid system that divides the world into cells of approximately 0.5° × 0.5° (roughly 50×50 km at the equator). Each cell is identified by a unique **GID** (Grid ID). This grid provides a consistent spatial unit for conflict research and prediction. - -### Administrative Boundaries -- **Country (GAUL Level 0)**: National boundaries using ISO A3 codes (e.g., 'USA', 'TZA', 'FRA') -- **Admin Level 1 (GAUL Level 1)**: First-level subdivisions (states, provinces, regions) -- **Admin Level 2 (GAUL Level 2)**: Second-level subdivisions (counties, districts) - -### The Mapping Challenge -PRIO-GRID cells do not align perfectly with administrative boundaries. A single grid cell may: -- Fall entirely within one country/admin region -- Span multiple countries (e.g., border regions) -- Partially overlap water bodies -- Have its centroid in a different region than its majority area - -This module implements sophisticated decision rules to handle these edge cases consistently. - ---- - -## Data Sources - -The module relies on the following shapefiles (paths configured relative to the module location): - -| Data Source | Path | Description | -|------------|------|-------------| -| Natural Earth | `shapefiles/ne_110m_admin_0_countries/` | Country boundaries (110m resolution) | -| PRIO-GRID | `shapefiles/priogrid_cellshp/` | Global grid cell definitions | -| GAUL Admin 1 | `shapefiles/GAUL_2024_L1/` | First-level administrative boundaries | -| GAUL Admin 2 | `shapefiles/GAUL_2024_L2/` | Second-level administrative boundaries | - -All shapefiles use **EPSG:4326** (WGS84) coordinate reference system. - ---- - -## Classes - -### `DistanceCache` - -A simple LRU (Least Recently Used) cache for distance calculations. - -```python -class DistanceCache: - def __init__(self, maxsize=1000) - def get(self, key) -> Optional[float] - def set(self, key, value) -> None -``` - -**Purpose**: Caches haversine distance calculations to avoid redundant computations when calculating distances between grid cells and capital cities. - ---- - -### `PriogridCountryMapper` - -The main class that handles all spatial mapping operations. - -#### Initialization - -```python -mapper = PriogridCountryMapper( - use_disk_cache=True, # Enable persistent disk caching - cache_dir=None, # Custom cache directory (default: ~/.priogrid_mapper_cache) - cache_ttl=None # Cache time-to-live in seconds (None = no expiration) -) -``` - -#### Attributes - -| Attribute | Type | Description | -|-----------|------|-------------| -| `countries_gdf` | GeoDataFrame | Natural Earth country boundaries | -| `priogrid_gdf` | GeoDataFrame | PRIO-GRID cell geometries with centroids | -| `admin1_gdf` | GeoDataFrame | GAUL Level 1 boundaries | -| `admin2_gdf` | GeoDataFrame | GAUL Level 2 boundaries | -| `priogrid_sindex` | SpatialIndex | R-tree index for fast PRIO-GRID queries | -| `admin1_sindex` | SpatialIndex | R-tree index for fast Admin1 queries | -| `admin2_sindex` | SpatialIndex | R-tree index for fast Admin2 queries | - ---- - -## Mapping Decision Logic - -### Country Assignment for a PRIO-GRID Cell - -The `find_country_for_gid()` method uses a simple, deterministic **largest overlap** approach: - -``` -┌─────────────────────────────────────────────────────────────┐ -│ INPUT: PRIO-GRID GID │ -└─────────────────────────────────────────────────────────────┘ - │ - ▼ -┌─────────────────────────────────────────────────────────────┐ -│ Step 1: Find all countries that intersect the grid cell │ -│ using spatial index for performance │ -└─────────────────────────────────────────────────────────────┘ - │ - ▼ -┌─────────────────────────────────────────────────────────────┐ -│ Step 2: Calculate overlap ratio for each country │ -│ overlap_ratio = intersection_area / cell_area │ -└─────────────────────────────────────────────────────────────┘ - │ - ▼ -┌─────────────────────────────────────────────────────────────┐ -│ Step 3: Sort by overlap ratio (descending) │ -└─────────────────────────────────────────────────────────────┘ - │ - ▼ -┌─────────────────────────────────────────────────────────────┐ -│ Step 4: Assign to country with LARGEST OVERLAP │ -│ Method tag: "largest overlap" │ -└─────────────────────────────────────────────────────────────┘ -``` - -#### Why Largest Overlap? - -The decision to use **only the largest overlap rule** (rather than a multi-rule hierarchy) provides several benefits: - -1. **Simplicity** - One consistent rule is easier to understand, explain, and maintain -2. **Determinism** - The same input always produces the same output, regardless of centroid position -3. **Intuitive for aggregation** - When aggregating predictions, assigning to the region with the most area coverage makes sense (more of the cell's "land" belongs there) -4. **Handles all edge cases** - Works correctly for: - - Cells entirely within one country (100% overlap) - - Border cells spanning multiple countries - - Coastal cells with partial water coverage (assigns to the country with most land area) - -### Admin Level Assignment - -The `find_admin1_for_gid()` and `find_admin2_for_gid()` methods use the same largest overlap approach: - -1. **First determine the country** - Uses `find_country_for_gid()` to ensure admin regions are filtered to the correct country -2. **Calculate overlaps** - For all admin regions in that country that intersect the grid cell -3. **Assign to largest overlap** - The admin region with the highest overlap ratio wins - -This approach ensures that: -- Admin regions are always consistent with the assigned country -- The assignment method is consistent across all administrative levels -- Results are deterministic and reproducible - ---- - -## Caching Architecture - -The module implements a sophisticated multi-layer caching system to optimize performance: - -### Cache Types - -``` -┌─────────────────────────────────────────────────────────────┐ -│ CACHING ARCHITECTURE │ -└─────────────────────────────────────────────────────────────┘ - -┌─────────────────────┐ ┌─────────────────────┐ -│ DISK CACHE │ │ MEMORY CACHE │ -│ (Persistent) │ │ (Session-based) │ -├─────────────────────┤ ├─────────────────────┤ -│ • Survives restarts │ │ • Faster access │ -│ • Shared across │ │ • Limited size │ -│ processes │ │ • LRU or TTL-based │ -│ • Uses joblib │ │ • Uses cachetools │ -│ Memory │ │ │ -└─────────────────────┘ └─────────────────────┘ - -Default location: ~/.priogrid_mapper_cache/ -├── country/ # GID → Country mappings -├── admin1/ # GID → Admin1 mappings -├── admin2/ # GID → Admin2 mappings -└── gid/ # Point → GID mappings -``` - -### Cache Configuration - -| Cache Type | Max Size | Purpose | -|------------|----------|---------| -| Country Cache | 100,000 | GID → Country mappings | -| GID Cache | 100,000 | Point → GID lookups | -| GIDs for Country | 10,000,000 | Country → [GIDs] reverse mapping | -| Admin1 Cache | 100,000 | GID → Admin1 mappings | -| Admin2 Cache | 100,000 | GID → Admin2 mappings | - -### Cache Methods - -```python -# Clear all or specific caches -mapper.clear_cache(cache_type="all") # "all", "country", "admin1", "admin2", "gid" - -# Get cache statistics -stats = mapper.get_cache_stats() - -# Pre-warm cache for frequently accessed data -mapper.warm_cache(gid_list=[123, 456, 789], iso_a3_list=['TZA', 'KEN']) -``` - ---- - -## Key Methods - -### Core Mapping Methods - -#### `find_country_for_gid(gid: int) -> dict` - -Find country information for a PRIO-GRID cell. - -**Returns:** -```python -{ - "gid": 123456, - "iso_a3": "TZA", - "country_name": "Tanzania", - "overlap_ratio": 0.85, - "method": "largest overlap", - # ... additional Natural Earth fields -} -``` - -#### `find_admin1_for_gid(gid: int) -> dict` - -Find Admin Level 1 region for a PRIO-GRID cell. - -**Returns:** -```python -{ - "gid": 123456, - "gaul1_code": 45678, - "gaul1_name": "Dar es Salaam", - "iso3_code": "TZA", - "method": "largest overlap", - "gaul0_code": 123, - "gaul0_name": "Tanzania", - # ... additional GAUL fields -} -``` - -#### `find_admin2_for_gid(gid: int) -> dict` - -Find Admin Level 2 region for a PRIO-GRID cell. - -**Returns:** -```python -{ - "gid": 123456, - "gaul2_code": 78901, - "gaul2_name": "Kinondoni", - "iso3_code": "TZA", - "method": "largest overlap", - "gaul1_code": 45678, - "gaul1_name": "Dar es Salaam", - # ... additional GAUL fields -} -``` - -#### `find_gids_for_country(iso_a3: str) -> List[int]` - -Find all PRIO-GRID cells belonging to a country (reverse mapping). - -```python -gids = mapper.find_gids_for_country("TZA") -# Returns: [123456, 123457, 123458, ...] -``` - -### Batch Processing Methods - -#### `batch_country_mapping(gid_list: List[int]) -> pd.DataFrame` - -Process multiple GIDs with cache optimization. - -```python -df = mapper.batch_country_mapping([123, 456, 789, ...]) -``` - -#### `batch_country_mapping_parallel(gid_list, batch_size=1000) -> pd.DataFrame` - -Process large datasets using multiprocessing. - -```python -df = mapper.batch_country_mapping_parallel(gid_list, batch_size=1000) -``` - -#### `batch_admin_mapping(gid_list, admin_level="admin1") -> pd.DataFrame` - -Batch mapping for admin regions. - -```python -df = mapper.batch_admin_mapping(gid_list, admin_level="admin2") -``` - -### DataFrame Enrichment Methods - -#### `enrich_dataframe_with_pg_info(df, ...)` - -Enrich a DataFrame containing PRIO-GRID IDs with geographic metadata. - -```python -enriched_df = mapper.enrich_dataframe_with_pg_info( - df, - pg_id_col="priogrid_id", - time_id_col="month_id", - include_country=True, - include_admin1=True, - include_admin2=True, - include_pg_info=True, - batch_size=1000, - use_multiprocessing=True, - show_progress=True -) -``` - -**Result columns added:** -- `country_iso_a3`, `country_name`, `country_overlap_ratio`, `country_method` -- `admin1_gaul1_code`, `admin1_gaul1_name`, `admin1_method` -- `admin2_gaul2_code`, `admin2_gaul2_name`, `admin2_method` -- `pg_xcoord`, `pg_ycoord`, `pg_geometry`, etc. - -### Visualization Methods - -#### `visualize_grid_and_country(gid: int)` - -Visualize a PRIO-GRID cell and its assigned country. - -```python -mapper.visualize_grid_and_country(123456) -``` - -#### `visualize_grid_and_admin(gid, admin_level="admin1", show_all_admins=False)` - -Visualize a PRIO-GRID cell with admin boundaries. - -```python -mapper.visualize_grid_and_admin(123456, admin_level="admin2", show_all_admins=True) -``` - -### Utility Methods - -#### `find_gid_for_point(point_geometry) -> int` - -Find which PRIO-GRID cell contains a point. - -```python -from shapely.geometry import Point -gid = mapper.find_gid_for_point(Point(36.8, -1.3)) # Nairobi -``` - -#### `find_country_by_iso_a3(iso_a3: str) -> dict` - -Get detailed country information by ISO code. - -```python -info = mapper.find_country_by_iso_a3("TZA") -``` - -#### `search_countries_by_name(name_pattern, exact_match=False) -> List[dict]` - -Search for countries by name pattern. - -```python -results = mapper.search_countries_by_name("tanzania") -results = mapper.search_countries_by_name("united", exact_match=False) -``` - ---- - -## Usage Examples - -### Basic Usage - -```python -from views_postprocessing.unfao.mapping.mapping import PriogridCountryMapper - -# Initialize mapper (loads shapefiles, builds spatial indices) -mapper = PriogridCountryMapper(use_disk_cache=True) - -# Map a single PRIO-GRID cell to country -result = mapper.find_country_for_gid(148345) -print(f"GID 148345 belongs to {result['country_name']} ({result['iso_a3']})") -print(f"Assignment method: {result['method']}") - -# Get all admin levels at once -all_info = mapper.find_all_admin_for_gid(148345) -``` - -### Batch Processing for Large Datasets - -```python -import pandas as pd - -# Your prediction data with PRIO-GRID IDs -predictions_df = pd.DataFrame({ - 'priogrid_id': [148345, 148346, 148347, ...], - 'month_id': [553, 553, 553, ...], - 'predicted_fatalities': [0.5, 1.2, 0.3, ...] -}) - -# Enrich with geographic metadata -enriched_df = mapper.enrich_dataframe_with_pg_info( - predictions_df, - pg_id_col='priogrid_id', - include_country=True, - include_admin1=True, - include_admin2=True, - use_multiprocessing=True, - show_progress=True -) - -# Now you can aggregate by country, admin1, or admin2 -country_totals = enriched_df.groupby('country_iso_a3')['predicted_fatalities'].sum() -``` - -### Using the Default Global Mapper - -```python -from views_postprocessing.unfao.mapping.mapping import get_default_mapper - -# A global mapper is automatically initialized when the module is imported -mapper = get_default_mapper() - -# Use it directly -result = mapper.find_country_for_gid(148345) -``` - -### Cache Management - -```python -# Check cache statistics -stats = mapper.get_cache_stats() -print(f"Cache type: {stats['cache_type']}") -print(f"Country cache size: {stats['country_cache_size']}") - -# Clear cache before fresh analysis -mapper.clear_cache(cache_type="all") - -# Pre-warm cache for a specific region -african_countries = ['TZA', 'KEN', 'UGA', 'RWA', 'BDI'] -mapper.warm_cache(iso_a3_list=african_countries) -``` - ---- - -## Performance Considerations - -### Spatial Indexing - -The module uses **R-tree spatial indices** for all geographic datasets: -- Reduces point-in-polygon queries from O(n) to O(log n) -- Critical for batch processing thousands of grid cells - -### Recommended Batch Sizes - -| Operation | Recommended Batch Size | Notes | -|-----------|----------------------|-------| -| Country mapping | 1,000 | Good balance of memory and parallelization | -| Admin mapping | 500-1,000 | Admin2 is more complex | -| DataFrame enrichment | 1,000 | Depends on available memory | - -### Memory Management - -- Enable disk caching (`use_disk_cache=True`) for large datasets -- Use `only_metadata=True` in `enrich_dataframe_with_pg_info()` to minimize memory -- Clear caches periodically for long-running processes - -### Multiprocessing vs Threading - -The module uses **ThreadPoolExecutor** instead of multiprocessing to avoid: -- Pickling issues with GeoDataFrames -- Memory duplication across processes -- Complex IPC overhead - -For CPU-bound tasks, consider running multiple single-threaded instances. - ---- - -## Configuration - -### Global Constants - -```python -# Cache directory (default: ~/.priogrid_mapper_cache) -CACHE_DIR = Path.home() / ".priogrid_mapper_cache" - -# Cache sizes -COUNTRY_CACHE_MAXSIZE = 100_000 -GID_CACHE_MAXSIZE = 100_000 -GIDS_FOR_COUNTRY_CACHE_MAXSIZE = 10_000_000 -ADMIN1_CACHE_MAXSIZE = 100_000 -ADMIN2_CACHE_MAXSIZE = 100_000 -``` - -### Shapefile Paths - -Paths are configured relative to the module location: - -```python -NATURAL_EARTH_COUNTRY_PATH = Path(__file__).parent.parent.parent / "shapefiles" / "ne_110m_admin_0_countries" / "ne_110m_admin_0_countries.shp" -PRIOGRID_SHAPEFILE_PATH = Path(__file__).parent.parent.parent / "shapefiles" / "priogrid_cellshp" / "priogrid_cell.shp" -ADM_1_SHAPEFILE_PATH = Path(__file__).parent.parent.parent / "shapefiles" / "GAUL_2024_L1" / "GAUL_2024_L1.shp" -ADM_2_SHAPEFILE_PATH = Path(__file__).parent.parent.parent / "shapefiles" / "GAUL_2024_L2" / "GAUL_2024_L2.shp" -``` - ---- - -## Helper Functions - -### `cached_haversine(lat1, lon1, lat2, lon2) -> float` - -Calculate great-circle distance between two points with caching. - -```python -distance_km = cached_haversine( - lat1=-6.8, lon1=39.3, # Dar es Salaam - lat2=-1.3, lon2=36.8 # Nairobi -) -``` - -### `set_default_mapper() -> PriogridCountryMapper` - -Initialize the global default mapper. - -### `get_default_mapper() -> PriogridCountryMapper` - -Retrieve the global default mapper instance. - ---- - -## Error Handling - -The module handles various edge cases: - -| Scenario | Behavior | -|----------|----------| -| Invalid GID | Returns `None` | -| GID in water (no country) | Returns `None` or uses water handling rule | -| Invalid ISO A3 code | Returns `None` with warning | -| Missing shapefile | Raises `FileNotFoundError` during initialization | -| Invalid geometry | Logged as warning, skipped in calculations | -| Cache miss | Computes and caches result | - ---- - -## Dependencies - -``` -pandas -geopandas -shapely -numpy -joblib -cachetools -matplotlib (optional, for visualization) -tqdm (optional, for progress bars) -``` diff --git a/views_postprocessing/unfao/mapping/__init__.py b/views_postprocessing/unfao/mapping/__init__.py deleted file mode 100644 index e69de29..0000000 diff --git a/views_postprocessing/unfao/mapping/mapping.py b/views_postprocessing/unfao/mapping/mapping.py deleted file mode 100644 index 1714af1..0000000 --- a/views_postprocessing/unfao/mapping/mapping.py +++ /dev/null @@ -1,3171 +0,0 @@ -import pandas as pd -import geopandas as gpd -from shapely.geometry import Point as ShapelyPoint -from shapely import wkt -from typing import Optional - -from functools import partial -from collections import OrderedDict -import math -import numpy as np -import logging -from pathlib import Path -import os - -from multiprocessing import Pool, cpu_count -from joblib import Memory -from cachetools import LRUCache, TTLCache - -logging.basicConfig(level=logging.INFO) -logger = logging.getLogger(__name__) - -NATURAL_EARTH_COUNTRY_PATH = ( - Path(__file__).parent.parent.parent - / "shapefiles" - / "ne_10m_admin_0_countries" - / "ne_10m_admin_0_countries.shp" -) - -PRIOGRID_SHAPEFILE_PATH = ( - Path(__file__).parent.parent.parent - / "shapefiles" - / "priogrid_cellshp" - / "priogrid_cell.shp" -) -ADM_1_SHAPEFILE_PATH = ( - Path(__file__).parent.parent.parent - / "shapefiles" - / "GAUL_2024_L1" - / "GAUL_2024_L1.shp" -) -ADM_2_SHAPEFILE_PATH = ( - Path(__file__).parent.parent.parent - / "shapefiles" - / "GAUL_2024_L2" - / "GAUL_2024_L2.shp" -) - -# Configure disk cache -CACHE_DIR = Path.home() / ".priogrid_mapper_cache" -os.makedirs(CACHE_DIR, exist_ok=True) - -# Cache sizes -COUNTRY_CACHE_MAXSIZE = 100000 -GID_CACHE_MAXSIZE = 100000 -GIDS_FOR_COUNTRY_CACHE_MAXSIZE = 10000000 -ADMIN1_CACHE_MAXSIZE = 100000 -ADMIN2_CACHE_MAXSIZE = 100000 - -# Global disk caches (shared across instances) -_disk_country_cache = Memory(location=CACHE_DIR / "country", verbose=0) -_disk_admin1_cache = Memory(location=CACHE_DIR / "admin1", verbose=0) -_disk_admin2_cache = Memory(location=CACHE_DIR / "admin2", verbose=0) -_disk_gid_cache = Memory(location=CACHE_DIR / "gid", verbose=0) - - -class DistanceCache: - """Simple LRU cache for distance calculations.""" - - def __init__(self, maxsize=1000): - self.cache = OrderedDict() - self.maxsize = maxsize - - def get(self, key): - if key in self.cache: - self.cache.move_to_end(key) - return self.cache[key] - return None - - def set(self, key, value): - self.cache[key] = value - self.cache.move_to_end(key) - if len(self.cache) > self.maxsize: - self.cache.popitem(last=False) - - -# Global distance cache -_distance_cache = DistanceCache() - - -def cached_haversine(lat1, lon1, lat2, lon2): - """Calculate haversine distance with caching.""" - # Validate coordinates - if ( - not (-90 <= lat1 <= 90) - or not (-180 <= lon1 <= 180) - or not (-90 <= lat2 <= 90) - or not (-180 <= lon2 <= 180) - ): - raise ValueError( - "Invalid coordinates: latitudes must be between -90 and 90, longitudes between -180 and 180" - ) - - # Create a cache key based on the coordinates (rounded to 6 decimal places) - cache_key = (round(lat1, 6), round(lon1, 6), round(lat2, 6), round(lon2, 6)) - - # Check cache first - cached_result = _distance_cache.get(cache_key) - if cached_result is not None: - return cached_result - - # Not in cache, perform the calculation - lat1, lon1, lat2, lon2 = map(math.radians, [lat1, lon1, lat2, lon2]) - dlon = lon2 - lon1 - dlat = lat2 - lat1 - a = ( - math.sin(dlat / 2) ** 2 - + math.cos(lat1) * math.cos(lat2) * math.sin(dlon / 2) ** 2 - ) - c = 2 * math.atan2(math.sqrt(a), math.sqrt(1 - a)) - result = 6371.0 * c # Earth radius in kilometers - - # Update cache - _distance_cache.set(cache_key, result) - return result - - -class PriogridCountryMapper: - """ - A comprehensive class for mapping between PRIO-GRID cells and countries using Natural Earth shapefiles. - Now with persistent disk caching for improved performance across sessions. - - Attributes: - countries_gdf (GeoDataFrame): Processed Natural Earth country data - priogrid_gdf (GeoDataFrame): Processed PRIO-GRID data with cell geometries and centroids - priogrid_sindex (SpatialIndex): Spatial index for faster PRIO-GRID queries - """ - - def __init__(self, use_disk_cache=True, cache_dir=None, cache_ttl=None): - """ - Initialize the PriogridCountryMapper with optional disk caching. - - Parameters: - use_disk_cache (bool): Whether to use persistent disk caching - cache_dir (str or Path): Custom cache directory path - cache_ttl (int): Time-to-live for cache entries in seconds (None for no expiration) - """ - # Initialize attributes used in __del__ immediately to prevent AttributeError - # if an exception occurs during initialization (e.g., missing shapefiles). - self._process_pool = None - self._max_workers = cpu_count() - - # Load Natural Earth country data - # NOTE: This is the line that was crashing in your traceback (line 134) - country_path = str(NATURAL_EARTH_COUNTRY_PATH) - self.countries_gdf = self._load_and_preprocess_naturalearth(country_path) - - # Load PRIO-GRID data - self.priogrid_gdf = self._load_priogrid(str(PRIOGRID_SHAPEFILE_PATH)) - - self.priogrid_sindex = ( - self.priogrid_gdf.sindex if hasattr(self.priogrid_gdf, "sindex") else None - ) - - # Configure caching - self.use_disk_cache = use_disk_cache - self.cache_ttl = cache_ttl - - if use_disk_cache: - # Set up custom cache directory if provided - if cache_dir: - self.cache_dir = Path(cache_dir) - os.makedirs(self.cache_dir, exist_ok=True) - else: - self.cache_dir = CACHE_DIR - - # Initialize disk caches - self._disk_country_cache = Memory( - location=self.cache_dir / "country", verbose=0 - ) - self._disk_admin1_cache = Memory( - location=self.cache_dir / "admin1", verbose=0 - ) - self._disk_admin2_cache = Memory( - location=self.cache_dir / "admin2", verbose=0 - ) - self._disk_gid_cache = Memory(location=self.cache_dir / "gid", verbose=0) - - logger.info(f"Using disk cache at: {self.cache_dir}") - else: - # Initialize in-memory caches as instance variables - if cache_ttl: - self._country_cache = TTLCache( - maxsize=COUNTRY_CACHE_MAXSIZE, ttl=cache_ttl - ) - self._gid_cache = TTLCache(maxsize=GID_CACHE_MAXSIZE, ttl=cache_ttl) - self._gids_for_country_cache = TTLCache( - maxsize=GIDS_FOR_COUNTRY_CACHE_MAXSIZE, ttl=cache_ttl - ) - self._admin1_cache = TTLCache( - maxsize=ADMIN1_CACHE_MAXSIZE, ttl=cache_ttl - ) - self._admin2_cache = TTLCache( - maxsize=ADMIN2_CACHE_MAXSIZE, ttl=cache_ttl - ) - else: - self._country_cache = LRUCache(maxsize=COUNTRY_CACHE_MAXSIZE) - self._gid_cache = LRUCache(maxsize=GID_CACHE_MAXSIZE) - self._gids_for_country_cache = LRUCache( - maxsize=GIDS_FOR_COUNTRY_CACHE_MAXSIZE - ) - self._admin1_cache = LRUCache(maxsize=ADMIN1_CACHE_MAXSIZE) - self._admin2_cache = LRUCache(maxsize=ADMIN2_CACHE_MAXSIZE) - - logger.info("Using in-memory caching") - - # Load admin1 and admin2 data if paths are provided - self.admin1_path = str(ADM_1_SHAPEFILE_PATH) - self.admin2_path = str(ADM_2_SHAPEFILE_PATH) - - self.admin1_gdf = None - self.admin2_gdf = None - - if self.admin1_path: - self.admin1_gdf = self._load_admin_data(self.admin1_path, "admin1") - self.admin1_sindex = ( - self.admin1_gdf.sindex if hasattr(self.admin1_gdf, "sindex") else None - ) - else: - self.admin1_sindex = None - - if self.admin2_path: - self.admin2_gdf = self._load_admin_data(self.admin2_path, "admin2") - self.admin2_sindex = ( - self.admin2_gdf.sindex if hasattr(self.admin2_gdf, "sindex") else None - ) - else: - self.admin2_sindex = None - - def _get_gid_cache_key(self, gid): - """Generate a consistent cache key for GID-based lookups""" - return f"gid_{gid}" - - def _get_point_cache_key(self, point_geometry): - """Generate a consistent cache key for point-based lookups""" - return f"point_{round(point_geometry.x, 6)}_{round(point_geometry.y, 6)}" - - def _get_country_cache_key(self, iso_a3): - """Generate a consistent cache key for country-based lookups""" - return f"country_{iso_a3.upper() if iso_a3 else iso_a3}" - - def _get_admin_cache_key(self, gid, admin_level): - """Generate a consistent cache key for admin-based lookups""" - return f"{admin_level}_{gid}" - - def clear_cache(self, cache_type="all"): - """ - Clear cached data. - - Parameters: - cache_type (str): Type of cache to clear ("all", "country", "admin1", "admin2", "gid") - """ - if self.use_disk_cache: - if cache_type == "all" or cache_type == "country": - self._disk_country_cache.clear() - logger.info("Cleared country cache") - - if cache_type == "all" or cache_type == "admin1": - self._disk_admin1_cache.clear() - logger.info("Cleared admin1 cache") - - if cache_type == "all" or cache_type == "admin2": - self._disk_admin2_cache.clear() - logger.info("Cleared admin2 cache") - - if cache_type == "all" or cache_type == "gid": - self._disk_gid_cache.clear() - logger.info("Cleared GID cache") - else: - if cache_type == "all" or cache_type == "country": - self._country_cache.clear() - logger.info("Cleared country cache") - - if cache_type == "all" or cache_type == "admin1": - self._admin1_cache.clear() - logger.info("Cleared admin1 cache") - - if cache_type == "all" or cache_type == "admin2": - self._admin2_cache.clear() - logger.info("Cleared admin2 cache") - - if cache_type == "all" or cache_type == "gid": - self._gid_cache.clear() - logger.info("Cleared GID cache") - - def get_cache_stats(self): - """ - Get statistics about cache usage. - - Returns: - dict: Cache statistics - """ - if self.use_disk_cache: - stats = { - "cache_type": "disk", - "cache_dir": str(self.cache_dir), - "country_cache_size": len(self._disk_country_cache), - "admin1_cache_size": len(self._disk_admin1_cache), - "admin2_cache_size": len(self._disk_admin2_cache), - "gid_cache_size": len(self._disk_gid_cache), - } - - # Calculate total cache size on disk - total_size = 0 - for cache_dir in [ - self.cache_dir / d for d in ["country", "admin1", "admin2", "gid"] - ]: - if cache_dir.exists(): - for file in cache_dir.glob("*"): - if file.is_file(): - total_size += file.stat().st_size - - stats["total_disk_size_mb"] = round(total_size / (1024 * 1024), 2) - else: - stats = { - "cache_type": "memory", - "country_cache_size": len(self._country_cache), - "country_cache_maxsize": self._country_cache.maxsize, - "gid_cache_size": len(self._gid_cache), - "gid_cache_maxsize": self._gid_cache.maxsize, - "admin1_cache_size": len(self._admin1_cache), - "admin1_cache_maxsize": self._admin1_cache.maxsize, - "admin2_cache_size": len(self._admin2_cache), - "admin2_cache_maxsize": self._admin2_cache.maxsize, - } - - if hasattr(self._country_cache, "currsize"): - stats["country_cache_hit_rate"] = ( - f"{self._country_cache.hits}/{self._country_cache.hits + self._country_cache.misses}" - ) - stats["gid_cache_hit_rate"] = ( - f"{self._gid_cache.hits}/{self._gid_cache.hits + self._gid_cache.misses}" - ) - - return stats - - def warm_cache(self, gid_list=None, iso_a3_list=None): - """ - Warm up the cache with frequently accessed data. - - Parameters: - gid_list (list): List of GIDs to pre-cache - iso_a3_list (list): List of ISO A3 codes to pre-cache - """ - if gid_list: - logger.info(f"Warming cache with {len(gid_list)} GIDs") - for gid in gid_list: - self.find_country_for_gid(gid) - self.find_admin1_for_gid(gid) - self.find_admin2_for_gid(gid) - - if iso_a3_list: - logger.info(f"Warming cache with {len(iso_a3_list)} countries") - for iso_a3 in iso_a3_list: - self.find_country_by_iso_a3(iso_a3) - self.find_gids_for_country(iso_a3) - - logger.info("Cache warming complete") - - def __del__(self): - if self._process_pool: - self._process_pool.close() - self._process_pool.join() - - def _init_process_pool(self): - """Initialize the process pool if not already done""" - if self._process_pool is None: - self._process_pool = Pool(processes=self._max_workers) - - def batch_country_mapping_parallel(self, gid_list, batch_size=1000): - """ - Find countries for multiple PRIO-GRID cells using multiprocessing. - - Parameters: - gid_list (list): List of PRIO-GRID cell IDs - batch_size (int): Number of GIDs to process in each batch - - Returns: - DataFrame: Results of the batch mapping - """ - self._init_process_pool() - - # Split into batches - batches = [ - gid_list[i : i + batch_size] for i in range(0, len(gid_list), batch_size) - ] - - # Create partial function for mapping - map_func = partial(self._process_gid_batch) - - # Process batches in parallel - results = [] - for batch_result in self._process_pool.imap_unordered(map_func, batches): - results.extend(batch_result) - - return pd.DataFrame(results) - - def _process_gid_batch(self, gid_batch): - """Process a batch of GIDs in a single process""" - batch_results = [] - for gid in gid_batch: - result = self.find_country_for_gid(gid) - batch_results.append(result) - return batch_results - - def _load_and_preprocess_naturalearth(self, country_path): - """ - Load and preprocess the Natural Earth country dataset from the provided path. - - Parameters: - country_path (str): Path to the Natural Earth country dataset file - - Returns: - GeoDataFrame: Processed Natural Earth data with proper geometry - - Raises: - FileNotFoundError: If the file cannot be found - ValueError: If the file format is unsupported or data cannot be parsed - """ - # Load Natural Earth data - if country_path.endswith(".shp"): - countries_gdf = gpd.read_file(country_path) - else: - # Assume it's a CSV with WKT geometry - countries_df = pd.read_csv(country_path) - - # Parse the WKT geometry, removing the SRID prefix if present - def parse_wkt(wkt_str): - if wkt_str.startswith("SRID=4326;"): - wkt_str = wkt_str.replace("SRID=4326;", "") - return wkt.loads(wkt_str) - - countries_df["geometry"] = countries_df["geometry"].apply(parse_wkt) - countries_gdf = gpd.GeoDataFrame(countries_df, geometry="geometry") - - # Set CRS if not already set - if countries_gdf.crs is None: - countries_gdf.set_crs(epsg=4326, inplace=True) - - # Fix invalid geometries before validation and use - countries_gdf["geometry"] = countries_gdf["geometry"].make_valid() - - # Validate data - self._validate_naturalearth_data(countries_gdf) - - return countries_gdf - - def _validate_naturalearth_data(self, countries_gdf): - """Validate Natural Earth data for consistency and correctness.""" - # Check for required columns - required_columns = ["ISO_A3", "NAME_EN", "geometry"] - missing_columns = [ - col for col in required_columns if col not in countries_gdf.columns - ] - if missing_columns: - err_msg = f"Missing required columns in Natural Earth data: {missing_columns}" - logger.error(err_msg) - raise ValueError(err_msg) - - # Check geometry validity - invalid_geometries = countries_gdf[~countries_gdf["geometry"].is_valid] - if not invalid_geometries.empty: - logger.warning( - f"Found {len(invalid_geometries)} invalid geometries in Natural Earth data" - ) - - def _load_priogrid(self, priogrid_path): - """ - Load and preprocess the PRIO-GRID data from the provided path. - - Parameters: - priogrid_path (str): Path to the PRIO-GRID dataset file - - Returns: - GeoDataFrame: Processed PRIO-GRID data with geometry and centroid fields - - Raises: - FileNotFoundError: If the file cannot be found - ValueError: If the file format is unsupported or data cannot be parsed - """ - # Load PRIO-GRID data - priogrid_gdf = gpd.read_file(priogrid_path) - - # Ensure CRS is set - if priogrid_gdf.crs is None: - priogrid_gdf.set_crs(epsg=4326, inplace=True) - - # Fix invalid geometries before use - priogrid_gdf["geometry"] = priogrid_gdf["geometry"].make_valid() - - import warnings - with warnings.catch_warnings(): - warnings.filterwarnings("ignore", message=".*Geometry is in a geographic CRS.*") - priogrid_gdf["centroid"] = priogrid_gdf["geometry"].centroid - - # Validate data - self._validate_priogrid_data(priogrid_gdf) - - return priogrid_gdf - - def _validate_priogrid_data(self, priogrid_gdf): - """Validate PRIO-GRID data for consistency and correctness.""" - # Check for required columns - required_columns = ["gid", "geometry", "centroid"] - missing_columns = [ - col for col in required_columns if col not in priogrid_gdf.columns - ] - if missing_columns: - err_msg = f"Missing required columns in PRIO-GRID data: {missing_columns}" - logger.error(err_msg) - raise ValueError(err_msg) - - # Check geometry validity - invalid_geometries = priogrid_gdf[~priogrid_gdf["geometry"].is_valid] - if not invalid_geometries.empty: - logger.warning( - f"Found {len(invalid_geometries)} invalid geometries in PRIO-GRID data" - ) - - def find_gid_for_point(self, point_geometry): - """ - Find the PRIO-GRID cell ID that contains a point using spatial index for faster query. - - Parameters: - point_geometry (shapely.geometry.Point): Point geometry to locate - - Returns: - int or None: PRIO-GRID cell ID containing the point, or None if not found - """ - if self.use_disk_cache: - # Use disk cache - @self._disk_gid_cache.cache() - def _find_gid_for_point_impl(point_geometry): - # Validate input geometry - if not point_geometry.is_valid: - raise ValueError("Invalid point geometry") - - # Not in cache, perform the expensive operation - if self.priogrid_sindex: - possible_matches_index = list( - self.priogrid_sindex.intersection(point_geometry.bounds) - ) - possible_matches = self.priogrid_gdf.iloc[possible_matches_index] - precise_matches = possible_matches[ - possible_matches.intersects(point_geometry) - ] - - if len(precise_matches) > 0: - result = precise_matches.iloc[0]["gid"] - else: - result = None - else: - # Fallback to iterative search - result = None - for _, grid_cell in self.priogrid_gdf.iterrows(): - if grid_cell["geometry"].contains(point_geometry): - result = grid_cell["gid"] - break - - return result - - return _find_gid_for_point_impl(point_geometry) - else: - # Use in-memory cache - # Validate input geometry - if not point_geometry.is_valid: - raise ValueError("Invalid point geometry") - - # Create a consistent cache key - cache_key = self._get_point_cache_key(point_geometry) - - # Check cache first - if cache_key in self._gid_cache: - return self._gid_cache[cache_key] - - # Not in cache, perform the expensive operation - if self.priogrid_sindex: - possible_matches_index = list( - self.priogrid_sindex.intersection(point_geometry.bounds) - ) - possible_matches = self.priogrid_gdf.iloc[possible_matches_index] - precise_matches = possible_matches[ - possible_matches.intersects(point_geometry) - ] - - if len(precise_matches) > 0: - result = precise_matches.iloc[0]["gid"] - else: - result = None - else: - # Fallback to iterative search - result = None - for _, grid_cell in self.priogrid_gdf.iterrows(): - if grid_cell["geometry"].contains(point_geometry): - result = grid_cell["gid"] - break - - # Update cache - self._gid_cache[cache_key] = result - return result - - def find_country_for_gid(self, gid): - """ - Find the country information for a PRIO-GRID cell using majority area-based method. - Now with persistent disk caching. - """ - if self.use_disk_cache: - # Use disk cache - @self._disk_country_cache.cache() - def _find_country_for_gid_impl(gid): - # Get the PRIO-GRID cell - grid_cell = self.priogrid_gdf[self.priogrid_gdf["gid"] == gid] - if len(grid_cell) == 0: - return None - - grid_geometry = grid_cell["geometry"].iloc[0] - - if grid_geometry.area == 0.0: - logger.warning(f"Zero-area grid geometry for GID {gid}, skipping country overlap") - return None - - # Use spatial index to find potentially intersecting countries - if hasattr(self.countries_gdf, "sindex"): - possible_countries_idx = list( - self.countries_gdf.sindex.intersection(grid_geometry.bounds) - ) - candidate_countries = self.countries_gdf.iloc[ - possible_countries_idx - ] - candidate_countries = candidate_countries[ - candidate_countries.intersects(grid_geometry) - ] - else: - candidate_countries = self.countries_gdf[ - self.countries_gdf.intersects(grid_geometry) - ] - - # Calculate area overlaps - overlaps = [] - for _, country in candidate_countries.iterrows(): - try: - intersection = country["geometry"].intersection(grid_geometry) - overlap_area = intersection.area - total_area = grid_geometry.area - overlap_ratio = overlap_area / total_area - - overlaps.append( - { - "country_data": country, - "overlap_area": overlap_area, - "overlap_ratio": overlap_ratio, - } - ) - except Exception as e: - logger.debug(f"Overlap calculation error for GID {gid}: {e}") - continue - - # Sort by overlap ratio (descending) - overlaps.sort(key=lambda x: x["overlap_ratio"], reverse=True) - - # Assignment rule: always use largest overlap - result = None - if overlaps: - country = overlaps[0]["country_data"] - method_used = "largest overlap" - - # Prepare result - if overlaps: - result = { - "gid": int(gid), - "iso_a3": country["ISO_A3"], - "country_name": country["NAME_EN"], - "overlap_ratio": ( - float(overlaps[0]["overlap_ratio"]) if overlaps else 0.0 - ), - "method": method_used, - } - - # Add all available country data - for col in country.index: - if col not in ["geometry", "ISO_A3", "NAME_EN"]: - result[col] = country[col] - - return result - - return _find_country_for_gid_impl(gid) - else: - # Use in-memory cache - # Generate consistent cache key - cache_key = self._get_gid_cache_key(gid) - - # Check cache first - if cache_key in self._country_cache: - cached_value = self._country_cache[cache_key] - if cached_value is None: - return None - result = cached_value.copy() - return result - - # Get the PRIO-GRID cell - grid_cell = self.priogrid_gdf[self.priogrid_gdf["gid"] == gid] - if len(grid_cell) == 0: - self._country_cache[cache_key] = None - return None - - grid_geometry = grid_cell["geometry"].iloc[0] - - if grid_geometry.area == 0.0: - logger.warning(f"Zero-area grid geometry for GID {gid}, skipping country overlap") - self._country_cache[cache_key] = None - return None - - # Use spatial index to find potentially intersecting countries - if hasattr(self.countries_gdf, "sindex"): - possible_countries_idx = list( - self.countries_gdf.sindex.intersection(grid_geometry.bounds) - ) - candidate_countries = self.countries_gdf.iloc[possible_countries_idx] - candidate_countries = candidate_countries[ - candidate_countries.intersects(grid_geometry) - ] - else: - candidate_countries = self.countries_gdf[ - self.countries_gdf.intersects(grid_geometry) - ] - - # Calculate area overlaps - overlaps = [] - for _, country in candidate_countries.iterrows(): - try: - intersection = country["geometry"].intersection(grid_geometry) - overlap_area = intersection.area - total_area = grid_geometry.area - overlap_ratio = overlap_area / total_area - - overlaps.append( - { - "country_data": country, - "overlap_area": overlap_area, - "overlap_ratio": overlap_ratio, - } - ) - except Exception as e: - logger.debug(f"Overlap calculation error for GID {gid}: {e}") - continue - - # Sort by overlap ratio (descending) - overlaps.sort(key=lambda x: x["overlap_ratio"], reverse=True) - - # Assignment rule: always use largest overlap - result = None - if overlaps: - country = overlaps[0]["country_data"] - method_used = "largest overlap" - - # Prepare result - if overlaps: - result = { - "gid": int(gid), - "iso_a3": country["ISO_A3"], - "country_name": country["NAME_EN"], - "overlap_ratio": ( - float(overlaps[0]["overlap_ratio"]) if overlaps else 0.0 - ), - "method": method_used, - } - - # Add all available country data - for col in country.index: - if col not in ["geometry", "ISO_A3", "NAME_EN"]: - result[col] = country[col] - - # Update cache - self._country_cache[cache_key] = result - return result - - def batch_country_mapping(self, gid_list): - """ - Optimized batch mapping using vectorized operations where possible. - """ - results = [] - cache_hits = 0 - - for gid in gid_list: - # Check cache first - cache_key = self._get_gid_cache_key(gid) - if cache_key in self._country_cache: - cached_value = self._country_cache[cache_key] - if cached_value is not None: - results.append(cached_value.copy()) - else: - results.append(None) - cache_hits += 1 - continue - - # Not in cache, process this GID - result = self.find_country_for_gid(gid) - results.append(result) - - if cache_hits > 0: - logger.info( - f"Cache hits: {cache_hits}/{len(gid_list)} ({cache_hits/len(gid_list)*100:.1f}%)" - ) - - return pd.DataFrame([r for r in results if r is not None]) - - def find_gids_for_country(self, iso_a3): - """ - Find all PRIO-GRID cells that belong to a specific country. - Optimized using spatial indexing and vectorized operations. - """ - if self.use_disk_cache: - # Use disk cache - @self._disk_country_cache.cache() - def _find_gids_for_country_impl(iso_a3): - # Filter Natural Earth data for the specific country - country = self.countries_gdf[self.countries_gdf["ISO_A3"] == iso_a3] - - if len(country) == 0: - return [] - else: - # Get the country geometry - country_geometry = country.iloc[0]["geometry"] - - # OPTIMIZATION: Use spatial index to quickly find intersecting grid cells - if self.priogrid_sindex: - possible_matches_index = list( - self.priogrid_sindex.intersection(country_geometry.bounds) - ) - candidate_grid_cells = self.priogrid_gdf.iloc[ - possible_matches_index - ] - intersecting_cells = candidate_grid_cells[ - candidate_grid_cells.intersects(country_geometry) - ] - else: - intersecting_cells = self.priogrid_gdf[ - self.priogrid_gdf.intersects(country_geometry) - ] - - if len(intersecting_cells) == 0: - result = [] - else: - # OPTIMIZATION: Process with early termination - matching_gids = self._find_dominant_country_gids( - intersecting_cells, country_geometry, iso_a3 - ) - result = matching_gids - - return result - - return _find_gids_for_country_impl(iso_a3) - else: - # Use in-memory cache - # Create consistent cache key - cache_key = self._get_country_cache_key(iso_a3) - - # Check cache first - if cache_key in self._gids_for_country_cache: - result = self._gids_for_country_cache[cache_key] - return result - - # Not in cache, perform optimized operation - # Filter Natural Earth data for the specific country - country = self.countries_gdf[self.countries_gdf["ISO_A3"] == iso_a3] - - if len(country) == 0: - result = [] - else: - # Get the country geometry - country_geometry = country.iloc[0]["geometry"] - - # OPTIMIZATION: Use spatial index to quickly find intersecting grid cells - if self.priogrid_sindex: - possible_matches_index = list( - self.priogrid_sindex.intersection(country_geometry.bounds) - ) - candidate_grid_cells = self.priogrid_gdf.iloc[ - possible_matches_index - ] - intersecting_cells = candidate_grid_cells[ - candidate_grid_cells.intersects(country_geometry) - ] - else: - intersecting_cells = self.priogrid_gdf[ - self.priogrid_gdf.intersects(country_geometry) - ] - - if len(intersecting_cells) == 0: - result = [] - else: - # OPTIMIZATION: Process with early termination - matching_gids = self._find_dominant_country_gids( - intersecting_cells, country_geometry, iso_a3 - ) - result = matching_gids - - # Update cache - self._gids_for_country_cache[cache_key] = result - return result - - def _find_dominant_country_gids(self, intersecting_cells, country_geometry, iso_a3): - """Find grid cells where target country has dominant area using optimized approach.""" - matching_gids = [] - - for _, grid_cell in intersecting_cells.iterrows(): - grid_geometry = grid_cell["geometry"] - grid_gid = grid_cell["gid"] - - # OPTIMIZATION 1: Early check for complete containment - if country_geometry.contains(grid_geometry): - matching_gids.append(int(grid_gid)) - continue - - # Calculate overlap with the target country - if grid_geometry.area == 0.0: - logger.warning(f"Zero-area grid geometry for GID {grid_gid}, skipping country reverse lookup") - continue - - try: - intersection = country_geometry.intersection(grid_geometry) - overlap_ratio = intersection.area / grid_geometry.area - - # If overlap is significant, include it - if overlap_ratio > 0.5: - matching_gids.append(int(grid_gid)) - except Exception as e: - logger.debug(f"Geometry error for GID {grid_gid}: {e}") - continue - - return matching_gids - - def get_all_iso_a3_codes(self): - """ - Get a list of all unique ISO A3 codes in the Natural Earth dataset. - - Returns: - list: List of unique ISO A3 codes - """ - return self.countries_gdf["ISO_A3"].unique().tolist() - - def visualize_grid_and_country(self, gid): - """ - Visualize a PRIO-GRID cell and the country it belongs to. - - Parameters: - gid (int): PRIO-GRID cell ID - - Raises: - ImportError: If matplotlib is not available - ValueError: If the PRIO-GRID cell ID is not found - """ - try: - import matplotlib.pyplot as plt - except ImportError: - logger.error("Matplotlib is required for visualization") - return - - # Get the grid cell - grid_cell = self.priogrid_gdf[self.priogrid_gdf["gid"] == gid] - - if len(grid_cell) == 0: - raise ValueError(f"PRIO-GRID cell with ID {gid} not found") - - # Find which country contains the grid cell based on majority area - country_info = self.find_country_for_gid(gid) - containing_country = None - - if country_info: - iso_a3 = country_info["iso_a3"] - containing_country = self.countries_gdf[ - self.countries_gdf["ISO_A3"] == iso_a3 - ].iloc[0] - - # Create plot - fig, ax = plt.subplots(1, 1, figsize=(10, 8)) - - # Plot all countries - self.countries_gdf.plot(ax=ax, color="lightgray", edgecolor="white", alpha=0.5) - - # Highlight the containing country - if containing_country is not None: - gpd.GeoDataFrame([containing_country], geometry="geometry").plot( - ax=ax, color="lightblue", edgecolor="blue", alpha=0.7 - ) - - # Plot the grid cell - grid_cell.plot(ax=ax, color="red", alpha=0.7) - - # Add centroid marker - centroid = grid_cell["centroid"].iloc[0] - ax.scatter( - centroid.x, - centroid.y, - color="yellow", - s=100, - marker="*", - edgecolor="black", - label="Grid Centroid", - ) - - # Set title and legend - country_name = ( - containing_country["NAME_EN"] - if containing_country is not None - else "No Country" - ) - method = country_info.get("method", "unknown") if country_info else "unknown" - overlap_ratio = country_info.get("overlap_ratio", 0) if country_info else 0 - - ax.set_title( - f"PRIO-GRID Cell {gid} in {country_name}\n" - f"Method: {method} | Overlap Ratio: {overlap_ratio:.3f}" - ) - ax.legend() - - plt.tight_layout() - plt.show() - - def calculate_capital_distance(self, grid_lat, grid_lon, capital_coords): - """ - Calculate distance from grid cell to capital. - - Parameters: - grid_lat: Latitude of grid centroid - grid_lon: Longitude of grid centroid - capital_coords: Tuple of (capital_lon, capital_lat) or None - - Returns: - float or None: Distance in kilometers or None if calculation fails - """ - if capital_coords is None or None in capital_coords: - return None - - try: - capital_lon, capital_lat = capital_coords - return cached_haversine(grid_lat, grid_lon, capital_lat, capital_lon) - except Exception as e: - logger.error(f"Error calculating capital distance: {e}") - return None - - def get_gid_from_point(self, point: ShapelyPoint) -> Optional[int]: - """ - Find the PRIO-GRID cell ID that contains a point. - - Parameters: - point (ShapelyPoint): Point geometry to locate - - Returns: - int or None: PRIO-GRID cell ID containing the point, or None if not found - """ - for _, grid_cell in self.priogrid_gdf.iterrows(): - if grid_cell["geometry"].contains(point): - return grid_cell["gid"] - return None - - def get_point_from_gid(self, gid: int) -> Optional["Point"]: # noqa: F821 - """ - Get the centroid point of a PRIO-GRID cell. - - Parameters: - gid (int): PRIO-GRID cell ID - - Returns: - Point or None: Point object representing the centroid, or None if not found - """ - # from ..space import Point - grid_cell = self.priogrid_gdf[self.priogrid_gdf["gid"] == gid] - if len(grid_cell) == 0: - return None - centroid = grid_cell["centroid"].iloc[0] - return Point(lat=centroid.y, lon=centroid.x) # noqa: F821 - - def _load_admin_data(self, admin_path, admin_level): - """ - Load and preprocess admin1 or admin2 data from the provided path. - - Parameters: - admin_path (str): Path to the admin dataset file - admin_level (str): Either "admin1" or "admin2" - - Returns: - GeoDataFrame: Processed admin data with proper geometry - """ - # Load admin data - if admin_path.endswith(".shp"): - admin_gdf = gpd.read_file(admin_path) - else: - # Assume it's a CSV with WKT geometry - admin_df = pd.read_csv(admin_path) - - # Parse the WKT geometry, removing the SRID prefix if present - def parse_wkt(wkt_str): - if wkt_str.startswith("SRID=4326;"): - wkt_str = wkt_str.replace("SRID=4326;", "") - return wkt.loads(wkt_str) - - admin_df["geometry"] = admin_df["geometry"].apply(parse_wkt) - admin_gdf = gpd.GeoDataFrame(admin_df, geometry="geometry") - - # Set CRS if not already set - if admin_gdf.crs is None: - admin_gdf.set_crs(epsg=4326, inplace=True) - - # Fix invalid geometries before use - admin_gdf["geometry"] = admin_gdf["geometry"].make_valid() - - # Validate data - self._validate_admin_data(admin_gdf, admin_level) - - return admin_gdf - - def _validate_admin_data(self, admin_gdf, admin_level): - """Validate admin data for consistency and correctness.""" - # Check for required columns based on admin level - if admin_level == "admin1": - required_columns = ["gaul1_code", "gaul1_name", "iso3_code", "geometry"] - else: # admin2 - required_columns = ["gaul2_code", "gaul2_name", "iso3_code", "geometry"] - - missing_columns = [ - col for col in required_columns if col not in admin_gdf.columns - ] - if missing_columns: - err_msg = f"Missing required columns in {admin_level} data: {missing_columns}" - logger.error(err_msg) - raise ValueError(err_msg) - - # Check geometry validity - invalid_geometries = admin_gdf[~admin_gdf["geometry"].is_valid] - if not invalid_geometries.empty: - logger.warning( - f"Found {len(invalid_geometries)} invalid geometries in {admin_level} data" - ) - - def find_admin1_for_gid(self, gid): - """ - Find the admin1 information for a PRIO-GRID cell. - First determines the country, then finds the admin1 region within that country. - Now with persistent disk caching. - """ - if self.admin1_gdf is None: - logger.warning("Admin1 data not loaded. Cannot find admin1 for GID.") - return None - - if self.use_disk_cache: - # Use disk cache - @self._disk_admin1_cache.cache() - def _find_admin1_for_gid_impl(gid): - # First find the country for this GID - country_info = self.find_country_for_gid(gid) - if not country_info: - return None - - iso_a3 = country_info["iso_a3"] - - # Get the PRIO-GRID cell - grid_cell = self.priogrid_gdf[self.priogrid_gdf["gid"] == gid] - if len(grid_cell) == 0: - return None - - grid_geometry = grid_cell["geometry"].iloc[0] - - if grid_geometry.area == 0.0: - logger.warning(f"Zero-area grid geometry for GID {gid}, skipping admin1 overlap") - return None - - # Filter admin1 regions to only those in the same country - country_admin1 = self.admin1_gdf[self.admin1_gdf["iso3_code"] == iso_a3] - - if len(country_admin1) == 0: - return None - - # Use spatial index to find potentially intersecting admin1 regions - if self.admin1_sindex: - possible_admin1_idx = list( - self.admin1_sindex.intersection(grid_geometry.bounds) - ) - candidate_admin1 = self.admin1_gdf.iloc[possible_admin1_idx] - candidate_admin1 = candidate_admin1[ - candidate_admin1.intersects(grid_geometry) - ] - # Filter to only those in the same country - candidate_admin1 = candidate_admin1[ - candidate_admin1["iso3_code"] == iso_a3 - ] - else: - candidate_admin1 = country_admin1[ - country_admin1.intersects(grid_geometry) - ] - - if len(candidate_admin1) == 0: - return None - - # Calculate overlaps and use largest overlap - overlaps = [] - for _, admin1 in candidate_admin1.iterrows(): - try: - intersection = admin1["geometry"].intersection(grid_geometry) - overlap_area = intersection.area - total_area = grid_geometry.area - overlap_ratio = overlap_area / total_area - - overlaps.append( - { - "admin1_data": admin1, - "overlap_area": overlap_area, - "overlap_ratio": overlap_ratio, - } - ) - except Exception as e: - logger.debug(f"Overlap calculation error for GID {gid}: {e}") - continue - - # Sort by overlap ratio (descending) - overlaps.sort(key=lambda x: x["overlap_ratio"], reverse=True) - - if not overlaps: - return None - - admin1 = overlaps[0]["admin1_data"] - method_used = "largest overlap" - - # Prepare result - result = { - "gid": int(gid), - "gaul1_code": ( - int(admin1["gaul1_code"]) - if admin1["gaul1_code"] is not None - else None - ), - "gaul1_name": admin1["gaul1_name"], - "iso3_code": admin1["iso3_code"], - "method": method_used, - } - - # Add additional fields if available - if "gaul0_code" in admin1: - result["gaul0_code"] = ( - int(admin1["gaul0_code"]) - if admin1["gaul0_code"] is not None - else None - ) - if "gaul0_name" in admin1: - result["gaul0_name"] = admin1["gaul0_name"] - if "continent" in admin1: - result["continent"] = admin1["continent"] - if "disp_en" in admin1: - result["disp_en"] = admin1["disp_en"] - - return result - - return _find_admin1_for_gid_impl(gid) - else: - # Use in-memory cache - # Generate consistent cache key - cache_key = self._get_admin_cache_key(gid, "admin1") - - # Check cache first - if cache_key in self._admin1_cache: - cached_value = self._admin1_cache[cache_key] - if cached_value is not None: - result = cached_value.copy() - return result - else: - return None - - # First find the country for this GID - country_info = self.find_country_for_gid(gid) - if not country_info: - self._admin1_cache[cache_key] = None - return None - - iso_a3 = country_info["iso_a3"] - - # Get the PRIO-GRID cell - grid_cell = self.priogrid_gdf[self.priogrid_gdf["gid"] == gid] - if len(grid_cell) == 0: - self._admin1_cache[cache_key] = None - return None - - grid_geometry = grid_cell["geometry"].iloc[0] - - if grid_geometry.area == 0.0: - logger.warning(f"Zero-area grid geometry for GID {gid}, skipping admin1 overlap") - self._admin1_cache[cache_key] = None - return None - - # Filter admin1 regions to only those in the same country - country_admin1 = self.admin1_gdf[self.admin1_gdf["iso3_code"] == iso_a3] - - if len(country_admin1) == 0: - self._admin1_cache[cache_key] = None - return None - - # Use spatial index to find potentially intersecting admin1 regions - if self.admin1_sindex: - possible_admin1_idx = list( - self.admin1_sindex.intersection(grid_geometry.bounds) - ) - candidate_admin1 = self.admin1_gdf.iloc[possible_admin1_idx] - candidate_admin1 = candidate_admin1[ - candidate_admin1.intersects(grid_geometry) - ] - # Filter to only those in the same country - candidate_admin1 = candidate_admin1[ - candidate_admin1["iso3_code"] == iso_a3 - ] - else: - candidate_admin1 = country_admin1[ - country_admin1.intersects(grid_geometry) - ] - - if len(candidate_admin1) == 0: - self._admin1_cache[cache_key] = None - return None - - # Calculate overlaps and use largest overlap - overlaps = [] - for _, admin1 in candidate_admin1.iterrows(): - try: - intersection = admin1["geometry"].intersection(grid_geometry) - overlap_area = intersection.area - total_area = grid_geometry.area - overlap_ratio = overlap_area / total_area - - overlaps.append( - { - "admin1_data": admin1, - "overlap_area": overlap_area, - "overlap_ratio": overlap_ratio, - } - ) - except Exception as e: - logger.debug(f"Overlap calculation error for GID {gid}: {e}") - continue - - # Sort by overlap ratio (descending) - overlaps.sort(key=lambda x: x["overlap_ratio"], reverse=True) - - if not overlaps: - self._admin1_cache[cache_key] = None - return None - - admin1 = overlaps[0]["admin1_data"] - method_used = "largest overlap" - - # Prepare result - result = { - "gid": int(gid), - "gaul1_code": ( - int(admin1["gaul1_code"]) - if admin1["gaul1_code"] is not None - else None - ), - "gaul1_name": admin1["gaul1_name"], - "iso3_code": admin1["iso3_code"], - "method": method_used, - } - - # Add additional fields if available - if "gaul0_code" in admin1: - result["gaul0_code"] = ( - int(admin1["gaul0_code"]) - if admin1["gaul0_code"] is not None - else None - ) - if "gaul0_name" in admin1: - result["gaul0_name"] = admin1["gaul0_name"] - if "continent" in admin1: - result["continent"] = admin1["continent"] - if "disp_en" in admin1: - result["disp_en"] = admin1["disp_en"] - - # Update cache - self._admin1_cache[cache_key] = result - return result - - def find_admin2_for_gid(self, gid): - """ - Find the admin2 information for a PRIO-GRID cell. - First determines the country, then finds the admin2 region within that country. - Now with persistent disk caching. - """ - if self.admin2_gdf is None: - logger.warning("Admin2 data not loaded. Cannot find admin2 for GID.") - return None - - if self.use_disk_cache: - # Use disk cache - @self._disk_admin2_cache.cache() - def _find_admin2_for_gid_impl(gid): - # First find the country for this GID - country_info = self.find_country_for_gid(gid) - if not country_info: - return None - - iso_a3 = country_info["iso_a3"] - - # Get the PRIO-GRID cell - grid_cell = self.priogrid_gdf[self.priogrid_gdf["gid"] == gid] - if len(grid_cell) == 0: - return None - - grid_geometry = grid_cell["geometry"].iloc[0] - - if grid_geometry.area == 0.0: - logger.warning(f"Zero-area grid geometry for GID {gid}, skipping admin2 overlap") - return None - - # Filter admin2 regions to only those in the same country - country_admin2 = self.admin2_gdf[self.admin2_gdf["iso3_code"] == iso_a3] - - if len(country_admin2) == 0: - return None - - # Use spatial index to find potentially intersecting admin2 regions - if self.admin2_sindex: - possible_admin2_idx = list( - self.admin2_sindex.intersection(grid_geometry.bounds) - ) - candidate_admin2 = self.admin2_gdf.iloc[possible_admin2_idx] - candidate_admin2 = candidate_admin2[ - candidate_admin2.intersects(grid_geometry) - ] - # Filter to only those in the same country - candidate_admin2 = candidate_admin2[ - candidate_admin2["iso3_code"] == iso_a3 - ] - else: - candidate_admin2 = country_admin2[ - country_admin2.intersects(grid_geometry) - ] - - if len(candidate_admin2) == 0: - return None - - # Calculate overlaps and use largest overlap - overlaps = [] - for _, admin2 in candidate_admin2.iterrows(): - try: - intersection = admin2["geometry"].intersection(grid_geometry) - overlap_area = intersection.area - total_area = grid_geometry.area - overlap_ratio = overlap_area / total_area - - overlaps.append( - { - "admin2_data": admin2, - "overlap_area": overlap_area, - "overlap_ratio": overlap_ratio, - } - ) - except Exception as e: - logger.debug(f"Overlap calculation error for GID {gid}: {e}") - continue - - # Sort by overlap ratio (descending) - overlaps.sort(key=lambda x: x["overlap_ratio"], reverse=True) - - if not overlaps: - return None - - admin2 = overlaps[0]["admin2_data"] - method_used = "largest overlap" - - # Prepare result - result = { - "gid": int(gid), - "gaul2_code": ( - int(admin2["gaul2_code"]) - if admin2["gaul2_code"] is not None - else None - ), - "gaul2_name": admin2["gaul2_name"], - "iso3_code": admin2["iso3_code"], - "method": method_used, - } - - # Add additional fields if available - if "gaul0_code" in admin2: - result["gaul0_code"] = ( - int(admin2["gaul0_code"]) - if admin2["gaul0_code"] is not None - else None - ) - if "gaul0_name" in admin2: - result["gaul0_name"] = admin2["gaul0_name"] - if "gaul1_code" in admin2: - result["gaul1_code"] = ( - int(admin2["gaul1_code"]) - if admin2["gaul1_code"] is not None - else None - ) - if "gaul1_name" in admin2: - result["gaul1_name"] = admin2["gaul1_name"] - if "continent" in admin2: - result["continent"] = admin2["continent"] - if "disp_en" in admin2: - result["disp_en"] = admin2["disp_en"] - - return result - - return _find_admin2_for_gid_impl(gid) - else: - # Use in-memory cache - # Generate consistent cache key - cache_key = self._get_admin_cache_key(gid, "admin2") - - # Check cache first - if cache_key in self._admin2_cache: - cached_value = self._admin2_cache[cache_key] - if cached_value is not None: - result = cached_value.copy() - return result - else: - return None - - # First find the country for this GID - country_info = self.find_country_for_gid(gid) - if not country_info: - self._admin2_cache[cache_key] = None - return None - - iso_a3 = country_info["iso_a3"] - - # Get the PRIO-GRID cell - grid_cell = self.priogrid_gdf[self.priogrid_gdf["gid"] == gid] - if len(grid_cell) == 0: - self._admin2_cache[cache_key] = None - return None - - grid_geometry = grid_cell["geometry"].iloc[0] - - if grid_geometry.area == 0.0: - logger.warning(f"Zero-area grid geometry for GID {gid}, skipping admin2 overlap") - self._admin2_cache[cache_key] = None - return None - - # Filter admin2 regions to only those in the same country - country_admin2 = self.admin2_gdf[self.admin2_gdf["iso3_code"] == iso_a3] - - if len(country_admin2) == 0: - self._admin2_cache[cache_key] = None - return None - - # Use spatial index to find potentially intersecting admin2 regions - if self.admin2_sindex: - possible_admin2_idx = list( - self.admin2_sindex.intersection(grid_geometry.bounds) - ) - candidate_admin2 = self.admin2_gdf.iloc[possible_admin2_idx] - candidate_admin2 = candidate_admin2[ - candidate_admin2.intersects(grid_geometry) - ] - # Filter to only those in the same country - candidate_admin2 = candidate_admin2[ - candidate_admin2["iso3_code"] == iso_a3 - ] - else: - candidate_admin2 = country_admin2[ - country_admin2.intersects(grid_geometry) - ] - - if len(candidate_admin2) == 0: - self._admin2_cache[cache_key] = None - return None - - # Calculate overlaps and use largest overlap - overlaps = [] - for _, admin2 in candidate_admin2.iterrows(): - try: - intersection = admin2["geometry"].intersection(grid_geometry) - overlap_area = intersection.area - total_area = grid_geometry.area - overlap_ratio = overlap_area / total_area - - overlaps.append( - { - "admin2_data": admin2, - "overlap_area": overlap_area, - "overlap_ratio": overlap_ratio, - } - ) - except Exception as e: - logger.debug(f"Overlap calculation error for GID {gid}: {e}") - continue - - # Sort by overlap ratio (descending) - overlaps.sort(key=lambda x: x["overlap_ratio"], reverse=True) - - if not overlaps: - self._admin2_cache[cache_key] = None - return None - - admin2 = overlaps[0]["admin2_data"] - method_used = "largest overlap" - - # Prepare result - result = { - "gid": int(gid), - "gaul2_code": ( - int(admin2["gaul2_code"]) - if admin2["gaul2_code"] is not None - else None - ), - "gaul2_name": admin2["gaul2_name"], - "iso3_code": admin2["iso3_code"], - "method": method_used, - } - - # Add additional fields if available - if "gaul0_code" in admin2: - result["gaul0_code"] = ( - int(admin2["gaul0_code"]) - if admin2["gaul0_code"] is not None - else None - ) - if "gaul0_name" in admin2: - result["gaul0_name"] = admin2["gaul0_name"] - if "gaul1_code" in admin2: - result["gaul1_code"] = ( - int(admin2["gaul1_code"]) - if admin2["gaul1_code"] is not None - else None - ) - if "gaul1_name" in admin2: - result["gaul1_name"] = admin2["gaul1_name"] - if "continent" in admin2: - result["continent"] = admin2["continent"] - if "disp_en" in admin2: - result["disp_en"] = admin2["disp_en"] - - # Update cache - self._admin2_cache[cache_key] = result - return result - - def find_all_admin_for_gid(self, gid): - """ - Find country, admin1, and admin2 information for a PRIO-GRID cell. - - Parameters: - gid (int): PRIO-GRID cell ID - - Returns: - dict: Combined information for country, admin1, and admin2 - """ - result = {"gid": int(gid)} - - # Find country information - country_info = self.find_country_for_gid(gid) - if country_info: - result["country"] = country_info - - # Find admin1 information - admin1_info = self.find_admin1_for_gid(gid) - if admin1_info: - result["admin1"] = admin1_info - - # Find admin2 information - admin2_info = self.find_admin2_for_gid(gid) - if admin2_info: - result["admin2"] = admin2_info - - return result - - def batch_admin_mapping(self, gid_list, admin_level="admin1"): - """ - Batch mapping for admin1 or admin2 using vectorized operations. - - Parameters: - gid_list (list): List of PRIO-GRID cell IDs - admin_level (str): Either "admin1" or "admin2" - - Returns: - DataFrame: Results of the batch mapping - """ - if admin_level not in ["admin1", "admin2"]: - raise ValueError("admin_level must be either 'admin1' or 'admin2'") - - if admin_level == "admin1" and self.admin1_gdf is None: - raise ValueError("Admin1 data not loaded") - if admin_level == "admin2" and self.admin2_gdf is None: - raise ValueError("Admin2 data not loaded") - - # Select the appropriate find function - if admin_level == "admin1": - find_func = self.find_admin1_for_gid - cache = self._admin1_cache - else: # admin2 - find_func = self.find_admin2_for_gid - cache = self._admin2_cache - - results = [] - cache_hits = 0 - - for gid in gid_list: - # Check cache first - cache_key = self._get_admin_cache_key(gid, admin_level) - if cache_key in cache: - cached_value = cache[cache_key] - if cached_value is not None: - results.append(cached_value.copy()) - else: - results.append(None) - cache_hits += 1 - continue - - # Not in cache, process this GID - result = find_func(gid) - results.append(result) - - if cache_hits > 0: - logger.info( - f"Cache hits: {cache_hits}/{len(gid_list)} ({cache_hits/len(gid_list)*100:.1f}%)" - ) - - return pd.DataFrame([r for r in results if r is not None]) - - def batch_admin_mapping_parallel( - self, gid_list, admin_level="admin1", batch_size=1000 - ): - """ - Batch mapping for admin1 or admin2 using multiprocessing. - - Parameters: - gid_list (list): List of PRIO-GRID cell IDs - admin_level (str): Either "admin1" or "admin2" - batch_size (int): Number of GIDs to process in each batch - - Returns: - DataFrame: Results of the batch mapping - """ - if admin_level not in ["admin1", "admin2"]: - raise ValueError("admin_level must be either 'admin1' or 'admin2'") - - if admin_level == "admin1" and self.admin1_gdf is None: - raise ValueError("Admin1 data not loaded") - if admin_level == "admin2" and self.admin2_gdf is None: - raise ValueError("Admin2 data not loaded") - - self._init_process_pool() - - # Split into batches - batches = [ - gid_list[i : i + batch_size] for i in range(0, len(gid_list), batch_size) - ] - - # Select the appropriate find function - if admin_level == "admin1": - find_func = self.find_admin1_for_gid - else: # admin2 - find_func = self.find_admin2_for_gid - - # Create partial function for mapping - map_func = partial(self._process_gid_batch_admin, find_func=find_func) - - # Process batches in parallel - results = [] - for batch_result in self._process_pool.imap_unordered(map_func, batches): - results.extend(batch_result) - - return pd.DataFrame(results) - - def _process_gid_batch_admin(self, gid_batch, find_func): - """Process a batch of GIDs in a single process for admin mapping""" - batch_results = [] - for gid in gid_batch: - result = find_func(gid) - batch_results.append(result) - return batch_results - - def find_all_admin_for_point(self, point_geometry): - """ - Find country, admin1, and admin2 information for a point. - - Parameters: - point_geometry (shapely.geometry.Point): Point geometry to locate - - Returns: - dict: Combined information for country, admin1, and admin2 - """ - # Find the GID for the point - gid = self.find_gid_for_point(point_geometry) - if gid is None: - return None - - # Use the existing method to find all admin information - return self.find_all_admin_for_gid(gid) - - def extend_find_country_for_gid(self, gid): - """ - Extended version of find_country_for_gid that also includes admin1 and admin2 information. - - Parameters: - gid (int): PRIO-GRID cell ID - - Returns: - dict: Combined information for country, admin1, and admin2 - """ - # Get country information using the existing method - result = self.find_country_for_gid(gid) - - if result is None: - return None - - # Add admin1 information - admin1_info = self.find_admin1_for_gid(gid) - if admin1_info: - result["admin1"] = admin1_info - - # Add admin2 information - admin2_info = self.find_admin2_for_gid(gid) - if admin2_info: - result["admin2"] = admin2_info - - return result - - def visualize_grid_and_admin( - self, gid, admin_level="admin1", show_all_admins=False - ): - """ - Visualize a PRIO-GRID cell and its corresponding admin1 or admin2 boundary. - - Parameters: - gid (int): PRIO-GRID cell ID - admin_level (str): Either "admin1" or "admin2" - show_all_admins (bool): If True, shows all admin regions in the area, - otherwise only shows the one containing the grid cell - - Raises: - ImportError: If matplotlib is not available - ValueError: If the PRIO-GRID cell ID is not found or admin data is not loaded - """ - try: - import matplotlib.pyplot as plt - import matplotlib.patches as mpatches - except ImportError: - logger.error("Matplotlib is required for visualization") - return - - if admin_level not in ["admin1", "admin2"]: - raise ValueError("admin_level must be either 'admin1' or 'admin2'") - - # Check if admin data is loaded - if admin_level == "admin1" and self.admin1_gdf is None: - raise ValueError("Admin1 data not loaded. Cannot visualize admin1.") - if admin_level == "admin2" and self.admin2_gdf is None: - raise ValueError("Admin2 data not loaded. Cannot visualize admin2.") - - # Get the grid cell - grid_cell = self.priogrid_gdf[self.priogrid_gdf["gid"] == gid] - if len(grid_cell) == 0: - raise ValueError(f"PRIO-GRID cell with ID {gid} not found") - - grid_geometry = grid_cell["geometry"].iloc[0] - grid_centroid = grid_cell["centroid"].iloc[0] - - # Select appropriate admin data - if admin_level == "admin1": - admin_gdf = self.admin1_gdf - admin_code_col = "gaul1_code" - admin_name_col = "gaul1_name" - find_func = self.find_admin1_for_gid - else: # admin2 - admin_gdf = self.admin2_gdf - admin_code_col = "gaul2_code" - admin_name_col = "gaul2_name" - find_func = self.find_admin2_for_gid - - # Find the admin region containing the grid cell - admin_info = find_func(gid) - - # Create plot - fig, ax = plt.subplots(1, 1, figsize=(12, 10)) - - if show_all_admins: - # Show all admin regions that intersect with the grid cell's bounds - if admin_level == "admin1" and self.admin1_sindex: - possible_admins_idx = list( - self.admin1_sindex.intersection(grid_geometry.bounds) - ) - nearby_admins = self.admin1_gdf.iloc[possible_admins_idx] - elif admin_level == "admin2" and self.admin2_sindex: - possible_admins_idx = list( - self.admin2_sindex.intersection(grid_geometry.bounds) - ) - nearby_admins = self.admin2_gdf.iloc[possible_admins_idx] - else: - nearby_admins = admin_gdf[admin_gdf.intersects(grid_geometry.buffer(1))] - - # Plot all nearby admin regions - nearby_admins.plot( - ax=ax, color="lightgray", edgecolor="white", alpha=0.5, linewidth=0.5 - ) - - # Highlight the containing admin region if found - if admin_info: - containing_admin = admin_gdf[ - admin_gdf[admin_code_col] == admin_info[admin_code_col] - ] - if len(containing_admin) > 0: - containing_admin.plot( - ax=ax, - color="lightblue", - edgecolor="blue", - alpha=0.7, - linewidth=1.5, - ) - else: - # Only show the containing admin region - if admin_info: - containing_admin = admin_gdf[ - admin_gdf[admin_code_col] == admin_info[admin_code_col] - ] - if len(containing_admin) > 0: - containing_admin.plot( - ax=ax, - color="lightblue", - edgecolor="blue", - alpha=0.7, - linewidth=1.5, - ) - - # Plot the grid cell - grid_cell.plot(ax=ax, color="red", alpha=0.8, edgecolor="darkred", linewidth=2) - - # Add centroid marker - ax.scatter( - grid_centroid.x, - grid_centroid.y, - color="yellow", - s=150, - marker="*", - edgecolor="black", - linewidth=1, - label="Grid Centroid", - zorder=5, - ) - - # Set title and legend - if admin_info: - admin_name = admin_info[admin_name_col] - admin_code = admin_info[admin_code_col] - method = admin_info.get("method", "unknown") - - title = f"PRIO-GRID Cell {gid} in {admin_level.upper()} {admin_name} (Code: {admin_code})\n" - title += f"Method: {method}" - else: - title = f"PRIO-GRID Cell {gid} - No {admin_level.upper()} region found" - - ax.set_title(title, fontsize=14, fontweight="bold") - - # Create legend - legend_elements = [ - mpatches.Patch(color="red", alpha=0.8, label="PRIO-GRID Cell"), - mpatches.Patch( - color="lightblue", alpha=0.7, label=f"Containing {admin_level.upper()}" - ), - mpatches.Patch( - color="lightgray", - alpha=0.5, - label="Other Admin Regions" if show_all_admins else None, - ), - ] - - # Filter out None values - legend_elements = [e for e in legend_elements if e.get_label() is not None] - - ax.legend(handles=legend_elements, loc="upper right", bbox_to_anchor=(1, 1)) - - # Add grid coordinates as text - ax.text( - grid_centroid.x, - grid_centroid.y - 0.1, - f'GID: {gid}\n({grid_cell["xcoord"].iloc[0]:.1f}, {grid_cell["ycoord"].iloc[0]:.1f})', - ha="center", - va="top", - fontsize=9, - bbox=dict(boxstyle="round,pad=0.3", facecolor="white", alpha=0.8), - ) - - # Set aspect ratio and remove unnecessary axes - ax.set_aspect("equal") - ax.set_xlabel("Longitude") - ax.set_ylabel("Latitude") - ax.grid(True, alpha=0.3) - - plt.tight_layout() - plt.show() - - def visualize_admin_regions_for_gids( - self, - gid_list, - admin_level="admin1", - show_grid_cells=True, - color_by_country=True, - figsize=(15, 10), - ): - """ - Visualize admin1 or admin2 regions for multiple PRIO-GRID cells. - - Parameters: - gid_list (list): List of PRIO-GRID cell IDs - admin_level (str): Either "admin1" or "admin2" - show_grid_cells (bool): Whether to show the PRIO-GRID cells on the map - color_by_country (bool): If True, colors admin regions by country - figsize (tuple): Figure size (width, height) - - Raises: - ImportError: If matplotlib is not available - ValueError: If admin data is not loaded - """ - try: - import matplotlib.pyplot as plt - import matplotlib.patches as mpatches - from matplotlib.colors import LinearSegmentedColormap # noqa: F401 - import matplotlib.cm as cm # noqa: F401 - except ImportError: - logger.error("Matplotlib is required for visualization") - return - - if admin_level not in ["admin1", "admin2"]: - raise ValueError("admin_level must be either 'admin1' or 'admin2'") - - # Check if admin data is loaded - if admin_level == "admin1" and self.admin1_gdf is None: - raise ValueError("Admin1 data not loaded. Cannot visualize admin1.") - if admin_level == "admin2" and self.admin2_gdf is None: - raise ValueError("Admin2 data not loaded. Cannot visualize admin2.") - - # Get admin information for all GIDs - if admin_level == "admin1": - admin_results = self.batch_admin_mapping(gid_list, admin_level="admin1") - admin_gdf = self.admin1_gdf - admin_code_col = "gaul1_code" - admin_name_col = "gaul1_name" - else: - admin_results = self.batch_admin_mapping(gid_list, admin_level="admin2") - admin_gdf = self.admin2_gdf - admin_code_col = "gaul2_code" - admin_name_col = "gaul2_name" - - if admin_results.empty: - logger.warning(f"No {admin_level} regions found for the provided GIDs") - return - - # Get unique admin regions - unique_admin_codes = admin_results[admin_code_col].unique() - admin_regions = admin_gdf[admin_gdf[admin_code_col].isin(unique_admin_codes)] - - # Get grid cells - grid_cells = self.priogrid_gdf[self.priogrid_gdf["gid"].isin(gid_list)] - - # Create plot - fig, ax = plt.subplots(1, 1, figsize=figsize) - - # Plot all admin regions in the area - if color_by_country: - # Create a colormap for countries - unique_countries = admin_results["iso3_code"].unique() - country_colors = plt.cm.tab20(np.linspace(0, 1, len(unique_countries))) - country_color_map = { - country: color - for country, color in zip(unique_countries, country_colors) - } - - # Plot with color based on country - for _, admin in admin_regions.iterrows(): - admin_code = admin[admin_code_col] - admin_row = admin_results[admin_results[admin_code_col] == admin_code] - if not admin_row.empty: - country = admin_row["iso3_code"].iloc[0] - color = country_color_map[country] - admin_gdf[admin_gdf[admin_code_col] == admin_code].plot( - ax=ax, color=[color], edgecolor="black", alpha=0.7, linewidth=1 - ) - - # Create legend for countries - legend_elements = [ - mpatches.Patch(color=country_color_map[country], label=country) - for country in unique_countries - ] - else: - # Plot with uniform color - admin_regions.plot( - ax=ax, color="lightblue", edgecolor="black", alpha=0.7, linewidth=1 - ) - legend_elements = [ - mpatches.Patch( - color="lightblue", alpha=0.7, label=f"{admin_level.upper()} Regions" - ) - ] - - # Plot grid cells if requested - if show_grid_cells: - grid_cells.plot( - ax=ax, color="red", alpha=0.8, edgecolor="darkred", linewidth=1.5 - ) - - # Add labels for grid cells - for _, grid_cell in grid_cells.iterrows(): - centroid = grid_cell["centroid"] - ax.text( - centroid.x, - centroid.y, - str(grid_cell["gid"]), - ha="center", - va="center", - fontsize=8, - color="white", - fontweight="bold", - bbox=dict(boxstyle="circle,pad=0.2", facecolor="red", alpha=0.7), - ) - - # Add labels for admin regions - for _, admin in admin_regions.iterrows(): - centroid = admin["geometry"].centroid - admin_name = admin[admin_name_col] - admin_code = admin[admin_code_col] - - # Truncate long names - if len(admin_name) > 20: - admin_name = admin_name[:17] + "..." - - ax.text( - centroid.x, - centroid.y, - f"{admin_name}\n({admin_code})", - ha="center", - va="center", - fontsize=9, - bbox=dict(boxstyle="round,pad=0.3", facecolor="white", alpha=0.8), - ) - - # Set title - ax.set_title( - f"{admin_level.upper()} Regions for {len(gid_list)} PRIO-GRID Cells\n" - f"Found {len(unique_admin_codes)} unique {admin_level.upper()} regions", - fontsize=14, - fontweight="bold", - ) - - # Add grid cells to legend if shown - if show_grid_cells: - legend_elements.insert( - 0, mpatches.Patch(color="red", alpha=0.8, label="PRIO-GRID Cells") - ) - - ax.legend(handles=legend_elements, loc="upper right", bbox_to_anchor=(1, 1)) - - # Set aspect ratio and labels - ax.set_aspect("equal") - ax.set_xlabel("Longitude") - ax.set_ylabel("Latitude") - ax.grid(True, alpha=0.3) - - # Add summary statistics as text - stats_text = f"Total GIDs: {len(gid_list)}\n" - stats_text += ( - f"Unique {admin_level.upper()} regions: {len(unique_admin_codes)}\n" - ) - if color_by_country: - stats_text += f"Unique countries: {len(unique_countries)}" - - ax.text( - 0.02, - 0.98, - stats_text, - transform=ax.transAxes, - fontsize=10, - verticalalignment="top", - bbox=dict(boxstyle="round,pad=0.5", facecolor="white", alpha=0.8), - ) - - plt.tight_layout() - plt.show() - - def find_country_by_iso_a3(self, iso_a3): - """ - Find country information by ISO A3 code from the Natural Earth dataset. - - Parameters: - iso_a3 (str): ISO A3 code of the country (e.g., 'USA', 'FRA', 'TZA') - - Returns: - dict: Country information with all available fields from Natural Earth, - or None if country is not found - """ - # Validate input - if not isinstance(iso_a3, str) or len(iso_a3) != 3: - logger.warning( - f"Invalid ISO A3 code format: {iso_a3}. Expected 3-character string." - ) - return None - - # Convert to uppercase for consistency - iso_a3 = iso_a3.upper() - - if self.use_disk_cache: - # Use disk cache - @self._disk_country_cache.cache() - def _find_country_by_iso_a3_impl(iso_a3): - # Find the country in the Natural Earth dataset - country_rows = self.countries_gdf[ - self.countries_gdf["ISO_A3"] == iso_a3 - ] - - if len(country_rows) == 0: - return None - - # Get the first (and should be only) matching country - country = country_rows.iloc[0] - - # Prepare result with all available country data - result = { - "iso_a3": country["ISO_A3"], - "country_name": country["NAME_EN"], - } - - # Add all other available fields from Natural Earth - for col in country.index: - if col not in ["geometry", "ISO_A3", "NAME_EN"]: - # Convert numpy types to native Python types - value = country[col] - if pd.isna(value): - result[col] = None - elif isinstance(value, (np.integer, np.int64, np.int32)): - result[col] = int(value) - elif isinstance(value, (np.floating, np.float64, np.float32)): - result[col] = float(value) - elif isinstance(value, np.bool_): - result[col] = bool(value) - else: - result[col] = str(value) - - # Add geometry as WKT string - if hasattr(country["geometry"], "wkt"): - result["geometry_wkt"] = country["geometry"].wkt - - # Add some commonly useful derived fields - result["data_source"] = "Natural Earth" - result["centroid"] = { - "lon": float(country["geometry"].centroid.x), - "lat": float(country["geometry"].centroid.y), - } - - # Calculate bounding box - bounds = country["geometry"].bounds - result["bounds"] = { - "min_lon": float(bounds[0]), - "min_lat": float(bounds[1]), - "max_lon": float(bounds[2]), - "max_lat": float(bounds[3]), - } - - return result - - return _find_country_by_iso_a3_impl(iso_a3) - else: - # Use in-memory cache - # Generate consistent cache key - cache_key = self._get_country_cache_key(iso_a3) - - # Check cache first - if cache_key in self._country_cache: - cached_value = self._country_cache[cache_key] - if cached_value is None: - return None - result = cached_value.copy() - return result - - # Find the country in the Natural Earth dataset - country_rows = self.countries_gdf[self.countries_gdf["ISO_A3"] == iso_a3] - - if len(country_rows) == 0: - logger.warning( - f"Country with ISO A3 code '{iso_a3}' not found in Natural Earth dataset" - ) - # Update cache with None to avoid repeated lookups - self._country_cache[cache_key] = None - return None - - # Get the first (and should be only) matching country - country = country_rows.iloc[0] - - # Prepare result with all available country data - result = { - "iso_a3": country["ISO_A3"], - "country_name": country["NAME_EN"], - } - - # Add all other available fields from Natural Earth - for col in country.index: - if col not in ["geometry", "ISO_A3", "NAME_EN"]: - # Convert numpy types to native Python types - value = country[col] - if pd.isna(value): - result[col] = None - elif isinstance(value, (np.integer, np.int64, np.int32)): - result[col] = int(value) - elif isinstance(value, (np.floating, np.float64, np.float32)): - result[col] = float(value) - elif isinstance(value, np.bool_): - result[col] = bool(value) - else: - result[col] = str(value) - - # Add geometry as WKT string - if hasattr(country["geometry"], "wkt"): - result["geometry_wkt"] = country["geometry"].wkt - - # Add some commonly useful derived fields - result["data_source"] = "Natural Earth" - result["centroid"] = { - "lon": float(country["geometry"].centroid.x), - "lat": float(country["geometry"].centroid.y), - } - - # Calculate bounding box - bounds = country["geometry"].bounds - result["bounds"] = { - "min_lon": float(bounds[0]), - "min_lat": float(bounds[1]), - "max_lon": float(bounds[2]), - "max_lat": float(bounds[3]), - } - - # Update cache - self._country_cache[cache_key] = result - return result - - def find_multiple_countries_by_iso_a3(self, iso_a3_list): - """ - Find information for multiple countries by their ISO A3 codes. - - Parameters: - iso_a3_list (list): List of ISO A3 codes (e.g., ['USA', 'FRA', 'TZA']) - - Returns: - dict: Dictionary with ISO A3 codes as keys and country info as values. - Countries not found will have None as values. - """ - if not isinstance(iso_a3_list, (list, tuple, set)): - raise ValueError( - "iso_a3_list must be a list, tuple, or set of ISO A3 codes" - ) - - results = {} - cache_hits = 0 - - for iso_a3 in iso_a3_list: - # Check cache first - cache_key = self._get_country_cache_key(iso_a3.upper()) - if cache_key in self._country_cache: - cached_value = self._country_cache[cache_key] - results[iso_a3.upper()] = ( - cached_value.copy() if cached_value is not None else None - ) - cache_hits += 1 - else: - # Not in cache, fetch the country info - country_info = self.find_country_by_iso_a3(iso_a3) - results[iso_a3.upper()] = country_info - - if cache_hits > 0: - logger.info( - f"Cache hits: {cache_hits}/{len(iso_a3_list)} ({cache_hits/len(iso_a3_list)*100:.1f}%)" - ) - - return results - - def get_country_summary(self, iso_a3): - """ - Get a summary of country information including related grid cells and admin regions. - - Parameters: - iso_a3 (str): ISO A3 code of the country - - Returns: - dict: Summary including country info, grid cell count, and admin region counts - """ - # Get basic country information - country_info = self.find_country_by_iso_a3(iso_a3) - if not country_info: - return None - - # Get all grid cells for this country - gids = self.find_gids_for_country(iso_a3) - - # Count admin1 regions if available - admin1_count = 0 - if self.admin1_gdf is not None: - admin1_count = len(self.admin1_gdf[self.admin1_gdf["iso3_code"] == iso_a3]) - - # Count admin2 regions if available - admin2_count = 0 - if self.admin2_gdf is not None: - admin2_count = len(self.admin2_gdf[self.admin2_gdf["iso3_code"] == iso_a3]) - - # Create summary - summary = { - "country_info": country_info, - "priogrid": { - "gid_count": len(gids), - "gid_list": gids[:100], # Return first 100 GIDs to avoid huge responses - "has_more": len(gids) > 100, - }, - "admin_regions": { - "admin1_count": admin1_count, - "admin2_count": admin2_count, - }, - "data_source": "Natural Earth", - } - - return summary - - def search_countries_by_name( - self, name_pattern, exact_match=False, case_sensitive=False - ): - """ - Search for countries by name pattern. - - Parameters: - name_pattern (str): Name or pattern to search for - exact_match (bool): If True, only exact matches are returned - case_sensitive (bool): If True, search is case sensitive - - Returns: - list: List of matching countries with their information - """ - if not isinstance(name_pattern, str): - raise ValueError("name_pattern must be a string") - - # Prepare the search pattern - if not case_sensitive: - name_pattern = name_pattern.lower() - search_column = self.countries_gdf["NAME_EN"].str.lower() - else: - search_column = self.countries_gdf["NAME_EN"] - - # Find matches - if exact_match: - matching_indices = search_column == name_pattern - else: - matching_indices = search_column.str.contains(name_pattern, na=False) - - matching_countries = self.countries_gdf[matching_indices] - - if len(matching_countries) == 0: - logger.info(f"No countries found matching pattern: {name_pattern}") - return [] - - # Convert matches to list of dictionaries - results = [] - for _, country in matching_countries.iterrows(): - country_info = { - "iso_a3": country["ISO_A3"], - "country_name": country["NAME_EN"], - } - - # Add a few key fields - key_fields = [ - "CONTINENT", - "REGION_UN", - "SUBREGION", - "POP_EST", - "GDP_MD", - "INCOME_GRP", - ] - for field in key_fields: - if field in country: - country_info[field.lower()] = country[field] - - results.append(country_info) - - logger.info(f"Found {len(results)} countries matching pattern: {name_pattern}") - return results - - def enrich_dataframe_with_pg_info( - self, - df, - pg_id_col="priogrid_id", - time_id_col="month_id", - include_country=True, - include_admin1=True, - include_admin2=True, - include_pg_info=True, - country_cols=None, - admin1_cols=None, - admin2_cols=None, - pg_cols=None, - batch_size=1000, - use_multiprocessing=True, - show_progress=True, - only_metadata=True, - ): - """ - Enrich a DataFrame with country, admin1, admin2, and PRIO-GRID information based on PRIO-GRID IDs. - - Parameters: - df (pd.DataFrame): Input DataFrame containing PRIO-GRID IDs - pg_id_col (str): Name of the column containing PRIO-GRID IDs - include_country (bool): Whether to include country information - include_admin1 (bool): Whether to include admin1 information - include_admin2 (bool): Whether to include admin2 information - include_pg_info (bool): Whether to include PRIO-GRID information - country_cols (list): Specific country columns to include (None for all) - admin1_cols (list): Specific admin1 columns to include (None for all) - admin2_cols (list): Specific admin2 columns to include (None for all) - pg_cols (list): Specific PRIO-GRID columns to include (None for all) - batch_size (int): Number of PRIO-GRID IDs to process in each batch - use_multiprocessing (bool): Whether to use multiprocessing for faster processing - show_progress (bool): Whether to show a progress bar - - Returns: - pd.DataFrame: Enriched DataFrame with additional columns - - Raises: - ValueError: If pg_id_col is not found in the DataFrame - """ - # Validate input - if pg_id_col not in df.columns: - raise ValueError(f"Column '{pg_id_col}' not found in DataFrame") - - if only_metadata: - df = pd.DataFrame(df[[pg_id_col, time_id_col]]) - # Create a copy to avoid modifying the original DataFrame - result_df = df.copy() - - # Get unique PRIO-GRID IDs to process - unique_pg_ids = df[pg_id_col].unique() - total_ids = len(unique_pg_ids) - total_batches = (total_ids - 1) // batch_size + 1 - - logger.info( - f"Starting enrichment of {total_ids} unique PRIO-GRID IDs in {total_batches} batches" - ) - logger.info( - f"Parameters: country={include_country}, admin1={include_admin1}, admin2={include_admin2}, pg_info={include_pg_info}" - ) - - # Ensure cache directories exist - if self.use_disk_cache: - for cache_type in ["country", "admin1", "admin2", "gid"]: - cache_path = self.cache_dir / cache_type - if not cache_path.exists(): - os.makedirs(cache_path, exist_ok=True) - logger.info(f"Created cache directory: {cache_path}") - - # Process in batches - all_pg_data = {} - processed_count = 0 - failed_batches = 0 - failed_gids = [] # observability-only; not in return value (C-21) - - # Initialize progress bar if requested - pbar = None - if show_progress: - try: - from tqdm import tqdm - - pbar = tqdm( - total=total_ids, desc="Processing PRIO-GRID IDs", unit="ids" - ) - except ImportError: - logger.warning( - "tqdm not installed. Progress bar will not be shown. Install with: pip install tqdm" - ) - show_progress = False - - if use_multiprocessing and total_ids > batch_size: - # Use threading instead of multiprocessing to avoid pickling issues - from concurrent.futures import ThreadPoolExecutor, as_completed - - # Split into batches - batches = [ - unique_pg_ids[i : i + batch_size] - for i in range(0, total_ids, batch_size) - ] - - # Process batches in parallel using threads - with ThreadPoolExecutor(max_workers=self._max_workers) as executor: - # Submit all batches for processing - future_to_batch = { - executor.submit( - self._process_pg_batch, - batch_ids, - include_country, - include_admin1, - include_admin2, - include_pg_info, - country_cols, - admin1_cols, - admin2_cols, - pg_cols, - ): batch_idx - for batch_idx, batch_ids in enumerate(batches) - } - - # Collect results as they complete - for future in as_completed(future_to_batch): - batch_idx = future_to_batch[future] - try: - batch_result = future.result() - all_pg_data.update(batch_result) - processed_count += len(batches[batch_idx]) - - # Update progress - if pbar: - pbar.update(len(batches[batch_idx])) - - # Log progress every 10% or every 5 batches - if ( - batch_idx % max(5, total_batches // 10) == 0 - or batch_idx == total_batches - 1 - ): - progress_pct = (processed_count / total_ids) * 100 - logger.info( - f"Progress: {processed_count}/{total_ids} ({progress_pct:.1f}%) - Batch {batch_idx + 1}/{total_batches}" - ) - - except Exception as e: - failed_batches += 1 - failed_gids.extend(list(batches[batch_idx])) - logger.error(f"Batch {batch_idx} failed ({len(batches[batch_idx])} GIDs affected): {str(e)}") - continue - else: - # Process sequentially for smaller datasets - for i in range(0, total_ids, batch_size): - batch_ids = unique_pg_ids[i : i + batch_size] - batch_num = i // batch_size + 1 - - logger.info( - f"Processing batch {batch_num}/{total_batches} ({len(batch_ids)} IDs)" - ) - - try: - batch_result = self._process_pg_batch( - batch_ids, - include_country, - include_admin1, - include_admin2, - include_pg_info, - country_cols, - admin1_cols, - admin2_cols, - pg_cols, - ) - all_pg_data.update(batch_result) - processed_count += len(batch_ids) - - # Update progress - if pbar: - pbar.update(len(batch_ids)) - - # Log progress - progress_pct = (processed_count / total_ids) * 100 - logger.info( - f"Completed batch {batch_num}/{total_batches} - {progress_pct:.1f}% complete" - ) - - except Exception as e: - failed_batches += 1 - failed_gids.extend(list(batch_ids)) - logger.error(f"Batch {batch_num} failed ({len(batch_ids)} GIDs affected): {str(e)}") - continue - - if failed_batches > 0: - logger.error(f"Enrichment had {failed_batches} failed batch(es) affecting {len(failed_gids)} GIDs. Output will have NaN metadata for these GIDs.") - - # Close progress bar - if pbar: - pbar.close() - - # Convert the collected data to a DataFrame - logger.info("Converting collected data to DataFrame") - pg_info_df = pd.DataFrame.from_dict(all_pg_data, orient="index") - - # Merge with the original DataFrame - if not pg_info_df.empty: - logger.info("Merging enriched data with original DataFrame") - result_df = result_df.merge( - pg_info_df, left_on=pg_id_col, right_on="pg_id", how="left" - ) - - # Drop the redundant pg_id column - if "pg_id" in result_df.columns: - result_df.drop(columns=["pg_id"], inplace=True) - else: - logger.warning("No enrichment data was generated") - - logger.info( - f"Enrichment complete. Added {len(result_df.columns) - len(df.columns)} new columns" - ) - return result_df - - def _process_pg_batch( - self, - pg_ids, - include_country, - include_admin1, - include_admin2, - include_pg_info, - country_cols, - admin1_cols, - admin2_cols, - pg_cols, - ): - """Process a batch of PRIO-GRID IDs and return their information""" - batch_data = {} - - for pg_id in pg_ids: - pg_data = {"pg_id": pg_id} - - # Get PRIO-GRID information - if include_pg_info: - grid_cell = self.priogrid_gdf[self.priogrid_gdf["gid"] == pg_id] - if len(grid_cell) > 0: - grid_info = grid_cell.iloc[0].to_dict() - - # Select specific columns if requested - if pg_cols is not None: - for col in pg_cols: - if col in grid_info: - pg_data[f"pg_{col}"] = grid_info[col] - else: - # Include all grid info with pg_ prefix - for key, value in grid_info.items(): - if key != "gid": # Skip gid as it's already the key - pg_data[f"pg_{key}"] = value - - # Get country information - if include_country: - country_info = self.find_country_for_gid(pg_id) - if country_info: - # Select specific columns if requested - if country_cols is not None: - for col in country_cols: - if col in country_info: - pg_data[f"country_{col}"] = country_info[col] - else: - # Include all country info with country_ prefix - for key, value in country_info.items(): - if key != "gid": # Skip gid as it's already the key - pg_data[f"country_{key}"] = value - - # Get admin1 information - if include_admin1: - admin1_info = self.find_admin1_for_gid(pg_id) - if admin1_info: - # Select specific columns if requested - if admin1_cols is not None: - for col in admin1_cols: - if col in admin1_info: - pg_data[f"admin1_{col}"] = admin1_info[col] - else: - # Include all admin1 info with admin1_ prefix - for key, value in admin1_info.items(): - if key != "gid": # Skip gid as it's already the key - pg_data[f"admin1_{key}"] = value - - # Get admin2 information - if include_admin2: - admin2_info = self.find_admin2_for_gid(pg_id) - if admin2_info: - # Select specific columns if requested - if admin2_cols is not None: - for col in admin2_cols: - if col in admin2_info: - pg_data[f"admin2_{col}"] = admin2_info[col] - else: - # Include all admin2 info with admin2_ prefix - for key, value in admin2_info.items(): - if key != "gid": # Skip gid as it's already the key - pg_data[f"admin2_{key}"] = value - - batch_data[pg_id] = pg_data - - return batch_data - - def enrich_dataframe_with_country_info( - self, - df, - iso_a3_col="iso_a3", - include_admin1=True, - include_admin2=True, - country_cols=None, - admin1_cols=None, - admin2_cols=None, - batch_size=1000, - use_multiprocessing=True, - show_progress=True, - ): - """ - Enrich a DataFrame with country, admin1, and admin2 information based on ISO A3 codes. - - Parameters: - df (pd.DataFrame): Input DataFrame containing ISO A3 codes - iso_a3_col (str): Name of the column containing ISO A3 codes - include_admin1 (bool): Whether to include admin1 information - include_admin2 (bool): Whether to include admin2 information - country_cols (list): Specific country columns to include (None for all) - admin1_cols (list): Specific admin1 columns to include (None for all) - admin2_cols (list): Specific admin2 columns to include (None for all) - batch_size (int): Number of ISO A3 codes to process in each batch - use_multiprocessing (bool): Whether to use multiprocessing for faster processing - show_progress (bool): Whether to show a progress bar - - Returns: - pd.DataFrame: Enriched DataFrame with additional columns - - Raises: - ValueError: If iso_a3_col is not found in the DataFrame - """ - # Validate input - if iso_a3_col not in df.columns: - raise ValueError(f"Column '{iso_a3_col}' not found in DataFrame") - - # Create a copy to avoid modifying the original DataFrame - result_df = df.copy() - - # Get unique ISO A3 codes to process - unique_iso_a3 = df[iso_a3_col].unique() - total_codes = len(unique_iso_a3) - total_batches = (total_codes - 1) // batch_size + 1 - - logger.info( - f"Starting enrichment of {total_codes} unique ISO A3 codes in {total_batches} batches" - ) - logger.info(f"Parameters: admin1={include_admin1}, admin2={include_admin2}") - - # Prepare parameters for batch processing - process_params = { - "include_admin1": include_admin1, - "include_admin2": include_admin2, - "country_cols": country_cols, - "admin1_cols": admin1_cols, - "admin2_cols": admin2_cols, - } - - # Process in batches - all_country_data = {} - processed_count = 0 - failed_batches = 0 - failed_codes = [] # observability-only; not in return value (C-21) - - # Initialize progress bar if requested - pbar = None - if show_progress: - try: - from tqdm import tqdm - - pbar = tqdm( - total=total_codes, desc="Processing ISO A3 codes", unit="codes" - ) - except ImportError: - logger.warning( - "tqdm not installed. Progress bar will not be shown. Install with: pip install tqdm" - ) - show_progress = False - - if use_multiprocessing and total_codes > batch_size: - # Use threading instead of multiprocessing to avoid pickling issues - from concurrent.futures import ThreadPoolExecutor, as_completed - - # Split into batches - batches = [ - unique_iso_a3[i : i + batch_size] - for i in range(0, total_codes, batch_size) - ] - - # Process batches in parallel using threads - with ThreadPoolExecutor(max_workers=self._max_workers) as executor: - # Submit all batches for processing - future_to_batch = { - executor.submit( - self._process_country_batch, batch_codes, **process_params - ): batch_idx - for batch_idx, batch_codes in enumerate(batches) - } - - # Collect results as they complete - for future in as_completed(future_to_batch): - batch_idx = future_to_batch[future] - try: - batch_result = future.result() - all_country_data.update(batch_result) - processed_count += len(batches[batch_idx]) - - # Update progress - if pbar: - pbar.update(len(batches[batch_idx])) - - # Log progress every 10% or every 5 batches - if ( - batch_idx % max(5, total_batches // 10) == 0 - or batch_idx == total_batches - 1 - ): - progress_pct = (processed_count / total_codes) * 100 - logger.info( - f"Progress: {processed_count}/{total_codes} ({progress_pct:.1f}%) - Batch {batch_idx + 1}/{total_batches}" - ) - - except Exception as e: - failed_batches += 1 - failed_codes.extend(list(batches[batch_idx])) - logger.error(f"Batch {batch_idx} failed ({len(batches[batch_idx])} codes affected): {str(e)}") - continue - else: - # Process sequentially for smaller datasets - for i in range(0, total_codes, batch_size): - batch_codes = unique_iso_a3[i : i + batch_size] - batch_num = i // batch_size + 1 - - logger.info( - f"Processing batch {batch_num}/{total_batches} ({len(batch_codes)} codes)" - ) - - try: - batch_result = self._process_country_batch( - batch_codes, **process_params - ) - all_country_data.update(batch_result) - processed_count += len(batch_codes) - - # Update progress - if pbar: - pbar.update(len(batch_codes)) - - # Log progress - progress_pct = (processed_count / total_codes) * 100 - logger.info( - f"Completed batch {batch_num}/{total_batches} - {progress_pct:.1f}% complete" - ) - - except Exception as e: - failed_batches += 1 - failed_codes.extend(list(batch_codes)) - logger.error(f"Batch {batch_num} failed ({len(batch_codes)} codes affected): {str(e)}") - continue - - if failed_batches > 0: - logger.error(f"Enrichment had {failed_batches} failed batch(es) affecting {len(failed_codes)} ISO A3 codes. Output will have NaN metadata for these codes.") - - # Close progress bar - if pbar: - pbar.close() - - # Convert the collected data to a DataFrame - logger.info("Converting collected data to DataFrame") - country_info_df = pd.DataFrame.from_dict(all_country_data, orient="index") - - # Merge with the original DataFrame - if not country_info_df.empty: - logger.info("Merging enriched data with original DataFrame") - result_df = result_df.merge( - country_info_df, left_on=iso_a3_col, right_on="iso_a3_code", how="left" - ) - - # Drop the redundant iso_a3_code column - if "iso_a3_code" in result_df.columns: - result_df.drop(columns=["iso_a3_code"], inplace=True) - else: - logger.warning("No enrichment data was generated") - - logger.info( - f"Enrichment complete. Added {len(result_df.columns) - len(df.columns)} new columns" - ) - return result_df - - def _process_country_batch( - self, - iso_a3_codes, - include_admin1, - include_admin2, - country_cols, - admin1_cols, - admin2_cols, - ): - """Process a batch of ISO A3 codes and return their information""" - batch_data = {} - - for iso_a3 in iso_a3_codes: - # Skip empty or invalid codes - if pd.isna(iso_a3) or not isinstance(iso_a3, str) or len(iso_a3) != 3: - batch_data[iso_a3] = {"iso_a3_code": iso_a3} - continue - - country_data = {"iso_a3_code": iso_a3} - - # Get country information - country_info = self.find_country_by_iso_a3(iso_a3) - if country_info: - # Select specific columns if requested - if country_cols is not None: - for col in country_cols: - if col in country_info: - country_data[f"country_{col}"] = country_info[col] - else: - # Include all country info with country_ prefix - for key, value in country_info.items(): - if key != "iso_a3": # Skip iso_a3 as it's already the key - country_data[f"country_{key}"] = value - - # Get admin1 information - if include_admin1 and self.admin1_gdf is not None: - admin1_regions = self.admin1_gdf[self.admin1_gdf["iso3_code"] == iso_a3] - - if not admin1_regions.empty: - # Create a list of admin1 regions - admin1_list = [] - for _, admin1 in admin1_regions.iterrows(): - admin1_dict = {} - - # Select specific columns if requested - if admin1_cols is not None: - for col in admin1_cols: - if col in admin1: - admin1_dict[col] = admin1[col] - else: - # Include all admin1 info - for key, value in admin1.items(): - if ( - key != "iso3_code" - ): # Skip iso3_code as it's already known - admin1_dict[key] = value - - admin1_list.append(admin1_dict) - - country_data["admin1_regions"] = admin1_list - country_data["admin1_count"] = len(admin1_list) - - # Get admin2 information - if include_admin2 and self.admin2_gdf is not None: - admin2_regions = self.admin2_gdf[self.admin2_gdf["iso3_code"] == iso_a3] - - if not admin2_regions.empty: - # Create a list of admin2 regions - admin2_list = [] - for _, admin2 in admin2_regions.iterrows(): - admin2_dict = {} - - # Select specific columns if requested - if admin2_cols is not None: - for col in admin2_cols: - if col in admin2: - admin2_dict[col] = admin2[col] - else: - # Include all admin2 info - for key, value in admin2.items(): - if ( - key != "iso3_code" - ): # Skip iso3_code as it's already known - admin2_dict[key] = value - - admin2_list.append(admin2_dict) - - country_data["admin2_regions"] = admin2_list - country_data["admin2_count"] = len(admin2_list) - - batch_data[iso_a3] = country_data - - return batch_data - - def add_country_info_to_dataframe( - self, df, iso_a3_col="iso_a3", name_col="country_name", additional_cols=None - ): - """ - Add basic country information to a DataFrame based on ISO A3 codes. - - Parameters: - df (pd.DataFrame): Input DataFrame containing ISO A3 codes - iso_a3_col (str): Name of the column containing ISO A3 codes - name_col (str): Name of the column to create for country names - additional_cols (list): Additional country columns to include - - Returns: - pd.DataFrame: DataFrame with added country information - """ - # Create a copy to avoid modifying the original DataFrame - result_df = df.copy() - - # Get unique ISO A3 codes to process - unique_iso_a3 = df[iso_a3_col].unique() - logger.info( - f"Processing {len(unique_iso_a3)} unique ISO A3 codes for country names" - ) - - # Create a mapping of ISO A3 to country name - iso_to_name = {} - iso_to_additional = {col: {} for col in additional_cols or []} - - for iso_a3 in unique_iso_a3: - # Skip empty or invalid codes - if pd.isna(iso_a3) or not isinstance(iso_a3, str) or len(iso_a3) != 3: - iso_to_name[iso_a3] = None - for col in iso_to_additional: - iso_to_additional[col][iso_a3] = None - continue - - country_info = self.find_country_by_iso_a3(iso_a3) - if country_info: - iso_to_name[iso_a3] = country_info.get("country_name") - - # Add additional columns if requested - for col in iso_to_additional: - iso_to_additional[col][iso_a3] = country_info.get(col) - else: - iso_to_name[iso_a3] = None - for col in iso_to_additional: - iso_to_additional[col][iso_a3] = None - - # Add the country names to the DataFrame - result_df[name_col] = result_df[iso_a3_col].map(iso_to_name) - - # Add additional columns if requested - for col in iso_to_additional: - result_df[f"country_{col}"] = result_df[iso_a3_col].map( - iso_to_additional[col] - ) - - logger.info("Added country information to DataFrame") - return result_df - - -# Global default mapper -_DEFAULT_MAPPER = None - - -def set_default_mapper(): - """ - Set the default PriogridCountryMapper instance to be used by SpatialEntity classes. - - This function creates a global default mapper that will be automatically used - by Point, Priogrid, and Country classes when no explicit mapper is provided. - - Returns: - PriogridCountryMapper: The created default mapper instance - """ - global _DEFAULT_MAPPER - _DEFAULT_MAPPER = PriogridCountryMapper() - return _DEFAULT_MAPPER - - -def get_default_mapper(): - """ - Get the default PriogridCountryMapper instance. - - Returns: - PriogridCountryMapper: The default mapper instance - - Raises: - ValueError: If no default mapper has been set - """ - if _DEFAULT_MAPPER is None: - raise ValueError("No default mapper set. Call set_default_mapper() first.") - return _DEFAULT_MAPPER - - -# Set default mapper -set_default_mapper()