diff --git a/CEPM/batch-log.md b/CEPM/batch-log.md index 5fd24a885..fde6ee9f9 100644 --- a/CEPM/batch-log.md +++ b/CEPM/batch-log.md @@ -17,7 +17,7 @@ was run and what it showed. Loosely inspired by [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). ## TEMPLATE Batch entry - +**## Author: Tyler Fitch`** **## Batch name: TEMPLATE `{Batch name here, e.g., v20260903}`** Summary: What'd we change, why, what'd we find, what's next @@ -57,7 +57,55 @@ Summary: What'd we change, why, what'd we find, what's next - [ ] I edited LLM-generated text to keep this entry short and to the point --- +##Batch entry +**## Author: Gaby** +**## Batch name: v20261001gt`** + +Summary: Re-ran the st-AZNM baseline, limit-RE, and optimized cases to test a fix for a pandas 3.x issue that caused ReEDS to drop buildable onshore wind in landlocked regions. The fix ensures wind supply-curve bins retain the correct data type, preventing onshore wind from being removed during input processing. The corrected runs successfully restored onshore wind as a build option. + +### Batch details + +- **Built from:** 20260925-test +- **Space and Time:** st/AZ.NM, 2026–2032 every 3 years +- **Cases:** st-AZNM_baseline, st-AZNM_limitre, st-AZNM_optimized +- **Run comments:** Successful run of all three cases on fix/wind-rsc-pandas3. + +### Change log + +- **ReEDS change:** Changed reeds/input_processing/writesupplycurves.py in agg_supplycurve() so an empty supply curve's bin column is explicitly assigned int64 rather than being created from a bare []. +- **Reason for change:** Under pandas 3.x, an empty offshore-wind supply curve could cause the populated onshore-wind bin values to be upcast from integers to floats during pd.concat. This changed labels such as wsc1 to wsc1.0, causing ReEDS to silently drop the onshore-wind cap and cost rows. This was occurring for st/AZ.NM because the region has no offshore-wind resource while GSw_OfsWind=1. +- **Case change :No substantive case-setting changes; this batch re-runs the existing AZNM baseline, limit-RE, and optimized cases with the wind supply-curve fix. + +### Results +All three cases completed successfully. + +Compare-cases output: +runs/v20261001gt_st-AZNM_baseline/outputs/comparisons/ + +Onshore wind returned as a build option after the fix. Before the fix, all three AZNM cases remained at roughly 0.7 GW of existing onshore wind through 2032, with effectively no new wind builds. + +After the fix, substantial new onshore wind is built. By 2032, onshore wind capacity increases to roughly 11 GW in the baseline, 12 GW in limit-RE, and 20 GW in the optimized case. + +The effect is largest in the optimized case, where restoring the wind supply curve results in roughly 19 GW of additional onshore wind capacity by 2032 relative to the broken run. + +The corrected wind availability also changes the broader capacity mix, confirming that the pandas 3.x supply-curve bug was materially affecting the model's resource-selection results. + +### Documentation, decisions, next steps, & issues + +- **PROPOSED DECISION:** Merge fix into dev +- **DOCUMENTATION:** The pandas 3.x wind supply-curve issue and fix are documented in CEPM/known-reeds-issues.md and CEPM/reeds-to-cepm-log.md. +- **NEXT STEP:** Merge fix into dev +- **ISSUE:** The underlying bug also exists upstream and can affect other landlocked regions when offshore wind is enabled but the offshore supply curve is empty. +### Checklist + +- [X] I updated [`known-reeds-issues.md`](CEPM/known-reeds-issues.md) with any run-breaking issues I encountered +- [X] I updated [`reeds-to-cepm-log.md`](CEPM/reeds-to-cepm-log.md) with any changes to ReEDS files +- [X] I added any decisions to [`CEPM/decisions/`](CEPM/decisions/) and linked to them here +- [X] I moved batch results to the VM-Outputs folder +- [X] I added any issues we found to the JIRA issues epic +- [X] I edited LLM-generated text to keep this entry short and to the point +- ## Batch name: `v20260929_st-AZNM...`, `v20260929v2_st-MSALGA`, `v20260929v3_st-VA` This batch implements state-level regions (while still using z134 zones) to @@ -89,17 +137,17 @@ create a few additional testing geographies for st/MS.AL.GA and st/VA. - **ISSUE**: An underlying ReEDS issue is causing the VA-limitre runs to be infeasible. - **ISSUE**: If you run several multistep CEPM runs with the same batch name, by default the run_cepm.ps1's will compare all of them instead of isolating comparisons to the thrhee youu're looking at. + ### Checklist - [X] I updated [`known-reeds-issues.md`](CEPM/known-reeds-issues.md) with any run-breaking issues I encountered - [X] I updated [`reeds-to-cepm-log.md`](CEPM/reeds-to-cepm-log.md) with any changes to ReEDS files - [X] I added any decisions to [`CEPM/decisions/`](CEPM/decisions/) and linked to them here -- [/] I moved batch results to the VM-Outputs folder +- [X] I moved batch results to the VM-Outputs folder - [X] I added any issues we found to the JIRA issues epic - [X] I edited LLM-generated text to keep this entry short and to the point --- - ## Batch name: `v20260903qoff` / `v20260903h2off` / `v20260903h2qoff` These three batches explore the impacts of the interconnection queue penalty diff --git a/CEPM/known-reeds-issues.md b/CEPM/known-reeds-issues.md index 51e4b316f..470af95d7 100644 --- a/CEPM/known-reeds-issues.md +++ b/CEPM/known-reeds-issues.md @@ -551,14 +551,14 @@ same crash if run with `GSw_OfsWind=0`. Good candidate to contribute back, since it's the same fix pattern upstream already uses for `GSw_distpv`/`GSw_CSP` a few lines below. -## Onshore wind supply curve silently dropped under pandas 3 when a sibling curve is empty +## Onshore wind supply curve silently dropped under pandas 3.x when a sibling supply curve is empty (FIXED) -**Symptom:** a run completes normally — no error, no warning, `writesupplycurves.py` -logs `Starting`/`Finished` as usual — but builds **no new onshore wind at all**. -Total `wind-ons` capacity sits flat at the existing fleet for every solve year. The -tell is in `inputs_case/rsc_combined.csv`: `wind-ons` has only `cost_cap` and -`cost_trans` rows, and no `cap` or `cost` rows, so the model is handed zero buildable -wind resource. To check any run: +**Symptom:** a run completes normally, with no error or warning, but builds **no new +onshore wind at all** — total `wind-ons` capacity stays flat at its existing value for +every solve year. `writesupplycurves.py` logs `Starting`/`Finished` as usual. The tell +is in `inputs_case/rsc_combined.csv`: `wind-ons` has only `cost_cap` and `cost_trans` +rows, and no `cap` or `cost` rows at all, so the model is handed zero buildable wind +resource. To check any run: ```bash awk -F, 'NR>1{split($1,a,"_"); print a[1]"|"$3}' inputs_case/rsc_combined.csv \ @@ -566,32 +566,26 @@ awk -F, 'NR>1{split($1,a,"_"); print a[1]"|"$3}' inputs_case/rsc_combined.csv \ ``` A healthy run shows four categories (`cap`, `cost`, `cost_cap`, `cost_trans`) with -equal row counts. An affected run is missing `cap` and `cost` entirely: +equal row counts; an affected run is missing `cap` and `cost`. -``` - 377 cap 377 cost 377 cost_cap 377 cost_trans <- healthy (v20260916v4_st-AZNM_baseline) - 377 cost_cap 377 cost_trans <- broken (v20260924fix_st-AZNM_baseline) -``` - -First surfaced 2026-09-24 on `runs/v20260924fix_st-AZNM_*`, which built 0.7 GW of wind -in 2032 (the existing fleet, nothing new) against 14.2 GW (`baseline`/`limitre`) and -15.8 GW (`optimized`) in the otherwise-comparable `runs/v20260916v4_st-AZNM_*` eight -days earlier. Switches, `numbins_*` and siting scenarios were identical between the -two. +Surfaced 2026-09-24 on `runs/v20260924fix_st-AZNM_*`, which built 0.7 GW of wind (the +existing fleet, nothing new) in 2032 against 14.2 GW (`baseline`/`limitre`) and 15.8 GW +(`optimized`) in the otherwise-comparable `runs/v20260916v4_st-AZNM_*` eight days +earlier. Switches, `numbins_*`, and siting scenarios were identical between the two. **Root cause:** a dtype bug in `agg_supplycurve()` (`reeds/input_processing/writesupplycurves.py`) that only became live with pandas 3.0. `st-AZNM` is landlocked, so its `supplycurve_wind-ofs.csv` is header-only — but `GSw_OfsWind=1`, so offshore wind is still processed. For an empty input, -`agg_supplycurve()` takes its `if dfin.empty:` branch and assigns `dfin['bin'] = []`, -which pandas types as **float64**. pandas 3.0 stopped excluding empty frames from -dtype resolution in `pd.concat`, so `windall = pd.concat(wind, axis=0)` — which stacks -the onshore and offshore frames row-wise into a MultiIndex — upcasts the `bin` level -from int64 to float64. The next line, -`windall["bin"] = "wsc" + windall["bin"].astype(str)`, then produces `wsc1.0` instead -of `wsc1`, so the rename map `{"wsc1": "bin1", ...}` matches nothing and the wind bin -columns keep their `wsc*.0` names. +`agg_supplycurve()` took its `if dfin.empty:` branch and assigned `dfin['bin'] = []`, +which pandas types as **float64**. pandas 3.0 stopped excluding empty frames from dtype +resolution in `pd.concat`, so `windall = pd.concat(wind, axis=0)` — which stacks the +onshore and offshore frames row-wise into a MultiIndex — upcast the `bin` level from +int64 to float64. The next line, +`windall["bin"] = "wsc" + windall["bin"].astype(str)`, then produced `wsc1.0` instead +of `wsc1`, so the rename map `{"wsc1": "bin1", ...}` matched nothing and the wind bin +columns kept their `wsc*.0` names. From there the loss is silent by construction: `alloutcap.pivot(...)` selects its value columns with `[c for c in alloutcap.columns if c.startswith("bin")]`, which wind no @@ -600,46 +594,47 @@ longer has, so all wind capacity is dropped without complaint. The final rows too. Only `cost_cap`/`cost_trans` survive, because those are concatenated on *after* that pivot. -UPV, CSP, geohydro and EGS escape only incidentally — see Scope. `class` escapes for a -similar accidental reason: it is read from the CSV, which pandas types as object for an -empty file, rather than being assigned a bare `[]`. +UPV, CSP, geohydro and EGS escaped only incidentally: UPV is never concatenated with an +empty frame, so its `bin` stays int64. `class` escaped for a similar accidental reason — +it is read from the CSV, which pandas types as object for an empty file, rather than +being assigned a bare `[]`. -**Trigger:** this fork's pandas pin moved to `pandas==3.0.*` on 2026-09-10 (`129ecdbc`, +**Trigger:** the repo's pandas pin moved to `pandas==3.0.*` on 2026-09-10 (`129ecdbc`, "Bump Python to 3.14, realign packages with environment.yml"), and the local venv was rebuilt to pandas 3.0.5 on 2026-09-24 at 20:24 — 21 minutes before the first affected run started. `uv.lock` carried `pandas 2.0.3` from 2026-05-05 until that bump, so every run before 2026-09-24 was unaffected with identical inputs and switches. This is a -pre-existing upstream bug that the pandas upgrade activated, not a regression -introduced by the pin — and note upstream pinned `pandas=3.0` four months before we did -(see **Fixed upstream?** below). Do not "fix" it by pinning pandas back. +pre-existing upstream bug that the pandas upgrade activated, not a regression introduced +by the pin — and note upstream pinned `pandas=3.0` four months before we did (see +**Fixed upstream?** below). Do not "fix" it by pinning pandas back. **Scope — all four conditions must hold at once:** 1. **pandas >= 3.0.** The concat dtype-resolution change is what makes the latent bug live. Confirmed on 3.0.5: concatenating a populated int64-`bin` frame with an empty - float64-`bin` one yields `['wsc1.0', 'wsc2.0', ...]`; the populated frame alone - yields `['wsc1', 'wsc2', ...]`. + float64-`bin` one yields `['wsc1.0', 'wsc2.0', ...]`; the populated frame alone yields + `['wsc1', 'wsc2', ...]`. 2. **The code path stacks sub-techs with `pd.concat(dict, axis=0)` and then does `.astype(str)` on the `bin` level.** Only two sites qualify: wind (`pd.concat(wind, axis=0)` over `ons`/`ofs`, then `"wsc" + bin.astype(str)`) and geothermal (`pd.concat(geo, axis=0)` over `geohydro`/`egs`, then `"geosc" + bin.astype(str)`). **UPV and CSP cannot hit this** — they call - `agg_supplycurve()` standalone, so an empty result just stays empty with no - populated sibling to corrupt. -3. **One member of that dict is empty and the other is not.** Both populated, no - upcast; both empty, nothing to lose. Only the *mixed* case does damage, because the - empty frame's float64 silently rewrites the populated frame's bin labels. + `agg_supplycurve()` standalone, so an empty result just stays empty with no populated + sibling to corrupt. +3. **One member of that dict is empty and the other is not.** Both populated, no upcast; + both empty, nothing to lose. Only the *mixed* case does damage, because the empty + frame's float64 silently rewrites the populated frame's bin labels. 4. **The empty member is still being processed** — its switch is on, so it reaches the concat even though it has no resource. For wind that reduces to: **offshore wind enabled in a region with no offshore -resource.** Verified against the affected `st-AZNM` inputs — `GSw_OfsWind=1` yields -0 `cap`/0 `cost` wind rows, `GSw_OfsWind=0` yields 1640/1640. It is leaving the switch -*on* over an empty curve that bites; turning it *off* is safe. +resource.** Verified against the unfixed code on the same `st-AZNM` inputs — +`GSw_OfsWind=1` gives 0 `cap`/0 `cost` wind rows, `GSw_OfsWind=0` gives 1640/1640. It is +leaving the switch *on* over an empty curve that bites; turning it *off* is safe. `GSw_OfsWind` defaults to **1** in `cases.csv`, so every case satisfies condition 4 unless it explicitly opts out. Which regions satisfy condition 3, measured by offshore -supply curve rows across existing runs: +supply curve rows in existing runs: | Region | offshore rows | exposed? | |---|---|---| @@ -650,74 +645,72 @@ supply curve rows across existing runs: | `transreg/SERTP` | 2848 | no | | `country/USA` | 16191 | no | -**`nercr/WECC_SW` is exposed and has not been re-run since the pandas upgrade** — its -existing runs predate it and are clean, but the next WECC-SW run will lose its wind the -same way. +**`nercr/WECC_SW` is exposed and had not been re-run since the pandas upgrade** — its +existing runs predate it and are clean, but the next WECC-SW run on unfixed code would +have lost its wind the same way. A sweep of every run under `runs/` found only the three +`v20260924fix_st-AZNM_*` cases actually affected; `v20260925_SERTP_*` and +`v20260925_VA_*` also post-date the upgrade but have non-empty offshore curves, so their +wind `cap` rows are intact. -**The geothermal site is latent, not live.** `geoall = pd.concat(geo, axis=0)` followed -by `"geosc" + geoall["bin"].astype(str)` is the same construct, and fails condition 3 -only by accident: `rev_geo_types` is built from whichever of `geohydrosupplycurve` / +**The geothermal site is latent, not live.** `geoall = pd.concat(geo, axis=0)` followed by +`"geosc" + geoall["bin"].astype(str)` is the same construct, and fails condition 3 only by +accident: `rev_geo_types` is built from whichever of `geohydrosupplycurve` / `egssupplycurve` equals `reV`, and current switches set `egssupplycurve=reV` with `geohydrosupplycurve=ATB_2023`, leaving a single-element dict with nothing to concat against. Set both to `reV` in a region where one is empty and it would bite identically. +The fix covers it, being at the shared source. **Impact:** severe and silent — this is the dangerous kind. The run completes, every output file is written, and the results look plausible; wind is simply absent from the build. Anything downstream of capacity (generation mix, system cost, emissions, prices, PRAS) is wrong in a way no error surfaces. Because the failure mode is a missing input -rather than a crash, affected runs have to be identified by inspecting -`rsc_combined.csv` — scanning logs will not find them. +rather than a crash, affected runs must be identified by inspecting +`rsc_combined.csv`, not by scanning logs. -**Status:** not fixed. Diagnosed, with a validated candidate fix not yet applied. - -The candidate fix is one line — type the empty column to match the int64 that +**Status:** fixed, on `fix/wind-rsc-pandas3`. `agg_supplycurve()`'s empty branch now +assigns `pd.Series([], dtype='int64')` instead of `[]`, matching the int64 that `reeds.inputs.get_bin` produces in the non-empty branch, so the `bin` level never -upcasts: +upcasts. Fixed at the dtype source rather than at the `.astype(str)` call, because the +same trap sits one line below on `class` and because `agg_supplycurve()` also feeds the +geothermal `pd.concat(geo, axis=0)`, which is one switch change away from the same +failure (condition 3 above). -```python - if dfin.empty: -- dfin['bin'] = [] -+ dfin['bin'] = pd.Series([], dtype='int64') -``` +Verified against the real `v20260924fix_st-AZNM_baseline` inputs: restores 377 `cap` and +377 `cost` rows for `wind-ons` and 920.3 GW of buildable resource, with the `cap` rows +byte-identical to the pre-pandas-3 `v20260916v4` run. UPV output is unchanged before and +after the fix. **The three `v20260924fix_st-AZNM_*` runs need to be re-run**; their +results are not usable. -Fixing at the dtype source rather than at the `.astype(str)` call is deliberate: the -same trap sits one line below on `class`, and `agg_supplycurve()` also feeds the -geothermal concat above. Validated offline against the real -`v20260924fix_st-AZNM_baseline` inputs — restores 377 `cap` and 377 `cost` rows for -`wind-ons` and 920.3 GW of buildable resource, with the `cap` rows byte-identical to -the pre-pandas-3 `v20260916v4` run, and UPV output unchanged. Not yet committed to any -mainline branch. - -**Still recurring.** A sweep of every run under `runs/` finds four affected cases: the -three original `v20260924fix_st-AZNM_*` runs and `20260929_st-AZNM_baseline`, launched -2026-09-29 — after the bug was diagnosed but before any fix landed, and it lost its -wind the same way (0.7 GW in 2032). Every affected run needs re-running once a fix is -applied; their results are not usable. - -Worth noting for whoever fixes this: the reason it went unnoticed is that -`[c for c in alloutcap.columns if c.startswith("bin")]` drops non-matching columns -silently. Asserting that no `wsc*` columns survive the rename would turn this class of -failure loud rather than letting the pivot discard them. +**Files changed:** +- `reeds/input_processing/writesupplycurves.py` — in `agg_supplycurve()`, the + `if dfin.empty:` branch assigns `dfin['bin'] = pd.Series([], dtype='int64')` in place + of `dfin['bin'] = []`, with a comment recording the pandas-3 concat behaviour. No + other files touched. **Fixed upstream?** No — and, unlike most entries here, **the bug is live upstream right now, not latent.** `reeds/input_processing/writesupplycurves.py` has the identical bare `dfin['bin'] = []` at line 105 at tag `2026.08.03` and at line 121 on the current `upstream/main`, and upstream's `environment.yml` has pinned `pandas=3.0` since -2026-05-08 (`2ff493b5`, "update all python packages and use conda-forge for -everything") — four months before this fork moved to it. So upstream satisfies -conditions 1 and 2 already; they simply have not hit conditions 3-4, because their -default and test cases (`cendiv/Pacific`, `country/USA`) are coastal or national and -always have a populated offshore curve. Any upstream user running a landlocked region -with `GSw_OfsWind=1` on a current checkout gets silently zeroed wind. - -Not raised upstream as of 2026-09-29. Searched the `ReEDS-Model/ReEDS` issue tracker -(all 86 issues, open and closed) plus PR history for `writesupplycurves`, -`agg_supplycurve`, `rsc_combined`, `rscbin`/`wsc`, pandas/dtype, and -offshore/landlocked supply-curve terms. The nearest hits are unrelated: #27 ("Offshore -wind zones incompatible with custom regions") is about prescribed offshore builds under +2026-05-08 (`2ff493b5`, "update all python packages and use conda-forge for everything") +— four months before this fork moved to it. So upstream satisfies conditions 1 and 2 +already; they simply have not hit conditions 3-4, because their default and test cases +(`cendiv/Pacific`, `country/USA`) are coastal or national and always have a populated +offshore curve. Any upstream user running a landlocked region with `GSw_OfsWind=1` on a +current checkout gets silently zeroed wind. + +Not raised upstream as of 2026-09-25. Searched the `ReEDS-Model/ReEDS` issue tracker +(all 86 issues, open and closed) plus the PR history for `writesupplycurves`, +`agg_supplycurve`, `rsc_combined`, `rscbin`/`wsc`, pandas/dtype, and offshore/landlocked +supply-curve terms. The nearest hits are unrelated: #27 ("Offshore wind zones +incompatible with custom regions") is about prescribed offshore builds under `GSw_OffshoreZones=1`, and #28 is an offshore-zone TODO list. Strong candidate to report and contribute back — it is a one-line dtype correction that is a no-op under pandas 2. +**Worth knowing:** the reason this cost a week of runs is that +`[c for c in alloutcap.columns if c.startswith("bin")]` drops non-matching columns +silently. If this class of failure recurs, consider asserting that no `wsc*` columns +survive the rename rather than letting the pivot quietly discard them. + ## `startyear` must be old enough for historical hydro capacity factor data **Symptom:** confirmed live traceback, from `runs/v20260818_USA_optimized_mvp/gamslog.txt` diff --git a/CEPM/reeds-to-cepm-log.md b/CEPM/reeds-to-cepm-log.md index 822f30876..c7e8c1f94 100644 --- a/CEPM/reeds-to-cepm-log.md +++ b/CEPM/reeds-to-cepm-log.md @@ -35,6 +35,7 @@ Every upstream-owned path this fork has modified or added, as of the base above. | `reeds/core/setup/b_inputs.gms` | Modified | GAMS compatibility | | `reeds/input_processing/fuelcostprep.py` | Modified | Census divisions in fuelcostprep.py | | `reeds/input_processing/recf.py` | Modified | recf.py when offshore wind is disabled | +| `reeds/input_processing/writesupplycurves.py` | Modified | Empty supply curve drops onshore wind under pandas 3 | | `reeds/resource_adequacy/reeds2pras/README.md` | Modified | Minor and cosmetic | | `postprocessing/compare_cases.py` | Modified | Wrong module in compare_cases.py's "Flexibly Sited Demand" slide | | `reeds/report_utils.py` | Modified | parse_caselist TypeError with a prefix-glob caselist | @@ -168,6 +169,71 @@ input file. offshore wind via `techs_banned` instead of `GSw_OfsWind = 0` leaves the `eq_RPS_OFSWind` state mandate active and the model infeasible. +## Empty supply curve drops onshore wind under pandas 3 (writesupplycurves.py) + +### Description of issue: + +Runs completed cleanly but built **no new onshore wind**, because +`inputs_case/rsc_combined.csv` was written with only `cost_cap`/`cost_trans` rows for +`wind-ons` and no `cap`/`cost` rows — leaving the model with zero buildable wind +resource. No error, no warning; `writesupplycurves.py` logged `Finished` normally. +Caught 2026-09-24 on `runs/v20260924fix_st-AZNM_*` (0.7 GW of wind in 2032, all of it +pre-existing) against `runs/v20260916v4_st-AZNM_*` (14.2-15.8 GW) with identical +switches. + +`agg_supplycurve()` assigned `dfin['bin'] = []` for an empty input supply curve, which +pandas types as float64. `st-AZNM` is landlocked, so its offshore curve is header-only +while `GSw_OfsWind=1` still processes it. pandas 3.0 stopped excluding empty frames from +dtype resolution in `pd.concat`, so stacking the onshore and offshore frames upcast the +`bin` index level to float64, `"wsc" + bin.astype(str)` yielded `wsc1.0` instead of +`wsc1`, and the `{"wsc1": "bin1", ...}` rename matched nothing. The downstream pivot +selects value columns with `startswith("bin")` and silently dropped every wind row. + +The repo's pandas pin moved to `pandas==3.0.*` in `129ecdbc` (2026-09-10, Python 3.14 +bump); the local venv was rebuilt to 3.0.5 on 2026-09-24, 21 minutes before the first +affected run. This is a latent upstream bug the upgrade activated — not a reason to +un-pin pandas. + +### Files changed: + +- `reeds/input_processing/writesupplycurves.py` — in `agg_supplycurve()`, the + `if dfin.empty:` branch now assigns `pd.Series([], dtype='int64')` instead of `[]`, + matching the int64 `reeds.inputs.get_bin` produces in the non-empty branch, so the + `bin` level never upcasts through `pd.concat`. Fixed at the dtype source rather than + at the `.astype(str)` call: the same trap sits one line below on `class`, and + `agg_supplycurve()` also feeds the geothermal `pd.concat(geo, axis=0)`, which is one + switch change away from failing the same way. (UPV and CSP cannot hit it — they call + `agg_supplycurve()` standalone, with no populated sibling for an empty frame to + corrupt.) + +### Reference: [`known-reeds-issues.md`](known-reeds-issues.md) + +### What to test in new releases: + +- Does upstream still assign a bare `dfin['bin'] = []` in `agg_supplycurve()`'s empty + branch? It does at tag `2026.08.03` (line 105) and on `upstream/main` (line 121). + Note that upstream has pinned `pandas=3.0` since 2026-05-08 (`2ff493b5`) — four months + before this fork did — so **the bug is live upstream, not latent**. They have not hit + it only because their default and test cases (`cendiv/Pacific`, `country/USA`) are + coastal or national and always have a populated offshore curve. Not raised on their + issue tracker as of 2026-09-25 (all 86 issues searched). Worth reporting and + contributing back; it is a one-line dtype correction that is a no-op under pandas 2. + If they fix it themselves, drop this patch rather than merging it. +- Has upstream changed how `windall` is assembled, or how the `wsc{bin}` -> `bin{bin}` + rename is applied? Both sit within a few lines of the fix and either would change what + the patch needs to guarantee. +- After any pandas major-version bump, re-run one **landlocked** case with + `GSw_OfsWind=1` (an `st/AZ.NM`-style selection) and check that `rsc_combined.csv` has + all four `sc_cat` categories for `wind-ons` with equal row counts: + `awk -F, 'NR>1{split($1,a,"_"); print a[1]"|"$3}' inputs_case/rsc_combined.csv | sort | uniq -c | grep wind`. + A coastal case cannot surface this bug — its offshore curve is non-empty, so nothing + upcasts. +- More generally, this is the second empty-frame-into-`concat` bug in the input + processing chain (see *Resolving recf.py when offshore wind is disabled*). That + pattern is safe on `axis=1` (column-wise, dtypes stay per-column) and dangerous on + `axis=0` (the empty frame's dtypes participate). Check new upstream `pd.concat` calls + against that distinction. + ## Wrong module in compare_cases.py's "Flexibly Sited Demand" slide ### Description of issue: diff --git a/reeds/input_processing/writesupplycurves.py b/reeds/input_processing/writesupplycurves.py index a16838e4d..873919324 100644 --- a/reeds/input_processing/writesupplycurves.py +++ b/reeds/input_processing/writesupplycurves.py @@ -102,7 +102,11 @@ def agg_supplycurve( ### Assign bins if dfin.empty: - dfin['bin'] = [] + ### Type the empty 'bin' column to match reeds.inputs.get_bin's int64 output. + ### A bare [] gives float64, which pandas >=3.0 propagates through pd.concat + ### (empty frames now participate in dtype resolution), silently turning + ### downstream 'wsc{bin}' labels into 'wsc1.0' and dropping the supply curve. + dfin['bin'] = pd.Series([], dtype='int64') else: dfin = ( dfin