Skip to content

[Bugfix] Make CI robust against transient GMT download failures - #213

Merged
boriskaus merged 6 commits into
mainfrom
bk/fix-ci-gmt-download-retries
Sep 19, 2026
Merged

boriskaus merged 6 commits into
mainfrom
bk/fix-ci-gmt-download-retries

Conversation

@boriskaus

@boriskaus boriskaus commented Sep 18, 2026 •

Copy link
Copy Markdown
Member

The recurring CI failures (test_WaterFlow, test_GMT and the LaPalma tutorial) all raised the same error: "Could not download GMT topography data", triggered by GMT reporting

gmtread [ERROR]: Remote download is currently deactivated
gmtread [ERROR]: Unable to obtain remote file @earth_relief_03s

When the GMT data server is briefly unreachable, GMT sets auto_download = GMT_NO_DOWNLOAD and reports it with that misleading wording. The state clears once the server responds again, so retrying in-process does recover - but the existing backoff (5 s flat, ~20 s total) was too short to outlast the outage, so all attempts failed.

This is not a version issue: the passing and failing jobs ran identical GMT (v1.43.2 / GMT_jll 6.7.2) and the same Julia versions appear on both sides.

  • ext/GMT_utils.jl: use progressive backoff (~100 s total), fix the attempt counter (warnings reported "attempt 0/5") and report the file and underlying error when giving up.
  • src/IO.jl: same off-by-one in download_data, plus a truncated download could return a non-empty path and break out of the retry loop, showing up later as a confusing load error.
  • test/test_paraview_collection.jl: replace exact filesize assertions with structural checks. The .pvd embeds absolute paths and OS-dependent separators, so the byte count varies per machine (verified: 435 bytes from a longer path, not the asserted 317).
  • CI.yml: test Julia 1.13 explicitly and track 'pre'. Also wire up continue-on-error, as allow_failure was set on every matrix entry but never referenced, so the allow-failure job still failed the run.

Whats the purpose of this PR?

  • Bug fix
  • New feature
  • Documentation update
  • Other, please explain

Describe it in more detail below:

Checklist

  • The PR title is descriptive and starts with the appropriate tag: [BUGFIX], [ADDITION], [DOC], etc.
  • New tests (either assessing the correct behaviour of new internal functions or the correctness of a tutorial) were added, or old tests were updated
  • Affected tutorials have also been updated
  • The new feature was added in a way that does not break public API
  • New documentation related to the new feature was added
  • The new code follows the contributor guidelines, in particular the Runic Style

boriskaus and others added 4 commits September 18, 2026 22:40
The recurring CI failures (test_WaterFlow, test_GMT and the LaPalma
tutorial) all raised the same error: "Could not download GMT topography
data", triggered by GMT reporting

    gmtread [ERROR]: Remote download is currently deactivated
    gmtread [ERROR]: Unable to obtain remote file @earth_relief_03s

When the GMT data server is briefly unreachable, GMT sets
auto_download = GMT_NO_DOWNLOAD and reports it with that misleading
wording. The state clears once the server responds again, so retrying
in-process does recover - but the existing backoff (5 s flat, ~20 s
total) was too short to outlast the outage, so all attempts failed.

This is not a version issue: the passing and failing jobs ran identical
GMT (v1.43.2 / GMT_jll 6.7.2) and the same Julia versions appear on both
sides.

* ext/GMT_utils.jl: use progressive backoff (~100 s total), fix the
  attempt counter (warnings reported "attempt 0/5") and report the file
  and underlying error when giving up.
* src/IO.jl: same off-by-one in download_data, plus a truncated download
  could return a non-empty path and break out of the retry loop, showing
  up later as a confusing load error.
* test/test_paraview_collection.jl: replace exact filesize assertions
  with structural checks. The .pvd embeds absolute paths and
  OS-dependent separators, so the byte count varies per machine
  (verified: 435 bytes from a longer path, not the asserted 317).
* CI.yml: test Julia 1.13 explicitly and track 'pre'. Also wire up
  continue-on-error, as allow_failure was set on every matrix entry but
  never referenced, so the allow-failure job still failed the run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The CI failures in test_WaterFlow, test_GMT and the LaPalma tutorial all
came from GMT being unable to reach its data server:

    gmtread [ERROR]: Remote download is currently deactivated
    gmtread [ERROR]: Unable to obtain remote file @earth_relief_03s

Retrying the same server did not help, since the server was unreachable
for the whole retry window. Instead, cycle through the official mirrors
(https://www.generic-mapping-tools.org/mirrors) so that a single server
being offline no longer fails the download. The download is still
genuinely exercised - only *which* mirror serves it changes - so the
tests keep checking that downloading works.

Verified against a cold GMT cache: with two unreachable mirrors in front,
the download correctly falls through to a working one (GMT reports
"Remote data courtesy of GMT data server oceania"), and when every mirror
is unreachable import_topo still errors, so a real regression is not
masked.

Note that GMT reads GMT_DATA_SERVER when the session starts, so setting
the environment variable at runtime has no effect; `gmtset` is used
instead. All six configured mirrors were verified to serve the data
(singapore is omitted as it was unreachable).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The previous version called `gmtset(GMT_DATA_SERVER=...)` before *every*
download attempt, including the first. That was wrong in two ways:

`gmtset` changes GMT's global configuration by writing a gmt.conf in the
working directory. The test suite runs with ParallelTestRunner, whose
workers share that directory, so one test switching mirrors also changed
which server the other tests downloaded from. On macOS-aarch64 this made
test_WaterFlow build its grid from a different mirror than intended and
the values no longer matched the references:

    maximum(Topo_water.fields.area) ≈ 9.309204547276944e8
    Evaluated: 2.6336469421851286e6 ≈ 9.309204547276944e8

The failing jobs are exactly those that logged mirror-fallback warnings
(3 warnings on the failing macOS job, 0 on the passing ubuntu one).

Switching the server unconditionally also meant the configured default
was never used, so the mirrors were exercised even when nothing was wrong.

Now the first attempt uses whatever server GMT is already configured with
and touches no global state; mirrors are only tried after a failure, and
the original gmt.conf (or its absence) is restored afterwards. Verified
that the happy path leaves no gmt.conf behind, and that with an
unreachable default the download still falls back to a working mirror
("Remote data courtesy of GMT data server oceania") and cleans up after
itself.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
test_Gmsh intermittently failed on Windows with

    [ Error : Gmsh has not been initialized

`GmshDiscreteModel` initializes gmsh itself, but gmsh keeps global state
that is occasionally not in place yet on a freshly started worker process
under the parallel test runner. Initializing explicitly first makes that
state well-defined.

Note we deliberately do not finalize afterwards: GmshDiscreteModel already
finalizes the session it used, and a second finalize reproduces the very
same "Gmsh has not been initialized" error (confirmed locally - an earlier
version of this fix that finalized emitted exactly that message twice).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@boriskaus boriskaus changed the title Make CI robust against transient GMT download failures [Bugfix] Make CI robust against transient GMT download failures Sep 19, 2026
boriskaus and others added 2 commits September 19, 2026 14:49
…em in CI

The remaining CI failures were all GMT tile downloads timing out:

    grdblend [ERROR]: Libcurl Error: Timeout was reached
    grdblend [ERROR]: Unable to obtain remote file @N30E000.earth_relief_01m_g.nc

This is not the server blocking us (there is not a single 403/429 in the
logs; the server is contacted fine and the transfer then stalls). It is
contention: test_GMT, test_WaterFlow and the LaPalma tutorial run
concurrently under ParallelTestRunner and all pull tiles from the GMT
data server at the same time, each with GMT's own retries on top.

Download the tiles serially, once, before the workers are spawned
(test/topo_prefetch.jl, included from runtests.jl). The workers then
find them in GMT's cache (~/.gmt/server) and never touch the network.
Verified locally: after the prefetch, test_GMT and test_WaterFlow pass
with the data server *and* all mirrors unreachable.

On CI, ~/.gmt is additionally cached across runs (keyed on the hash of
topo_prefetch.jl, so changing the tile list invalidates it). The
scheduled run deliberately bypasses that cache, so the download itself
is still exercised regularly rather than only served from cache.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@boriskaus
boriskaus merged commit 768b20d into main Sep 19, 2026
20 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant