Skip to content

Make the daily scraper survive refused connections from govjobs.publi… - #13

Merged
masterries merged 1 commit into
mainfrom
claude/github-actions-scraping-failures-709d30
Sep 28, 2026
Merged

masterries merged 1 commit into
mainfrom
claude/github-actions-scraping-failures-709d30

Conversation

@masterries

Copy link
Copy Markdown
Owner

…c.lu

Recent runs died with an uncaught "[Errno 111] Connection refused" in the middle of the listing crawl, so nothing was saved, committed or notified. The site refuses TCP connections when a runner IP opens too many of them, and every bare requests.get() opened a new TCP+TLS connection.

job_scraper.py:

  • one keep-alive requests.Session for all requests (the live test ran 41 requests over a single connection), timeouts, urllib3 retries with backoff for refused connections and 429/5xx (Retry-After capped at 60s)
  • crawl the listing first, then the missing detail pages; on errors stop and save what was collected (::warning::), exit 1 only if nothing was scraped; missing details are fetched on the next run
  • circuit breaker after 3 consecutive failed detail pages, max 100 detail fetches and a 20 minute time budget per run
  • check listing status codes, dedupe links, max page guard, compare the collected count with the count shown on the site

telegram.py:

  • HTML parse mode, split messages between job blocks instead of every 4000 characters (cut links made Telegram reject the chunk)
  • timeout, retries, exit 1 when a message can't be sent

Workflow:

  • checkout@v7 / setup-python@v7 (Node 24), Python 3.14, pip cache, install from requirements.txt (drop unused firebase-admin)
  • cron at 05:17 and 15:47 UTC instead of 00:00, concurrency group, timeout-minutes, contents: write, push with rebase retry
  • notify and commit only on main, notify before committing so a failed notification is resent by the next run

…c.lu

Recent runs died with an uncaught "[Errno 111] Connection refused" in the
middle of the listing crawl, so nothing was saved, committed or notified.
The site refuses TCP connections when a runner IP opens too many of them,
and every bare requests.get() opened a new TCP+TLS connection.

job_scraper.py:
- one keep-alive requests.Session for all requests (the live test ran 41
  requests over a single connection), timeouts, urllib3 retries with
  backoff for refused connections and 429/5xx (Retry-After capped at 60s)
- crawl the listing first, then the missing detail pages; on errors stop
  and save what was collected (::warning::), exit 1 only if nothing was
  scraped; missing details are fetched on the next run
- circuit breaker after 3 consecutive failed detail pages, max 100 detail
  fetches and a 20 minute time budget per run
- check listing status codes, dedupe links, max page guard, compare the
  collected count with the count shown on the site

telegram.py:
- HTML parse mode, split messages between job blocks instead of every 4000
  characters (cut links made Telegram reject the chunk)
- timeout, retries, exit 1 when a message can't be sent

Workflow:
- checkout@v7 / setup-python@v7 (Node 24), Python 3.14, pip cache,
  install from requirements.txt (drop unused firebase-admin)
- cron at 05:17 and 15:47 UTC instead of 00:00, concurrency group,
  timeout-minutes, contents: write, push with rebase retry
- notify and commit only on main, notify before committing so a failed
  notification is resent by the next run

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@masterries
masterries merged commit d2dc63e into main Sep 28, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant