Skip to content

Fix gateway jobs: deliver incoming updates, build Telethon events, resolve @chat refs, start jobs at boot - #43

Merged
erfnzdeh merged 6 commits into
mainfrom
fix/daemon-jobs
Oct 3, 2026
Merged

erfnzdeh merged 6 commits into
mainfrom
fix/daemon-jobs

Conversation

@erfnzdeh

@erfnzdeh erfnzdeh commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

Gateway jobs have never matched a message under the v2 daemon. On the live install every job stopped with matched=0 skipped=0, and after a restart on 2026-09-30 no jobs ran at all. Four faults stacked up, and each one alone was enough to stop every job.

What was wrong

  1. No incoming update reached the event bus. SessionManager.ensure() called the ready hook straight after AccountSession.start(), which only schedules the supervisor. session.client was still None, so _attach_handlers returned early. Jobs, webhooks and watch saw the daemon's own sends only. Evidence: the live Neo account's events.state is {"seq": 1}.
  2. Jobs got the wrong event shape. The bus passes the raw TL update (UpdateNewChannelMessage), but every filter and action reads a Telethon high-level event (event.chat_id, is_private, reply()).
  3. chat_id: "@name" never matched. filter_chat_id skips string refs, expecting the gateway to resolve them, and nothing did, in v1 either.
  4. Jobs did not start at boot. reload_jobs() ran only on tlgr job reload.

Also fixed along the way:

  • apply_difference is read-only on every account's first connect. MessageBox has __slots__, so the too-long hook failed, cost one reconnect, and was then skipped.
  • job list showed - for enabled and running because JobState omitted fields at their defaults.

Changes

  • daemon/session.py, daemon/sessions.py, daemon/app.py: the session calls an on_client hook from the supervisor once it builds the client. That happens once per session, since the client survives reconnects.
  • core/telethon_compat.py: the too-long hook moves the box onto a per-client subclass with no extra slots, instead of assigning to the instance.
  • gateway/engine.py: the bus path builds the event with the same builder a Telethon handler uses and applies its filter, which restores NewMessage(incoming=True) so auto-replies skip the account's own messages. setup() resolves @chat refs, and a ref that failed (account offline) is retried at most once a minute.
  • daemon/app.py: run() starts the jobs in a background task after accounts connect, and shutdown cancels it.
  • models/daemon.py: JobState uses omit_defaults=False, like DaemonStatus.

Tests

New tests use real Telethon TL types and builders. The fake client gains _mb_entity_cache, _self_id and parse_mode, which Telethon events read. Each new test file fails against the old code.

  • tests/test_daemon_updates.py: an incoming update is numbered on the bus, and the handler is attached exactly once across reconnects.
  • tests/test_telethon_compat.py: the hook installs on a real MessageBox, reports a gap, and does not leak to another client.
  • tests/test_gateway_bus.py: forward, processed copy, account isolation, private auto-reply, no reply to own messages, ref resolution and retry.
  • tests/test_daemon_jobs.py: enabled jobs run after boot, and a broken jobs.yaml does not stop the daemon.

make lint typecheck test docs parity: all pass, 13324 tests.

Not in this PR

WebhookPusher.should_push evaluates webhook filters against the raw TL update too, the same shape bug as item 2. That is left for a follow-up.

Deploying

After merging, the live daemon needs a restart (tlgr daemon restart) to pick this up. The jobs then start on their own.

…ng updates reach the bus

SessionManager.ensure() called the ready hook straight after
AccountSession.start(), which only schedules the supervisor. The client did
not exist yet, so _attach_handlers returned early and no incoming update ever
reached the bus: jobs, webhooks and watch saw the daemon's own sends only, and
every gateway job stopped with matched=0 skipped=0.

The session now calls an on_client hook from the supervisor right after it
builds the client, before the private-API hooks that can raise. The client is
kept across reconnects, so the handler is attached once.
MessageBox declares __slots__, so assigning apply_difference on the instance
raised AttributeError on every account's first connect. The account logged
'degraded', burned one reconnect, and then ran without the hook, because the
retry reuses the client and skips _install_hooks.

The box now moves onto a subclass built for that client. It adds no slots,
so __class__ assignment is allowed, and other clients keep the plain class.
The new tests use Telethon's real MessageBox; the fake client has none, which
is why this was never caught.
…s, so jobs can match

Two faults hid behind the dispatch bug fixed earlier:

- The bus hands a job the raw TL update (UpdateNewChannelMessage), but every
  filter and action reads a high-level event: event.chat_id, is_private,
  reply(). The job now builds the event with the builder a Telethon handler
  would have used, attaches entities and client the way _dispatch_update
  does, and applies the builder's filter. That also restores
  NewMessage(incoming=True), so an auto-reply job no longer answers the
  account's own messages.
- filter_chat_id skips '@name' refs on the promise that the gateway resolves
  them, and nothing did, so 'chat_id: "@channel"' never matched, in v1
  either. setup() now resolves them through the job client, and logs a ref it
  cannot resolve.

The fake client gains _mb_entity_cache, _self_id and parse_mode, which
Telethon's events read, so the new tests drive real builders and TL updates.
…t failed while offline

reload_jobs() ran only on 'tlgr job reload', so every daemon restart left the
gateway with no jobs until someone noticed. On the live install that was from
2026-09-30 to 2026-10-03, across seven restarts.

run() now starts the jobs in a background task once accounts are connected,
so readiness does not wait for a job's account, and shutdown cancels it.
Because a job can now start while its account is still offline, a '@chat'
ref that fails to resolve is retried from the event path, at most once a
minute, instead of leaving the job unable to match for good.
…ing '-' for both

JobState inherited omit_defaults, so an enabled job dropped 'enabled' (True
is the default) and a stopped one dropped 'running' (False is the default).
The table renders a missing field as '-', so an enabled, stopped job looked
exactly like a disabled one. Same fix and reasoning as DaemonStatus.
@erfnzdeh
erfnzdeh merged commit f4dd540 into main Oct 3, 2026
13 checks passed
@erfnzdeh
erfnzdeh deleted the fix/daemon-jobs branch October 3, 2026 16:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant