Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
15 commits
Select commit Hold shift + click to select a range
09b6f2a
gateway: parse the shared action knobs (delay, percent, presence, tak…
erfnzdeh Oct 3, 2026
ba53542
gateway: a pacer per action kind, with a spacing floor, upward jitter…
erfnzdeh Oct 3, 2026
7c3ed2f
filters: await coroutine filters in the job engine, and add sender_is…
erfnzdeh Oct 3, 2026
cae4151
gateway: a pending-action store, a presence manager and an in-process…
erfnzdeh Oct 3, 2026
f75ee35
actions: plan on the bus lane and execute through the op layer on a p…
erfnzdeh Oct 3, 2026
dc09da3
daemon: one action scheduler per account, saved at shutdown and resum…
erfnzdeh Oct 3, 2026
52efe1b
gateway: test the approved jobs.yaml syntax, the pacing block and eac…
erfnzdeh Oct 3, 2026
97c61ea
job: queue list and cancel, per-action counters in job list and get, …
erfnzdeh Oct 3, 2026
6d149ad
config validate: report job problems with the daemon's own jobs.yaml …
erfnzdeh Oct 3, 2026
9ef6b8d
docs: job actions, knobs, pacing, presence, takeover and the queue; a…
erfnzdeh Oct 3, 2026
7b0f821
gateway: keep retrying a forward or read through a long reconnect, in…
erfnzdeh Oct 3, 2026
a6a24c0
gateway: read the saved queue before any job submits, so an early rep…
erfnzdeh Oct 3, 2026
1c902e4
tests: job disable and job remove drop the job's pending actions
erfnzdeh Oct 3, 2026
ff6ebca
gateway: keep an action worker alive through an unexpected error inst…
erfnzdeh Oct 3, 2026
09180ce
gateway: survive a request Telethon cancels on disconnect, scope take…
erfnzdeh Oct 3, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
54 changes: 54 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,39 @@ codes documented in `AGENT.md` are the public API.

### Added

- **Job actions `react`, `read` and `view`.** `react` puts an emoji (one,
a random pick from a list, or a weighted pick) on the message, replacing
the account's previous reaction; it reads the chat up to that message
first, so nobody sees a reaction on a message still shown unread, and an
album gets one reaction. `read` marks the chat read up to the triggering
message with the right RPC for the peer (forum topics included), coalesces
per chat and waits a reading time on top of its delay; `mentions` and
`reactions` clear those badges. `view` counts a channel view, or marks a
voice or round note listened in a DM or group, and never consumes
view-once media unless told to.
- **Knobs on every action, or on the job as a default:** `delay` (a random
delay from when the event arrived, which never blocks the bus), `percent`,
`presence` (`leave`, `blip`, `session`, with optional `quiet_hours`),
`on_takeover` (what to drop when you act in the chat yourself from another
device) and `dry_run`.
- **`reply` shows "typing..." first** (about 40 characters a second, 2-15 s),
on by default, `typing: false` to turn it off. Long form
`reply: {text, typing, delay, filters, processors}`.
- **Pacing per account and action kind**, configurable in a top-level
`pacing:` block of `jobs.yaml`, with cautious defaults (a reaction every
4 s and at most 300 an hour; reads and views every 2 s; forwards and
replies every 1.5 s). FLOOD_WAIT reschedules and slows the queue,
transient errors retry with backoff, permanent ones are counted.
- **Pending actions survive restarts and crashes** in
`accounts/<alias>/pending.json` and resume at their original due time.
React, view and reply expire 24 h after the event (configurable); read and
forward never do. A message replayed after a crash is not acted on twice.
- **Filters `sender_is_contact` and `chat_is_new`.**
- **`tlgr job queue`** lists pending actions with their due time, and
`tlgr job queue cancel` drops them by id, `--chat`, `--job` or `--all`.
`job list`, `job get` and `daemon status` report per-action counters:
done, skipped, superseded, expired, pending and errors.

- **`chat poster list` can walk deeper than one call.** A single call stops
at `--max-messages 20000` and at the operation deadline, so a longer
history could not be harvested at all: a bigger number was refused and
Expand All @@ -29,6 +62,27 @@ codes documented in `AGENT.md` are the public API.
messages were paced, so a chain of small calls would have read history
unthrottled.

### Changed

- **Every job action now runs through the op layer and is paced.** `forward`
and `reply` used to call Telethon directly from the bus handler; they now
run as `message.forward`, `message.send` and `media.upload` operations, so
a job obeys the policy allow/deny list, the rate limiter and the flood
budget like the CLI does. Their meaning is unchanged: a native forward
(with `drop_author`), or with processors a re-send of the processed text
or media caption; a reply to the matched message with processors applied.
Forwards on one account are now spaced at least 1.5 s apart, a
link-preview post is re-sent as text instead of failing, and an album gets
one reply instead of one per photo.
- **`jobs.yaml` is validated strictly.** An unknown key, a bad duration, a
percent outside 0-100 or an unknown presence mode is reported with the
job's name and the action's position, by `job add`, `job reload
--validate-only` and `config validate`. A broken job is skipped at load
while the others run, and a job whose edit broke it keeps running on
`job reload`.
- **`job get` runs in the daemon** so it can report live counters, and
`job disable` also drops what the job had queued.

### Fixed

- **Webhook filters work.** Both halves of `[webhook.filters]` were dead.
Expand Down
207 changes: 207 additions & 0 deletions docs/design/JOB_ACTIONS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,207 @@
# Job actions: one execution model

Status: implemented (feat/job-actions). Reference docs: `tlgr/gateway/README.md`
and `tlgr/actions/README.md`. This note records the choices the spec left
open, and why.

## Shape

A gateway job no longer runs actions in its bus handler. The handler
evaluates filters, rolls `percent`, calls the action's `plan()` (processors,
emoji pick, nothing awaited) and queues `PendingItem`s on the account's
`ActionScheduler` (`gateway/scheduler.py`). The scheduler owns delays,
pacing, quiet hours, expiry, takeover, retries and persistence, and calls the
action's `execute()`, which runs operations through `dispatch.execute` via
`DaemonOpRunner` (`gateway/executor.py`). One scheduler per account, shared by
every job on it; one worker and one `Pacer` per action kind.

Every built-in action, `forward` and `reply` included, goes through this
path. A plain function registered with `@register_action` still runs at once,
outside the scheduler, so third-party actions written against the old
interface keep working.

## Decisions

### Pacing and jitter

* `every` is a floor. The gap is drawn from `[every, 1.5 x every]`, a mean of
1.25x and about +-20% around it. A window centred on `every` (the spec's
"about +-25%") would put reactions 3 s apart under a 4 s rule; a floor
keeps the promise "never closer than `every`".
* `per_hour` is a rolling hour of consumed slots. An absent `per_hour` in
`pacing:` keeps the default cap (300 for react), `per_hour: null` lifts it.
* Pacer state (the rolling hour, the slow-down) is not persisted; a restart
starts the hour afresh. The server-side flood memory (`flood.json`) still
survives, which is the part that matters.
* Due times are wall-clock, sleeps are monotonic, so a worker never sleeps
more than 60 s before looking at the clock again; a delay that spans a Mac
sleep fires at most a minute late.
* A request cancelled by Telethon (it cancels everything in flight when the
client disconnects) is a transient failure of that action, never the end of
the worker; a worker that ends anyway is restarted.
* Dry-run items skip the pacer and presence: they never touch Telegram, so
they should not slow the real actions on the same account.

### Failures

Classification goes through `core.errors.classify`, the same table the CLI
uses.

* `RATE_LIMITED` (FLOOD_WAIT, slow mode): a job op sleeps a wait of up to
10 s inside the request (`flood_wait_max=10`); a longer one comes back and
the item is rescheduled at `now + wait + 1..5 s`. The kind's pacer is held
past the wait and its spacing doubles (up to 8x), recovering after ten
minutes without another flood. An item gives up after 10 floods (an error).
Floods are counted apart from transient failures, so one does not use up
the other's retries.
* Retryable (`RETRYABLE`: network, timeouts, server errors, disconnected):
three retries after about 5 s, 30 s and 2 min (each x1-1.5). An action
that never expires (forward, read) then keeps retrying every ten minutes,
up to 12 attempts in all (about 1.5 h), so a relay rides out an account
that is reconnecting instead of dropping posts.
* Everything else is permanent and counted under `errors` with `last_error`
(`USAGE: REACTION_INVALID: ...`): REACTION_INVALID, MESSAGE_ID_INVALID,
CHAT_WRITE_FORBIDDEN, policy denials, PEER_FLOOD and frozen accounts. The
last two are deliberately not retried: the account has been told to stop.

### Persistence

* A JSON file per account, `accounts/<alias>/pending.json`, written with
`write_private` (temp file, chmod 0600, fsync, rename). Chosen over SQLite
because the queue is small, it matches `flood.json` beside it, and it can be
read with `jq` when diagnosing. Writes are debounced (1 s) and forced at
shutdown.
* Bounded at 10 000 items per account; past that a new item is refused and
counted as an error, logged once.
* Items carry everything needed to run without the Telethon event: chat,
message, peer kind, forum topic, `grouped_id`, presence, quiet hours,
takeover mode, dry run, and the payload (emoji, processed text, forward
destination, view mode).
* The file also keeps the keys of the last 5 000 finished items. A key is
`job|index|event|chat|msg|extra` (the edit time for an edit event, the
destination for a forward), so an update replayed by catch-up after a crash
(the update state is saved once a minute) is not acted on twice. This also
protects forwards, which could duplicate before.
* Resume happens once per scheduler, when its account's jobs are created at
boot. Items whose job no longer exists (or is disabled) are dropped and
logged. An item that was running when the daemon died runs again
(at-least-once); a clean shutdown stops starting new actions, lets a
running one finish within the drain, then saves.
* `job remove` drops the job's pending items and counters; `job disable`
drops its pending items (counted as superseded). A job updated by
`job reload` keeps its queue.

### Expiry

React, view and reply expire 24 h after the event arrived; read and forward
never. `pacing.<alias>.expire` overrides per kind with a duration, or
`never`/`null`/`0`.

### react

* Custom emoji are supported as `custom:<document id>`, the spelling
`reaction add` already uses.
* Album target: the message carrying the caption, else the first one. That is
where Telegram Desktop attaches an album's reactions
(`HistoryView::GroupedMedia::itemForText()`: the caption item, else the
first part) and what Telegram for Android uses
(`MessageObject.GroupedMessages.findPrimaryMessageObject()`). If the
caption message arrives after the first one while the reaction is still
pending, the pending reaction is moved to it.
* An album is one `percent` roll, not one per photo. Album siblings count as
`skipped`.
* React implies read through `ensure_read`: paced on the read queue, read up
to the reacted message, and any pending read for that chat at or below it
finishes as done (coalesced). If the chat is already known read that far
(by tlgr or by you elsewhere) no read is sent. If the read fails the
reaction does not go out.

### reply

* One reply per album as well (to the caption message), the same rule as
react. Replying once per photo was never useful in a DM, which is the
main use case.
* Typing is a two-step item: `chat.typing` is started in the background and
the item is re-queued for when typing ends, then `message.send` goes out.
The reply queue is never blocked for the typing time.
* Text is sent as markdown, which is what `event.reply()` did with the
client's default parse mode.

### forward

* One pending item per destination, so one refusing destination does not
stop the others and each is retried or counted on its own.
* With processors: a photo or document is re-sent with `media.upload
--from-message` and the processed caption; a link-preview post is re-sent
as text (Telegram regenerates the preview), where the old code failed on
it; other media with text is re-sent as text, without text it is skipped.
Text is taken as `message.text` (markdown under the client's parse mode)
and sent with `--parse md`, so formatting survives as it did.
* Albums are still forwarded message by message, as before; grouping them
would need a hold window and changes what the destination sees.

### read

Reading time is 250 words a minute plus 3 s for media, capped at 60 s, added
to the `delay` draw.

### view

A broadcast channel post (the `post` flag, or a broadcast chat entity) is a
view. In private chats and groups only voice and round notes are something
to view; view-once media needs `include_view_once: true`.

### Presence

* The rule for two jobs with different modes on one account: the most online
request wins while its actions run. `leave` never sends anything and never
ends an online stretch another job started; offline is sent only once no
`blip`/`session` action is running and, for `session`, none is due within
5 s. Linger after the last action: 5 s.
* `session` goes online when its first action executes, not ahead of it.
* If the account-level `[presence] mode` is not `off`, the daemon already
manages the status and job presence is silent.
* Quiet hours use `[defaults] timezone` (an IANA name) when set, else the
daemon's local time. Held items are released at the end of the window,
spread over five minutes, then go through the pacer.

### Manual takeover

* A read is tlgr's own when its `max_id` is at or below the highest id tlgr
asked to read in that chat; the mark is set before the RPC, because the
echo can arrive before the answer.
* An outgoing message is tlgr's own when its id is one tlgr recorded sending,
or when tlgr sent in that chat less than 10 s earlier.
* A read that arrives within 10 s of tlgr sending in that chat is also tlgr's
own: sending marks the chat read on the server.
* Takeover drops items in that chat with a message id up to the read or sent
id, by each item's own `on_takeover`; forwards are never dropped. A read in
a forum topic only drops items of that topic (forum ids span the chat).
* `job queue cancel` counts what it drops as `superseded` too.

### Filters

* `sender_is_contact` trusts the `contact` flag on the sender entity the
update carried; without one it fetches the sender and caches the answer for
10 minutes.
* `chat_is_new` probes the history once per chat (`get_messages(limit=1,
offset_id=msg)`) and caches the first message id (or "not new"), in an LRU
of 5 000 chats per process.
* Both are coroutines, so filter evaluation gained `evaluate_async`; the
synchronous `evaluate` (webhook filters) rejects them with a reason.

### Validation and the CLI

* Validation is strict per job: unknown keys, bad durations, percent outside
0-100, unknown presence or takeover modes, unknown actions, processors on
an action that cannot use them. At load a broken job is skipped and logged
while the others run; on `job reload` a job whose edit broke it keeps
running in its last good form.
* `tlgr job queue` lists and `tlgr job queue cancel` cancels. The registry
forbids an op id that is also a group, so the list op is `job.queue.list`
tagged `group-default`, and its CLI group runs it when called bare. Options
go on the leaf (`tlgr job queue list --job dm-ack`).
* `job.queue.cancel` is destructive as a whole, so every variant needs
`--yes` off a terminal, like `job remove`.
* `job.get` moved from the local surface to the daemon so it can report live
counters.
6 changes: 3 additions & 3 deletions docs/reference/PARITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@ Coverage against the Telegram feature catalog, computed from the registry: every
`covered` is implemented today. `acct%` is covered **plus** waived — an id this build genuinely cannot cover, named in `tlgr/data/parity_waivers.toml` with the reason and the MTProto method that is missing. Ids whose feasibility is `not-applicable` or `prohibited` are excluded from the denominator once and never counted again.

```
catalog 2026-09-02 — 678 operations, 951 invocable paths
catalog 2026-09-02 — 680 operations, 953 invocable paths

domain covered req % acct% ops
auth_sessions_security 89 89 100.0% 100.0% 45
Expand All @@ -21,7 +21,7 @@ messages_core 167 167 100.0% 100.0% 61
polls_reactions_content 173 174 99.4% 100.0% 99
profile_settings_privacy 178 178 100.0% 100.0% 109
stories 120 120 100.0% 100.0% 48
updates_sync_network 189 189 100.0% 100.0% 68
updates_sync_network 189 189 100.0% 100.0% 70

priority covered req % acct%
P0 178 178 100.0% 100.0%
Expand Down Expand Up @@ -49,7 +49,7 @@ uncovered: 9 (9 waived with a reason)
| `polls_reactions_content` | 173 | 174 | 99.4% | 100.0% | 99 |
| `profile_settings_privacy` | 178 | 178 | 100.0% | 100.0% | 109 |
| `stories` | 120 | 120 | 100.0% | 100.0% | 48 |
| `updates_sync_network` | 189 | 189 | 100.0% | 100.0% | 68 |
| `updates_sync_network` | 189 | 189 | 100.0% | 100.0% | 70 |

## By priority

Expand Down
4 changes: 2 additions & 2 deletions docs/reference/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

# Command reference

678 operations across 47 groups, generated from the operation registry. Groups still served by v1's hand-written commands are not listed here; they arrive with their own PR.
680 operations across 47 groups, generated from the operation registry. Groups still served by v1's hand-written commands are not listed here; they arrive with their own PR.

| Group | Operations | Reference |
|---|---:|---|
Expand All @@ -27,7 +27,7 @@
| `gift` | 19 | [gift.md](gift.md) |
| `giveaway` | 6 | [giveaway.md](giveaway.md) |
| `inline` | 7 | [inline.md](inline.md) |
| `job` | 8 | [job.md](job.md) |
| `job` | 10 | [job.md](job.md) |
| `location` | 9 | [location.md](location.md) |
| `media` | 28 | [media.md](media.md) |
| `message` | 39 | [message.md](message.md) |
Expand Down
Loading
Loading