Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [Unreleased]

### Fixed
- **Transient connection failures no longer silently drop replies**: Telegram progress, final-response, and global error messages now retry transport failures with bounded exponential backoff while invalid formatting still falls back immediately to plain text. Explicit proxies are applied to the separate `getUpdates` client as well as normal Bot API requests. Claude SDK streams must now end with a `ResultMessage`; an empty stream is retried, while a partially observed stream is not replayed because its tool calls may already have caused side effects (#241)
- **OpenRouter empty-result fallback**: Anthropic-compatible providers that return an empty `ResultMessage.result` now fall back to the text in `AssistantMessage` instead of producing "No content to display" (#171)
- **A run that stops early no longer reports success**: Claude produces no final text when the CLI kills a run at the turn limit, so the bot fell through to its "✅ Task completed" placeholder and told the user the work was done. A run that died at turn 10 mid-task and a run that finished were indistinguishable in Telegram (#172). `ResultMessage.subtype` is now read alongside the cost and session id, the placeholder claims completion only for `success`, and the reply carries a footer saying why the run ended — turn limit, cost budget, cancellation, or an unrecognised reason named by its raw subtype — with the turn count and a prompt to send another message to continue. Every place the bot renders a Claude reply carries it — typed messages, document, photo and voice input in agentic mode, the same four in classic mode, `/continue`, the Continue Session button and the quick-action buttons, whose `✅ … Complete` heading now reads `⚠️ … Stopped` when the run was cut short. Webhook-triggered and scheduled runs carry it too: nobody is watching those, so the reason a nightly job came back short is the whole of what its notification can say about it (#230)
- **`ClaudeResponse.num_turns` is the turn count the CLI reports**: it was derived by counting `UserMessage` and `AssistantMessage` objects, which over-reports — every tool result arrives as another user message — so a run stopped at turn 10 could be recorded as having taken roughly twice that. It only reached the logs and the session store before; the stop-reason footer now shows it to the user, which made the approximation worth removing. `ResultMessage.num_turns` is used where the CLI supplies it, with the message count kept as the fallback for a result that carries none
- **Inline code spans accept backtick runs of any length**: `markdown_to_telegram_html` matched a code span only between single backticks with no backtick inside, so ``` ``a`b`` ``` rendered as the span `a` followed by loose text. A run of N backticks now opens a span that closes on a run of N, and one space is stripped from each end when both are present, as CommonMark specifies. Pairing the runs is done by scanning rather than by a backreferencing regex: `` (`+)([^\n]*?)\1 `` re-scans the rest of the line for every opener that never closes, which is superlinear on a reply whose backtick runs are all of different lengths — 2.5 seconds on a 256KB reply, on the event loop, against under a millisecond for the scan. This is what lets the blocked-tool-call line print an argument containing backticks: `` Bash(`echo `whoami``) `` names the command that was actually refused, where dropping the backticks would have named a different one
Expand Down
6 changes: 5 additions & 1 deletion src/bot/core.py
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,7 @@
from ..exceptions import ClaudeCodeTelegramError
from .features.registry import FeatureRegistry
from .orchestrator import MessageOrchestrator
from .utils.telegram_retry import retry_telegram_network

logger = structlog.get_logger()

Expand Down Expand Up @@ -105,6 +106,7 @@ async def initialize(self) -> None:
proxy_url = os.environ.get("HTTPS_PROXY") or os.environ.get("HTTP_PROXY")
if proxy_url:
builder.proxy(proxy_url)
builder.get_updates_proxy(proxy_url)
logger.info("Proxy configured", proxy=_redact_proxy_url(proxy_url))

self.app = builder.build()
Expand Down Expand Up @@ -343,7 +345,9 @@ async def _error_handler(
# Try to notify user
if update and update.effective_message:
try:
await update.effective_message.reply_text(user_message)
await retry_telegram_network(
lambda: update.effective_message.reply_text(user_message)
)
except Exception:
logger.exception("Failed to send error message to user")

Expand Down
91 changes: 58 additions & 33 deletions src/bot/handlers/message.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@

import structlog
from telegram import InputMediaPhoto, Update
from telegram.error import BadRequest, NetworkError
from telegram.ext import ContextTypes

from ...claude.exceptions import (
Expand All @@ -26,6 +27,7 @@
should_send_as_photo,
validate_image_path,
)
from ..utils.telegram_retry import retry_telegram_network

logger = structlog.get_logger()

Expand Down Expand Up @@ -305,6 +307,7 @@ async def handle_text_message(
# Get services
rate_limiter: Optional[RateLimiter] = context.bot_data.get("rate_limiter")
audit_logger: Optional[AuditLogger] = context.bot_data.get("audit_logger")
progress_msg = None

logger.info(
"Processing text message", user_id=user_id, message_length=len(message_text)
Expand All @@ -323,12 +326,18 @@ async def handle_text_message(
return

# Send typing indicator
await update.message.chat.send_action("typing")
try:
await update.message.chat.send_action("typing")
except NetworkError as exc:
# Typing indicators are cosmetic and must not abort a real request.
logger.debug("Failed to send typing action, ignoring", error=str(exc))

# Create progress message
progress_msg = await update.message.reply_text(
"🤔 Processing your request...",
reply_to_message_id=update.message.message_id,
progress_msg = await retry_telegram_network(
lambda: update.message.reply_text(
"🤔 Processing your request...",
reply_to_message_id=update.message.message_id,
)
)

# Get Claude integration and storage from context
Expand Down Expand Up @@ -438,7 +447,10 @@ async def stream_handler(update_obj):
]

# Delete progress message
await progress_msg.delete()
try:
await progress_msg.delete()
except Exception as exc:
logger.debug("Failed to delete progress message, ignoring", error=str(exc))

# Use MCP-collected images (from send_image_to_user tool calls)
images: list[ImageAttachment] = mcp_images
Expand Down Expand Up @@ -494,42 +506,49 @@ async def stream_handler(update_obj):
# Send formatted responses (may be multiple messages)
for i, message in enumerate(formatted_messages):
try:
await update.message.reply_text(
message.text,
parse_mode=message.parse_mode,
reply_markup=message.reply_markup,
reply_to_message_id=(
update.message.message_id if i == 0 else None
),
await retry_telegram_network(
lambda: update.message.reply_text(
message.text,
parse_mode=message.parse_mode,
reply_markup=message.reply_markup,
reply_to_message_id=(
update.message.message_id if i == 0 else None
),
)
)
if i < len(formatted_messages) - 1:
await asyncio.sleep(0.5)
except Exception as send_err:
except BadRequest as send_err:
logger.warning(
"Failed to send HTML response, retrying as plain text",
"Failed to send formatted response, retrying as plain text",
error=str(send_err),
message_index=i,
)
try:
await update.message.reply_text(
message.text,
reply_markup=message.reply_markup,
reply_to_message_id=(
update.message.message_id if i == 0 else None
),
await retry_telegram_network(
lambda: update.message.reply_text(
message.text,
reply_markup=message.reply_markup,
reply_to_message_id=(
update.message.message_id if i == 0 else None
),
)
)
except Exception as plain_err:
plain_error_text = str(plain_err)[:150]
logger.error(
"Failed to send plain text fallback response",
error=str(plain_err),
error=plain_error_text,
)
await update.message.reply_text(
f"Failed to deliver response "
f"(Telegram error: {str(plain_err)[:150]}). "
f"Please try again.",
reply_to_message_id=(
update.message.message_id if i == 0 else None
),
await retry_telegram_network(
lambda: update.message.reply_text(
f"Failed to deliver response "
f"(Telegram error: {plain_error_text}). "
f"Please try again.",
reply_to_message_id=(
update.message.message_id if i == 0 else None
),
)
)

# Send images separately
Expand Down Expand Up @@ -635,12 +654,18 @@ async def stream_handler(update_obj):

except Exception as e:
# Clean up progress message if it exists
try:
await progress_msg.delete()
except Exception as delete_error:
logger.debug("Failed to delete progress message", error=str(delete_error))
if progress_msg is not None:
try:
await progress_msg.delete()
except Exception as delete_error:
logger.debug(
"Failed to delete progress message", error=str(delete_error)
)

await update.message.reply_text(_format_error_message(e), parse_mode="HTML")
error_message = _format_error_message(e)
await retry_telegram_network(
lambda: update.message.reply_text(error_message, parse_mode="HTML")
)

# Log failed processing
if audit_logger:
Expand Down
63 changes: 38 additions & 25 deletions src/bot/orchestrator.py
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,7 @@
InputMediaPhoto,
Update,
)
from telegram.error import BadRequest, NetworkError
from telegram.ext import (
Application,
CallbackQueryHandler,
Expand All @@ -40,6 +41,7 @@
should_send_as_photo,
validate_image_path,
)
from .utils.telegram_retry import retry_telegram_network

logger = structlog.get_logger()

Expand Down Expand Up @@ -993,7 +995,11 @@ async def agentic_text(
return

chat = update.message.chat
await chat.send_action("typing")
try:
await chat.send_action("typing")
except NetworkError as exc:
# Typing indicators are cosmetic and must not abort a real request.
logger.debug("Failed to send typing action, ignoring", error=str(exc))

verbose_level = self._get_verbose_level(context)

Expand All @@ -1002,8 +1008,8 @@ async def agentic_text(
stop_kb = InlineKeyboardMarkup(
[[InlineKeyboardButton("Stop", callback_data=f"stop:{user_id}")]]
)
progress_msg = await update.message.reply_text(
"Working...", reply_markup=stop_kb
progress_msg = await retry_telegram_network(
lambda: update.message.reply_text("Working...", reply_markup=stop_kb)
)

# Register active request for stop callback
Expand Down Expand Up @@ -1171,38 +1177,45 @@ async def agentic_text(
if not message.text or not message.text.strip():
continue
try:
await update.message.reply_text(
message.text,
parse_mode=message.parse_mode,
reply_markup=None, # No keyboards in agentic mode
reply_to_message_id=(
update.message.message_id if i == 0 else None
),
await retry_telegram_network(
lambda: update.message.reply_text(
message.text,
parse_mode=message.parse_mode,
reply_markup=None, # No keyboards in agentic mode
reply_to_message_id=(
update.message.message_id if i == 0 else None
),
)
)
if i < len(formatted_messages) - 1:
await asyncio.sleep(0.5)
except Exception as send_err:
except BadRequest as send_err:
logger.warning(
"Failed to send HTML response, retrying as plain text",
"Failed to send formatted response, retrying as plain text",
error=str(send_err),
message_index=i,
)
try:
await update.message.reply_text(
message.text,
reply_markup=None,
reply_to_message_id=(
update.message.message_id if i == 0 else None
),
await retry_telegram_network(
lambda: update.message.reply_text(
message.text,
reply_markup=None,
reply_to_message_id=(
update.message.message_id if i == 0 else None
),
)
)
except Exception as plain_err:
await update.message.reply_text(
f"Failed to deliver response "
f"(Telegram error: {str(plain_err)[:150]}). "
f"Please try again.",
reply_to_message_id=(
update.message.message_id if i == 0 else None
),
plain_error_text = str(plain_err)[:150]
await retry_telegram_network(
lambda: update.message.reply_text(
f"Failed to deliver response "
f"(Telegram error: {plain_error_text}). "
f"Please try again.",
reply_to_message_id=(
update.message.message_id if i == 0 else None
),
)
)

# Send images separately if caption wasn't used
Expand Down
50 changes: 50 additions & 0 deletions src/bot/utils/telegram_retry.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,50 @@
"""Retry helpers for transient Telegram transport failures."""

import asyncio
from typing import Awaitable, Callable, TypeVar

import structlog
from telegram.error import BadRequest, NetworkError

logger = structlog.get_logger()

T = TypeVar("T")


async def retry_telegram_network(
operation: Callable[[], Awaitable[T]],
*,
max_attempts: int = 3,
base_delay: float = 0.5,
) -> T:
"""Run a Telegram request again after transient transport failures.

``BadRequest`` inherits from PTB's ``NetworkError`` even though it represents
a permanent request problem (for example invalid HTML), so it must never be
retried here. Callers can handle it separately with a formatting fallback.
"""
if max_attempts < 1:
raise ValueError("max_attempts must be at least 1")
if base_delay < 0:
raise ValueError("base_delay must not be negative")

for attempt in range(max_attempts):
try:
return await operation()
except BadRequest:
raise
except NetworkError as exc:
if attempt == max_attempts - 1:
raise

delay = base_delay * (2**attempt)
logger.warning(
"Transient Telegram request failure, retrying",
attempt=attempt + 1,
max_attempts=max_attempts,
delay_seconds=delay,
error_type=type(exc).__name__,
)
await asyncio.sleep(delay)

raise AssertionError("Telegram retry loop exited unexpectedly")
Loading
Loading