diff --git a/README.md b/README.md index c2b233f..212d35d 100644 --- a/README.md +++ b/README.md @@ -105,11 +105,26 @@ pytest The suite needs **no credentials, no network and no `.env`** — it mocks the Discord gateway and never calls yt-dlp. It covers queue and player state, the -`_advance` state machine (loop modes, skip, replay, autoplay, idle timeout), -the `services.media` helpers, the voice-state guards, and the command edge -cases. CI runs it on every push and pull request against Python 3.11 and 3.12. - -### Deploy to AWS (t4g.micro, ~$6/mo or free tier) -See **[deploy/README.md](deploy/README.md)** for a full walkthrough: launch script, -provisioning (`deploy/setup.sh`) and a `systemd` service that auto-restarts and -starts on boot. +`_advance` state machine (loop modes, skip, replay, autoplay, idle timeout), the +read-ahead buffer and stream teardown, the `services.media` helpers, the +voice-state guards, the command edge cases, the generated `!help`, the Genius +client, and startup/shutdown behaviour — including that a login survives a DNS +outage and that the process stops inside the time systemd allows. + +A few tests render audio through **real FFmpeg** to measure the pitch and speed +filters, since a wrong filter string looks perfectly reasonable and only shows up +as the wrong playback speed. Those skip automatically if FFmpeg is not installed. + +CI runs everything on every push and pull request against Python 3.11 and 3.12. + +### Deploy to a server +See **[deploy/README.md](deploy/README.md)** for the full walkthrough: +provisioning (`deploy/setup.sh`), a `systemd` service that auto-restarts and +starts on boot, self-updating with `deploy/update.sh`, and a daily timer that +keeps `yt-dlp` current. + +The hosted instance runs on a **self-hosted Linux box on a residential +connection**, not a cloud VM. That is a YouTube decision, not a cost one: from a +datacenter IP the bot needed a cookies file that expired every few weeks, and +from a residential IP the same requests work with none. Deploying to a cloud VM +is still supported and documented, with that caveat. diff --git a/deploy/README.md b/deploy/README.md index 95ea7e2..0876229 100644 --- a/deploy/README.md +++ b/deploy/README.md @@ -1,38 +1,34 @@ -# Deployment — AWS EC2 (t4g.micro, ARM) +# Deployment -This bot is designed to run on the cheapest always-on AWS instance: a -**t4g.micro** (ARM Graviton, 1 GB RAM, free-tier eligible). FFmpeg audio -streaming is CPU-light and memory-frugal thanks to lazy stream resolution, so -1 GB is enough for a single-guild-at-a-time music bot. +The bot runs on a **self-hosted Linux box on a residential connection** — +a spare laptop or mini PC with Ubuntu Server is plenty. It was previously +deployed to an AWS `t4g.micro`; that path still works and is documented at the +bottom, but the move off it was not about cost. -## What gets created +## Why not a cloud VM -| Resource | Value | -|-----------------|--------------------------------------------------| -| Instance type | `t4g.micro` (ARM) | -| AMI | Ubuntu 24.04 LTS (arm64) | -| Disk | 8 GB gp3 | -| Security group | inbound SSH (22) from your IP only; all outbound | -| Service | `loopify-bot` (systemd, auto-restart, boot-start)| +YouTube treats **datacenter IP ranges** (AWS, GCP, …) as suspect. From EC2 the +bot needed a cookies file exported from a logged-in account, and those cookies +expired every few weeks — a recurring chore with a dead bot at the end of it +whenever it was forgotten. -The bot makes only **outbound** connections (Discord, YouTube, SoundCloud), so no -inbound ports beyond SSH are required. +From a residential IP the same requests work **with no cookies at all**, and +time-to-first-byte measured 4.3 s against 9.1 s on EC2. A residential IP does not +*remove* YouTube's bot-checking (see the player-client note below), but combined +with the right client chain it removes the need for credentials. -## 1. Launch the instance +The trade is that the host is now your problem: power, network, and the disk it +boots from. FFmpeg audio streaming is CPU-light and memory-frugal, so the machine +itself barely notices — the service is capped at 768 MB and rarely peaks past +half of that. -The exact AWS CLI commands used to launch and tag the instance live in -[`launch_ec2.sh`](launch_ec2.sh). Run it from a machine with the AWS CLI -configured, or follow it step by step. +## 1. Provision the host -## 2. Provision it - -**Clone** the repo on the instance — do not copy the files over. A git checkout -is what makes `deploy/update.sh` work later and lets the bot report which commit -it is running: +Any Ubuntu 24.04 machine works. **Clone** the repo — do not copy the files over. +A git checkout is what makes `deploy/update.sh` work later, and what lets the bot +report which commit it is running: ```bash -ssh -i .pem ubuntu@ - git clone https://github.com/Isma-L154/LoopifyBot.git ~/LoopifyBot cd ~/LoopifyBot bash deploy/setup.sh @@ -42,33 +38,37 @@ bash deploy/setup.sh systemd units: the `loopify-bot` service and a daily `loopify-ytdlp-update` timer. It is idempotent, so re-running it is safe. -## 3. Add secrets and start +For an always-on box, also worth doing: ignore the lid if it is a laptop +(`logind.conf.d`), mask the suspend targets, and enable `unattended-upgrades`. + +## 2. Add secrets and start -Secrets are **never** committed. Create the `.env` directly on the instance: +Secrets are **never** committed. Create the `.env` directly on the host: ```bash cp .env.example .env -nano .env # fill in DISCORD_TOKEN (and Spotify/Genius if used) +nano .env # fill in DISCORD_TOKEN (and GENIUS_TOKEN for !lyrics) sudo systemctl start loopify-bot -sudo systemctl status loopify-bot -sudo journalctl -u loopify-bot -f # live logs — look for "Logged in as ..." +sudo journalctl -u loopify-bot -f # look for "Logged in as ..." ``` -## Updating the bot later +## 3. Updating later ```bash -ssh -i .pem ubuntu@ cd ~/LoopifyBot && bash deploy/update.sh ``` -`update.sh` pulls, reinstalls dependencies **only if `requirements.txt` -changed**, restarts the service, and then verifies it actually came back up — -printing recent logs and failing loudly if it did not. +`update.sh` pulls, reinstalls dependencies **only if `requirements.txt` changed**, +reinstalls the systemd units, restarts the service, and then verifies it actually +came back — printing recent logs and failing loudly if it did not. + +The units are reinstalled every time on purpose: a pull can change how the bot is +*run* (sandboxing, resource caps, stop timeouts), and restarting alone would keep +the old configuration while the repo claimed otherwise. ### If the host was deployed by copying files instead of cloning -`update.sh` refuses to run and tells you how to convert it in place. The short -version, from the app directory: +`update.sh` refuses to run and tells you how to convert it in place: ```bash git init -b main @@ -79,22 +79,21 @@ git reset --hard origin/main # discards local edits — check first ``` The `--set-upstream-to` line matters: without it `update.sh` has nothing to pull -from and stops with an explanation. - -`.env` and `cookies.txt` are gitignored, so they survive this untouched. +from and stops with an explanation. `.env` and `cookies.txt` are gitignored, so +they survive untouched. ### Knowing what is actually running The bot logs its versions at startup, so `journalctl` answers this directly: -``` -Running commit 0a90877 — yt-dlp 2026.08.19, FFmpeg 6.1.1-3ubuntu5, Python 3.12.3 -``` - ```bash sudo journalctl -u loopify-bot | grep "Running commit" | tail -1 ``` +``` +Running commit 1ea2431 — yt-dlp 2026.08.19, FFmpeg 6.1.1-3ubuntu5, Python 3.12.3 +``` + ## Keeping yt-dlp current — automatically `yt-dlp` is the only dependency deliberately left unpinned. YouTube changes its @@ -116,43 +115,63 @@ fails — a newer dependency is never worth trading a running bot for. `Persistent=true` means it catches up after downtime rather than silently skipping, which matters on a machine that is not on 24/7. -## 🎬 YouTube from cloud IPs — how it's made to work +## 🎬 What actually makes YouTube work -YouTube fights bots on **datacenter IPs** (AWS, GCP…) on two fronts, and the bot -handles both so playback works from EC2: +Four things, in order of how much they matter: -1. **"Sign in to confirm you're not a bot"** → defeated with **cookies** from a - logged-in account (`COOKIES_PATH`). -2. **JS signature ("nsig") challenge** on the web player → solved with **Deno**, - which `setup.sh` installs automatically. -3. **Session-bound stream URLs** (which 403 if FFmpeg fetches them directly) → - avoided by having **yt-dlp stream the audio and pipe it into FFmpeg**, so - yt-dlp (with the cookies/session) does the fetching. This is built into the bot. +1. **The player-client chain.** yt-dlp can impersonate several YouTube clients, + and most of them are bot-checked. Measured from a residential IP against the + same video: `web_embedded` works (~3.2 s), `mweb` works but is slow (~9.3 s), + and `default`, `web`, `android_vr`, `tv`, `ios` and `android_music` are **all** + bot-checked. The chain the bot uses is `web_embedded,mweb,tv_embedded` — + getting this right is what removed the need for cookies. +2. **Deno**, for the JS signature (`nsig`) challenge on the web player. + `setup.sh` installs it. +3. **yt-dlp does the fetching**, streaming the audio to stdout and piping it into + FFmpeg. Handing a `googlevideo` URL straight to FFmpeg gets a 403, because + those URLs are bound to the session that requested them. +4. **A residential IP**, which reduces the bot-checking but does not remove it. -### Keeping YouTube working: refresh the cookies +### Cookies: supported, no longer needed -Cookies are the one thing that expires. When YouTube starts getting blocked, -export fresh ones and drop them in: +`COOKIES_PATH` still works if you want it, and helps with age- or region-gated +videos. It is no longer part of normal operation. If you do use one, export it +from a **throwaway** account — cookies grant access to it — and `chmod 600` it. -```bash -# Export from a browser logged into a THROWAWAY YouTube account, Netscape format -# (e.g. the "Get cookies.txt LOCALLY" extension), then: -scp -i .pem cookies.txt ubuntu@:~/LoopifyBot/cookies.txt -ssh -i .pem ubuntu@ "chmod 600 ~/LoopifyBot/cookies.txt && sudo systemctl restart loopify-bot" -``` +### Always-available fallback: SoundCloud -Use a throwaway account — cookies grant access to it. Refresh every few weeks. +SoundCloud has none of these blocks and needs no credentials: `!play sc: `, +or paste a track/set URL. If YouTube ever blocks a track the bot does not crash — +it posts a message suggesting SoundCloud and moves on to the next track. -### Always-available fallback: SoundCloud +## Operational notes + +**Daily upgrades restart the bot.** With `unattended-upgrades` enabled, +`needrestart` restarts the service whenever it upgrades something the bot links +against. If that upgrade is glibc, the DNS resolver is being replaced at the same +moment, so a login can land on a resolver that is briefly unavailable. The bot +retries the login with backoff rather than exiting (see `utils/startup.py`). -SoundCloud has none of these blocks and needs no cookies: `!play sc: ` or -paste a SoundCloud track/set URL. If YouTube ever blocks a track, the bot doesn't -crash — it posts a message suggesting SoundCloud and moves to the next track. +**Stop timeouts are coupled to the voice timeout.** `TimeoutStopSec` must stay +above `cogs.music.VOICE_CONNECT_TIMEOUT`, because discord.py reuses the voice +*connect* timeout as the deadline for Discord to confirm a *departure* while +closing. With the two the wrong way round, systemd SIGKILLs a shutdown that was +going to finish — and a SIGKILL skips the reaping of the FFmpeg and yt-dlp +children. `tests/test_shutdown.py` enforces the ordering. -## Cost & teardown +## Running on a cloud VM instead -- ~**$6/month** on-demand (or free under the 12-month free tier: 750 h/month). -- To stop billing entirely, terminate the instance: - ```bash - aws ec2 terminate-instances --instance-ids - ``` +Still supported, with the caveat above: from a datacenter IP you will probably +need a cookies file, and it will expire. + +[`launch_ec2.sh`](launch_ec2.sh) holds the AWS CLI commands to launch a +`t4g.micro` (ARM Graviton, 1 GB RAM, free-tier eligible) with a security group +allowing inbound SSH from your IP only. The bot makes only **outbound** +connections, so nothing else needs opening. Then provision it exactly as above. + +Cost is roughly **$6/month** on-demand, or free for 12 months under the free tier +(750 h/month). To stop billing entirely, terminate it: + +```bash +aws ec2 terminate-instances --instance-ids +```