diff --git a/doc/3_configuration.md b/doc/3_configuration.md index 4f5f9c4..16af249 100644 --- a/doc/3_configuration.md +++ b/doc/3_configuration.md @@ -59,7 +59,7 @@ Changing `IMGFORGE_DEFAULT_FORMAT` uses a separate cache namespace for format-le | Variable | Default | Description & tips | | ------------------------- | ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `IMGFORGE_BASE_URL` | unset | Prefix prepended to source references that carry no scheme, so URLs can name only a path (`/unsafe/rs:fit:200:0/aW1hZ2UucG5n`). A reference that already names a scheme is used as-is. | -| `IMGFORGE_ALLOWED_SOURCES` | unset | Comma-separated allowlist of URL prefixes. `https://*.example.com/` permits exactly one subdomain label — `images.example.com` but not `example.com` or `a.b.example.com`. Anything not matching fails with `400 Bad Request`. | +| `IMGFORGE_ALLOWED_SOURCES` | unset | Comma-separated allowlist of URL prefixes. `https://*.example.com/` permits exactly one subdomain label — `images.example.com` but not `example.com` or `a.b.example.com`. The wildcard has to *end* the host, so `example.com.attacker.test` is rejected however the entry is spelled, and a host smuggled past `user@` is too. Anything not matching fails with `400 Bad Request`. | | `IMGFORGE_USER_AGENT` | `imgforge/` | `User-Agent` sent when fetching a source image. Some origins rate-limit or block by user agent. | | `IMGFORGE_MAX_REDIRECTS` | `10` | How many redirects a source fetch may follow. A source that redirects forever otherwise ties up a worker for the whole download timeout. | @@ -114,7 +114,7 @@ Each of these sets the starting value for a processing option that a URL can sti | Variable | Default | Description & tips | | ------------------------------ | --------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `IMGFORGE_PATH_PREFIX` | unset | Mounts every route under a prefix (`/imgforge`), for sharing a hostname with another service. Leading and trailing slashes are normalised away. | -| `IMGFORGE_HEALTH_CHECK_PATH` | `/health` | Path for the liveness endpoint. `/status` always answers as well, so imgforge's own name and imgproxy's both work without configuration. | +| `IMGFORGE_HEALTH_CHECK_PATH` | `/health` | Path for the liveness endpoint. `/status` always answers as well, so imgforge's own name and imgproxy's both work without configuration. Surrounding slashes are optional. Pointing it at a route imgforge already serves — `/metrics` or `/info` — is refused at startup with a configuration error rather than taken silently. | ## Cache configuration diff --git a/doc/6_request_lifecycle.md b/doc/6_request_lifecycle.md index d61046a..a43f4ce 100644 --- a/doc/6_request_lifecycle.md +++ b/doc/6_request_lifecycle.md @@ -3,20 +3,26 @@ What happens between an incoming request and the returned image, and where each failure surfaces. ``` - request - │ + request ┌──────────────────────────────────────────┐ + │ │ IMGFORGE_TIMEOUT wraps everything below, │ + │ │ so a 408 can come from any stage ──▶ 408 │ + │ └──────────────────────────────────────────┘ ├─ 1. routing & middleware ───── rate limit ──▶ 429 ├─ 2. parse path & authenticate ─ bad layout ──▶ 400 │ bad signature/token ──▶ 403 + ├─ 3. parse options ── invalid value ──▶ 400 + ├─ 4. negotiate ────── Accept, client hints + │ source not allowed ──▶ 400 │ - ├─ 3. cache lookup ─── hit ────────────────────────────┐ + ├─ 5. cache lookup ─── hit ────────────────────────────┐ │ │ miss │ - ├─ 4. fetch source ─── too large / wrong MIME ──▶ 400 │ - ├─ 5. parse options ── invalid value ──▶ 400 │ - ├─ 6. transform ────── over IMGFORGE_TIMEOUT ──▶ 408 │ - ├─ 7. populate cache │ + ├─ 6. fetch source ─── upstream status ──▶ 400 │ + │ over max_src_file_size ──▶ 400 │ + ├─ 7. transform ────── wrong MIME ──▶ 400 │ + │ over max_src_resolution ──▶ 400 │ + ├─ 8. populate cache │ │ │ - └─ 8. respond ◀────────────────────────────────────────┘ + └─ 9. respond ◀───── validator matches ──▶ 304 ────────┘ ``` ## 1. Routing & middleware @@ -29,33 +35,45 @@ The path splits into signature, processing directives, and source segment; an in Unless the signature segment is the literal `unsafe`, imgforge recomputes the HMAC from `IMGFORGE_KEY` and `IMGFORGE_SALT` and returns `403 Forbidden` on a mismatch — see [Signing a URL](4_url_structure.md#signing-a-url). When `IMGFORGE_SECRET` is set, image and info endpoints additionally require `Authorization: Bearer `; a missing or wrong token also returns `403 Forbidden`. -## 3. Cache lookup +## 3. Option parsing & 4. Negotiation + +Directives are parsed on top of the server's configured processing defaults, so the URL always wins over the configuration. Content negotiation then reads `Accept` and may replace the output format; client hints may supply a width or DPR the URL left open. See [Configuration](3_configuration.md). + +The source URL is decoded here, `IMGFORGE_BASE_URL` is applied, and `IMGFORGE_ALLOWED_SOURCES` is checked — before the cache lookup, so a source that is no longer permitted stops being served out of the cache. -With caching enabled, imgforge hashes the full request path and checks the configured backend. A hit returns the stored bytes immediately, skipping stages 4 through 7. `cache_hits_total` and `cache_misses_total` record the outcome. See [Caching](7_caching.md). +## 5. Cache lookup -## 4. Source acquisition +With caching enabled, imgforge checks the configured backend. A hit returns the stored bytes immediately, skipping stages 6 through 8. `cache_hits_total` and `cache_misses_total` record the outcome. See [Caching](7_caching.md). + +## 6. Source acquisition Unless the `raw` option is set, the request first takes a worker permit (at most `IMGFORGE_WORKERS` image operations run at once). The source is then fetched within `IMGFORGE_DOWNLOAD_TIMEOUT` seconds. -`IMGFORGE_MAX_SRC_FILE_SIZE` (or a per-request override) is enforced while the body streams, so an oversized source is abandoned mid-download. `IMGFORGE_ALLOWED_MIME_TYPES` and `IMGFORGE_MAX_SRC_RESOLUTION` are checked at the start of stage 6 instead — the resolution check needs the dimensions, so libvips has already opened the buffer by then. Treat them as limits on what gets *processed*, not as a barrier in front of the decoder. +A non-success status from the origin fails the request with `400 Bad Request` naming that status, rather than handing an error page to the decoder. -A watermark named by `watermark_url` or `IMGFORGE_WATERMARK_PATH` is fetched alongside the source. Failures here return `400 Bad Request` with a short reason; see [Error Troubleshooting](8_error_troubleshooting.md). +`IMGFORGE_MAX_SRC_FILE_SIZE` (or a per-request override) is enforced while the body streams, so an oversized source is abandoned mid-download — that one really does belong to the fetch. -## 5. Option parsing +`IMGFORGE_ALLOWED_MIME_TYPES` and `IMGFORGE_MAX_SRC_RESOLUTION` are checked in stage 7, not stage 6: the resolution check needs the dimensions, so libvips has opened the buffer and a worker slot has already been taken by the time either runs. Treat them as limits on what gets *processed*, not as a barrier in front of the decoder. The exception is `raw` and a matching `skip_processing`, which never enter stage 7 at all — those two run both checks in stage 6, before returning the source bytes, so opting out of processing is not a way around them. That holds on a cache hit as well: both limits are part of the cache key, so a tightened policy retires the entries stored under the looser one rather than being outrun by them. -Directives are parsed into a structured plan. Out-of-range values, invalid booleans, and malformed numbers return `400 Bad Request`. Unrecognised option *names* are logged at debug level and ignored, so a typo silently drops that transformation. The full catalogue is in [Processing Options](5_processing_options.md). +A watermark named by `watermark_url` or `IMGFORGE_WATERMARK_PATH` is fetched alongside the source. Failures here return `400 Bad Request` with a short reason; see [Error Troubleshooting](8_error_troubleshooting.md). -## 6. Image transformation +## 7. Image transformation -Decoding, transformation, and encoding run on Tokio's blocking pool so image work never occupies the async runtime. The stages — DPR scaling, load and EXIF orientation, geometry, canvas, effects, encode — are detailed in [Image Processing Pipeline](12_image_processing_pipeline.md). +Decoding, transformation, and encoding run on Tokio's blocking pool so image work never occupies the async runtime. The stages — DPR scaling, load, colour management, frame split, geometry, canvas, effects, encode — are detailed in [Image Processing Pipeline](12_image_processing_pipeline.md). Timing splits across four metrics: `image_operation_semaphore_wait_duration_seconds` (waiting for a worker permit), `image_operation_blocking_queue_duration_seconds` (waiting for a blocking thread), `image_operation_execution_duration_seconds` (the whole blocking section), and `image_processing_duration_seconds` (the transformation itself). -## 7. Response & caching +`raw` and a matching `skip_processing` both return the source bytes without entering this stage at all. + +## 8. Response & caching On success the bytes are inserted into the cache; a failed write is logged but does not affect the response. imgforge replies `200 OK` with the encoded bytes, the matching `Content-Type`, and an `X-Request-ID` header for log correlation. -## 8. Metrics & logging +The delivery headers are attached here — `Cache-Control`, `ETag`, `Last-Modified`, the canonical `Link`, `Vary`, and the CORS origin, each when configured. `Vary` lists whichever request headers the response actually depends on: `Accept` when format negotiation is enabled, and `Sec-CH-Width`, `Width`, `Sec-CH-DPR` and `DPR` when client hints are. With negotiation off and hints on it therefore carries no `Accept` at all. + +Either validator can short-circuit the response with `304 Not Modified` and no body: an `If-None-Match` that matches the `ETag`, or — when the request sends no `If-None-Match` — an `If-Modified-Since` at or after the `Last-Modified` being sent. The date is compared chronologically rather than as text, so a client whose copy is newer than the origin's timestamp still gets its `304`. `If-None-Match` takes precedence when both are present, so a request carrying a stale entity tag receives the body even if its date would have matched. The bytes were produced either way, so the saving is bandwidth rather than work. + +## 9. Metrics & logging Fetch durations feed `source_image_fetch_duration_seconds` and `source_images_fetched_total`, labelled by outcome. `/metrics` exposes every counter and histogram — see [Prometheus Monitoring](11_prometheus_monitoring.md). @@ -63,9 +81,11 @@ Fetch durations feed `source_image_fetch_duration_seconds` and `source_images_fe | Response | Cause | | ------------------------- | --------------------------------------------------------------------------- | +| `304` | A validator matched: `If-None-Match` against the `ETag` (requires `IMGFORGE_USE_ETAG`), or an `If-Modified-Since` at or after `Last-Modified` (requires `IMGFORGE_LAST_MODIFIED_ENABLED`). Either alone is enough. | | `403` | Invalid signature, an unsigned URL while unsigned mode is off, or a missing/invalid bearer token. | -| `400` | Invalid path, invalid option value, source rejected by a limit, or a failed watermark fetch. Body carries a plain-text reason. | -| `408` | The request exceeded `IMGFORGE_TIMEOUT`. imgforge never returns `504` itself — that comes from a proxy in front of it. | +| `404` | The URL's `expires` timestamp has passed. | +| `400` | Invalid path, invalid option value, source rejected by a limit or by `IMGFORGE_ALLOWED_SOURCES`, a non-success status from the origin, an output format this libvips cannot encode, or a failed watermark fetch. Body carries a plain-text reason; `IMGFORGE_DEVELOPMENT_ERRORS_MODE` appends the underlying error. | +| `408` | The request exceeded `IMGFORGE_TIMEOUT`. The timeout layer wraps the whole router, so this covers the source fetch, the wait for a worker slot, and a watermark fetch as readily as the transform itself — a `408` is not by itself evidence that processing was slow. imgforge never returns `504` itself; that comes from a proxy in front of it. | | `429` | Rate limiter depleted. | | `500` | Unhandled error; logged at `error` level with context. | diff --git a/doc/7_caching.md b/doc/7_caching.md index acf3c77..2918b6d 100644 --- a/doc/7_caching.md +++ b/doc/7_caching.md @@ -1,6 +1,6 @@ # 7. Caching -A cache hit skips fetching and processing entirely, returning stored bytes straight from stage 3 of the [request lifecycle](6_request_lifecycle.md). imgforge uses the [Foyer](https://foyer-rs.github.io/foyer/) cache engine and offers three backends. +A cache hit skips fetching and processing entirely, returning stored bytes straight from stage 5 of the [request lifecycle](6_request_lifecycle.md). imgforge uses the [Foyer](https://foyer-rs.github.io/foyer/) cache engine and offers three backends. | Backend | Storage | Survives restart | Suits | | -------- | --------------------------- | ---------------- | ------------------------------ | @@ -11,9 +11,38 @@ A cache hit skips fetching and processing entirely, returning stored bytes strai ## How caching works -- **Key derivation**: the cache key is the full request path — processing options, `cachebuster`, and output format included. Any difference in the path is a different entry. +- **Key derivation**: the cache key is the full request path — processing options, `cachebuster`, and output format included. Any difference in the path is a different entry. Several things outside the path join it too, because each changes the bytes without changing the URL: + + | Input | Joins the key when | + | ----- | ------------------ | + | Output version | Every processed response. Bumped by any release that changes the bytes an unchanged URL produces, so an upgrade retires entries rather than serving the old output indefinitely. A `raw` response is the one exception: it hands back the origin's own bytes, which an imgforge release does not change, so its key carries no version and its entries survive an upgrade. | + | `IMGFORGE_DEFAULT_FORMAT` | The URL names no format and none was negotiated. | + | Negotiated format | Content negotiation chose one from `Accept`. | + | Client hints | `IMGFORGE_ENABLE_CLIENT_HINTS` is on. The requested width and DPR both join, so a `Width: 320` request and a `Width: 1280` request cannot share an entry. | + | `max_result_dimension` | A ceiling is in force. | + | `max_animation_frames` | A ceiling is in force. | + | `max_animation_frame_resolution` | A ceiling is in force. | + | Resolved source URL | It differs from the request path — that is, `IMGFORGE_BASE_URL` is set. A relative reference names a different image the moment that setting changes. Joins a `raw` key too. | + | `max_src_resolution` | A ceiling is in force. Joins a `raw` key too. | + | `max_src_file_size` | A ceiling is in force. Joins a `raw` key too. | + | `IMGFORGE_ALLOWED_MIME_TYPES` | The list is set. A short digest of the sorted list joins the key, so reordering the variable is not treated as a change. Joins a `raw` key too. | + | `IMGFORGE_WATERMARK_PATH` | The request composites the server-side watermark — that is, it uses `watermark` without a `watermark_url` of its own. Repointing the setting retires the entries composited with the old logo. | + | `IMGFORGE_QUALITY` and the other option defaults | Any of them differs from imgforge's built-in defaults. They seed the parse, so they change the bytes exactly as a URL option would. | + + The last two are configuration rather than a ceiling, and they are in the key for the same reason the ceilings are: **a config change carries no version bump.** `OUTPUT_VERSION` retires entries when a *release* changes the output, which is why it does not have to name every processing detail — but nothing retires them when an *operator* changes what the server produces. Lowering `IMGFORGE_QUALITY` or swapping the logo would otherwise keep serving the old bytes until eviction. + + One residual is worth knowing: the watermark joins the key by its **path**, not its contents. Replacing the file at the same path is invisible to the key. Within a running process it is also invisible to imgforge — the watermark is loaded once and held — so this only matters across a restart with a persistent cache. Change the path, or clear the cache, when you replace the image in place. + + **Every ceiling is checked *after* the cache is consulted**, which is why each one has to be part of the key. Without that, an entry stored while a limit was loose keeps being served once the limit is tightened: the request is answered before it ever reaches the check it should have failed. The last three cover the case where that matters most — `raw` and a matching `skip_processing` hand back origin bytes with no processing between them and the client, so their source limits are the only thing standing between a tightened policy and an entry already in the cache. + + So: **changing any of these retires every entry stored under the previous setting**, and that is the point rather than a cost to avoid. Enabling one that was previously unset does the same thing, because an unset limit contributes nothing to the key and a set one contributes its prefix. Expect a cold cache for the affected URLs after any such change; the orphaned entries are unreachable and age out by eviction. A deployment that sets none of them keeps the keys it already had. - **Population**: rendered bytes are inserted after a successful response. A failed write is logged and does not affect the response. - **Invalidation**: there is no explicit purge. Caches are capacity-limited and evict least-recently-used entries; change the `cachebuster` token to force a miss when an upstream asset changes. +- **What a hit reproduces**: a hit never makes the source request, so each entry stores the origin headers the response depends on. The origin's `Cache-Control` (under `IMGFORGE_CACHE_CONTROL_PASSTHROUGH`) and `Last-Modified` are recorded when the entry is written and served identically on every hit — a `no-store` the origin sent keeps being said after the cache starts answering. `ETag`, the canonical `Link`, and a TTL-derived `Cache-Control` are derived from the bytes or the configuration and are identical either way. + +## Client-side caching + +The cache above saves imgforge work. The headers in [Response delivery](3_configuration.md#response-delivery) save the request entirely: `IMGFORGE_TTL` lets a browser or CDN hold the image without asking again, and `IMGFORGE_USE_ETAG` turns the requests it does make into `304 Not Modified` with no body. They are complementary — a CDN in front of imgforge is usually worth more than either. Metrics: