Skip to content

fix(cloudfront): hold a function to a CPU-time budget, not wall-clock time - #2523

Merged
vieiralucas merged 5 commits into
mainfrom
fix/cloudfront-function-cpu-time-limit
Sep 14, 2026
Merged

vieiralucas merged 5 commits into
mainfrom
fix/cloudfront-function-cpu-time-limit

Conversation

@vieiralucas

@vieiralucas vieiralucas commented Sep 14, 2026 •

Copy link
Copy Markdown
Member

Summary

cloudfront_test_function e2e tests needed a retry in 4 of the last 26 E2E runs. With #2520's retried-failure output, the first attempt's error is now on record, from test_function_echoes_request:

unexpected error: Some("function execution exceeded the 250ms time limit")

That handler is a trivial echo. It "exceeded" the limit because the limit was 250ms of wall-clock time on the caller, covering thread spawn, JS runtime setup, and any time a loaded CI runner simply didn't run the worker thread. Locally, setup takes 1-3ms even under full CPU load, which is why it never reproduced here.

Real CloudFront Functions are bounded by compute (~1ms of CPU per request), not elapsed time.

Change

crates/fakecloud-cloudfront/src/js_runtime.rs:

  • The 250ms budget is CPU time on the worker thread. The worker hands the caller its CPU clock and starting reading as it starts (pthread_getcpuclockid on Linux, the thread's Mach port via thread_info on macOS; elapsed time on platforms with neither). The caller polls every 5ms and abandons the run with the same exceeded the 250ms time limit error the moment the budget is spent, so a CPU-bound handler is still cut off at ~250ms. The worker re-checks its own CPU time when it finishes, for a run that crossed the budget between polls; that path keeps the handler's logs.
  • Nothing is charged before the worker starts. A thread the host is slow to schedule has used no compute.
  • A 5s wall-clock safety net (was the 250ms limit itself) abandons a run that never reports back.
  • ComputeUtilization is derived from the same CPU time.
  • The wait runs off the async runtime. TestFunction / TestConnectionFunction are async and await the runner through spawn_blocking, so a stalled worker cannot pin a tokio worker thread for the safety-net duration; their state lock is scoped so no guard crosses the await.

boa's loop-iteration and recursion caps are unchanged and still stop while(1){}.

Surfaces

  • Docs: website/content/docs/services/cloudfront.md describes the compute budget and safety net.
  • No API shape, SDK, conformance or count change. libc added as a cfg(unix) dependency of fakecloud-cloudfront (already in the lockfile).

Test plan

New unit tests in js_runtime.rs, each checked to fail against the behavior it guards:

Test Guards Fails when
a_cpu_bound_handler_is_cut_off_at_its_budget catastrophic regex backtracking (no loop iterations, so boa's caps never trip) returns in ~0.26s watchdog check disabled: waits the full 5s
a_worker_slow_to_start_is_not_charged_for_the_wait a worker delayed longer than the whole budget still runs a trivial handler pre-start wall time charged: exceeded the 250ms time limit
a_descheduled_thread_accrues_no_compute a sleeping thread uses none of the budget thread CPU clock replaced by elapsed time
waiting_on_a_handler_leaves_the_runtime_free a current-thread tokio runtime keeps ticking while a handler runs to its budget runner called directly: 0 ticks
a_run_over_its_compute_budget_reports_the_time_limit, a_worker_that_does_not_report_back_hits_the_wall_clock_safety_net, busy_work_accrues_compute budget, safety net, clock
  • fakecloud-cloudfront: 81/81 tests; cloudfront_test_function e2e 4/4; clippy -D warnings and fmt clean.
  • The Linux clock path is compiled and exercised by CI (only macOS was available locally).

Summary by cubic

Holds CloudFront Function test handlers to a CPU-time budget instead of wall-clock time, fixing false failures on loaded CI runners where a trivial echo handler exceeded the old 250ms elapsed-time limit.

  • The 250ms limit is now computed on the worker thread (via thread CPU clocks on Linux/macOS, elapsed time elsewhere), so scheduling delays don't count against it.
  • A 5s wall-clock safety net replaces the old 250ms wall-clock timeout for runs that never report back.
  • The wait runs off the async runtime via spawn_blocking, so a stalled worker no longer pins a tokio worker thread.
  • ComputeUtilization is derived from the same CPU time.
  • Docs updated to describe the compute budget and safety net.

Written for commit f32caab. Summary will update on new commits.

Review in cubic

… time

TestFunction and TestConnectionFunction ran the handler on a worker thread
and failed it with "function execution exceeded the 250ms time limit" when
the thread had not answered within 250ms of wall-clock time, counting thread
spawn, runtime setup and any time the host spent not running the thread.
On a loaded CI runner a trivial echo handler hit that limit: the
cloudfront_test_function e2e tests needed a retry in 4 of the last 26 E2E
runs, and nextest's retried-failure output (#2520) captured the error on
test_function_echoes_request.

Real CloudFront Functions are bounded by compute (~1ms of CPU per request),
not elapsed time. The worker now measures its own CPU time
(CLOCK_THREAD_CPUTIME_ID; elapsed time where the platform has no thread CPU
clock) and a run that used more than 250ms of it fails with the same error,
keeping whatever it logged. ComputeUtilization is derived from that CPU time
too. The caller's recv_timeout becomes a 10s wall-clock safety net for a run
that never returns. boa's loop and recursion caps still stop `while(1){}`.
The budget was only checked after the handler returned, so a CPU-bound
handler that boa's loop and recursion caps cannot stop (catastrophic regex
backtracking inside one builtin call) held the caller until the wall-clock
safety net instead of being cut off at 250ms.

The worker now hands the caller a handle to its CPU clock as it starts
(pthread_getcpuclockid on Linux, the thread's Mach port on macOS), and the
caller polls it every 5ms, abandoning the run with the time-limit error as
soon as it passes the budget. Where no thread CPU clock exists, elapsed
time stands in. The worker's own final check stays for a run that crosses
the budget between polls; the safety net drops to 5s.

Tests: a backtracking regex that runs for seconds is cut off in ~0.26s
(verified to fail, waiting the full 5s, with the watchdog check disabled),
and the busy-work clock test is bounded by work done rather than elapsed
time so a loaded runner cannot flake it.
…ding

- Charge nothing against the budget until the worker reports in: a thread
  the host is slow to schedule has used no compute, and only the wall-clock
  safety net applies before it starts. Previously the caller counted
  elapsed time from before the spawn, moving the false alarm into thread
  startup.
- Send the worker's own starting CPU reading with its clock, so CPU it
  burns before the caller reads the clock is counted and both checks
  measure from the same point.
- Test: a worker delayed longer than the whole budget before starting
  still runs a trivial handler successfully (fails with the time-limit
  error when pre-start wall time is charged).
TestFunction and TestConnectionFunction called the blocking runner from the
async request path, holding a tokio worker thread for the whole wait. With
the compute budget measured in CPU time, that wait can reach the 5s
wall-clock safety net when the host stalls the worker, and on a
current-thread runtime it stalls everything else.

Both handlers are async now and run the wait through spawn_blocking; their
state lock is scoped so no guard is held across the await. Test: a
current-thread runtime keeps ticking while a CPU-bound handler runs to its
budget (ticks 0 times when the runner is called directly).
@vieiralucas
vieiralucas merged commit 3d71e0a into main Sep 14, 2026
157 checks passed
@vieiralucas
vieiralucas deleted the fix/cloudfront-function-cpu-time-limit branch September 14, 2026 08:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant