fix(cloudfront): hold a function to a CPU-time budget, not wall-clock time - #2523
Merged
Merged
Conversation
… time TestFunction and TestConnectionFunction ran the handler on a worker thread and failed it with "function execution exceeded the 250ms time limit" when the thread had not answered within 250ms of wall-clock time, counting thread spawn, runtime setup and any time the host spent not running the thread. On a loaded CI runner a trivial echo handler hit that limit: the cloudfront_test_function e2e tests needed a retry in 4 of the last 26 E2E runs, and nextest's retried-failure output (#2520) captured the error on test_function_echoes_request. Real CloudFront Functions are bounded by compute (~1ms of CPU per request), not elapsed time. The worker now measures its own CPU time (CLOCK_THREAD_CPUTIME_ID; elapsed time where the platform has no thread CPU clock) and a run that used more than 250ms of it fails with the same error, keeping whatever it logged. ComputeUtilization is derived from that CPU time too. The caller's recv_timeout becomes a 10s wall-clock safety net for a run that never returns. boa's loop and recursion caps still stop `while(1){}`.
The budget was only checked after the handler returned, so a CPU-bound handler that boa's loop and recursion caps cannot stop (catastrophic regex backtracking inside one builtin call) held the caller until the wall-clock safety net instead of being cut off at 250ms. The worker now hands the caller a handle to its CPU clock as it starts (pthread_getcpuclockid on Linux, the thread's Mach port on macOS), and the caller polls it every 5ms, abandoning the run with the time-limit error as soon as it passes the budget. Where no thread CPU clock exists, elapsed time stands in. The worker's own final check stays for a run that crosses the budget between polls; the safety net drops to 5s. Tests: a backtracking regex that runs for seconds is cut off in ~0.26s (verified to fail, waiting the full 5s, with the watchdog check disabled), and the busy-work clock test is bounded by work done rather than elapsed time so a loaded runner cannot flake it.
…ding - Charge nothing against the budget until the worker reports in: a thread the host is slow to schedule has used no compute, and only the wall-clock safety net applies before it starts. Previously the caller counted elapsed time from before the spawn, moving the false alarm into thread startup. - Send the worker's own starting CPU reading with its clock, so CPU it burns before the caller reads the clock is counted and both checks measure from the same point. - Test: a worker delayed longer than the whole budget before starting still runs a trivial handler successfully (fails with the time-limit error when pre-start wall time is charged).
TestFunction and TestConnectionFunction called the blocking runner from the async request path, holding a tokio worker thread for the whole wait. With the compute budget measured in CPU time, that wait can reach the 5s wall-clock safety net when the host stalls the worker, and on a current-thread runtime it stalls everything else. Both handlers are async now and run the wait through spawn_blocking; their state lock is scoped so no guard is held across the await. Test: a current-thread runtime keeps ticking while a CPU-bound handler runs to its budget (ticks 0 times when the runner is called directly).
This was referenced Sep 14, 2026
…on-cpu-time-limit
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
cloudfront_test_functione2e tests needed a retry in 4 of the last 26 E2E runs. With #2520's retried-failure output, the first attempt's error is now on record, fromtest_function_echoes_request:That handler is a trivial echo. It "exceeded" the limit because the limit was 250ms of wall-clock time on the caller, covering thread spawn, JS runtime setup, and any time a loaded CI runner simply didn't run the worker thread. Locally, setup takes 1-3ms even under full CPU load, which is why it never reproduced here.
Real CloudFront Functions are bounded by compute (~1ms of CPU per request), not elapsed time.
Change
crates/fakecloud-cloudfront/src/js_runtime.rs:pthread_getcpuclockidon Linux, the thread's Mach port viathread_infoon macOS; elapsed time on platforms with neither). The caller polls every 5ms and abandons the run with the sameexceeded the 250ms time limiterror the moment the budget is spent, so a CPU-bound handler is still cut off at ~250ms. The worker re-checks its own CPU time when it finishes, for a run that crossed the budget between polls; that path keeps the handler's logs.ComputeUtilizationis derived from the same CPU time.TestFunction/TestConnectionFunctionare async and await the runner throughspawn_blocking, so a stalled worker cannot pin a tokio worker thread for the safety-net duration; their state lock is scoped so no guard crosses the await.boa's loop-iteration and recursion caps are unchanged and still stop
while(1){}.Surfaces
website/content/docs/services/cloudfront.mddescribes the compute budget and safety net.libcadded as acfg(unix)dependency offakecloud-cloudfront(already in the lockfile).Test plan
New unit tests in
js_runtime.rs, each checked to fail against the behavior it guards:a_cpu_bound_handler_is_cut_off_at_its_budgeta_worker_slow_to_start_is_not_charged_for_the_waitexceeded the 250ms time limita_descheduled_thread_accrues_no_computewaiting_on_a_handler_leaves_the_runtime_freea_run_over_its_compute_budget_reports_the_time_limit,a_worker_that_does_not_report_back_hits_the_wall_clock_safety_net,busy_work_accrues_computefakecloud-cloudfront: 81/81 tests;cloudfront_test_functione2e 4/4; clippy-D warningsand fmt clean.Summary by cubic
Holds CloudFront Function test handlers to a CPU-time budget instead of wall-clock time, fixing false failures on loaded CI runners where a trivial echo handler exceeded the old 250ms elapsed-time limit.
spawn_blocking, so a stalled worker no longer pins a tokio worker thread.ComputeUtilizationis derived from the same CPU time.Written for commit f32caab. Summary will update on new commits.