Skip to content

app: unsaved work survives a crash - #361

Merged
foxnne merged 1 commit into
mainfrom
app/crash-checkpoint
Oct 11, 2026
Merged

foxnne merged 1 commit into
mainfrom
app/crash-checkpoint

Conversation

@foxnne

@foxnne foxnne commented Oct 10, 2026 •

Copy link
Copy Markdown
Collaborator

Part of plans/RESTART_AND_CRASHES_PLAN.md, step 4.

What changes

A crash no longer loses unsaved work. While anything is unsaved, fizzy keeps a checkpoint of every open document in <config>/checkpoint/. It uses the same format a restart keeps them in (#346: contents, caret, scroll, dirty). A clean quit deletes it. So a checkpoint found at launch means the last run ended without quitting: a crash, a kill -9, a power cut. Its documents open from it as a restart's do, and a toast says "fizzy did not quit cleanly last time. Its unsaved changes are back."

src/editor/Checkpoint.zig:

  • When: the documents are looked at once a second, which reads a dirty flag each and captures nothing.
    • With unsaved work, a checkpoint is taken: about a second after input, never within 3 s of the last one, and at least every 30 s regardless. An agent's edits arrive with no input, so the 30 s rule covers them.
    • With nothing unsaved, the checkpoint is deleted at once.
  • Waking an idle app: fizzy draws only when something happens, so a dvui timer asks for a frame when the next look is due. Without it, typing that landed just after a look waited for the next unrelated event; I hit exactly that while testing. With everything saved there is no timer, so it costs nothing.
  • Off the frame: the state is captured on the UI thread (only the owners can), then written on a thread of its own.
  • In place, by generation (KeptDocuments.saveGeneration): each checkpoint's files are named <generation>-<index>.state. They're written beside the last one's, then session.zon is written to .part and renamed over the last, which is the moment the new checkpoint counts. Then the last generation's files are removed. A crash at any point leaves one whole generation.
    • The folder itself is never swapped or deleted while fizzy runs. A clean quit, or saving everything, empties it. The settings watcher watches the config folder, and on Linux a folder that appears and is renamed or deleted before inotify adds its watch logged an error (the first CI run of this PR failed on that).
    • The settings watcher ignores fizzy's own state folders (checkpoint/, session/, handover/, matched as top-level folders of the config folder). A checkpoint write no longer wakes the app or re-reads settings.zon. That was a cost on every platform, found by the crash-reports session.
  • At launch (Checkpoint.recover, in Editor.init): a restart's session is newer and wins, and the checkpoint is dropped. Otherwise the checkpoint becomes editor.session, read once.
  • For the crash notice (step 3's next PR, coordinated with that session): KeptDocuments.recoveredFromCheckpoint() says this run recovered unsaved work, so a crash and its recovery can show as one message.

SDK impact

  • None

Verified

  • macOS, in a sandbox profile driven through the agent plugin:
    • Recovery: typed into a file with the app then left idle; the checkpoint appeared 1.0 s later. kill -9, relaunch: the typed line was back, unsaved, the caret where it was, and the toast showed (screenshot of the relaunched window).
    • Saving removed the checkpoint about 100 ms later.
    • A clean quit (SIGTERM with nothing unsaved) left no checkpoint.
    • Across a restart's handover (app: a restart hands the window over instead of closing it #356): the edit came through, and the new instance took a fresh checkpoint.
    • Cost: typing into a 5.3 MB file and checkpointing it, the worst frame over the window was 4.0 ms against a 3.6 ms average, at 120 fps.
    • Test: test-integration (371/371) has a new one: two generations leave only the second's files; a checkpoint is read once; a folder with no session.zon is neither read nor kept; a restart's session wins. SettingsWatcher's classify test covers the state folders, and that a plugin named checkpoint is still a plugin.
    • Live: two writes left 1-* then only 2-* files; kill -9 then relaunch recovered both edits, unsaved; saving left the folder empty.
    • Gates: zig build, zig build test (531/532, 1 skipped), check-web, test-sdk-version.
  • Windows: CI cross-compile. The code is the same on every desktop; rename of the folder is the one OS call that differs.
  • Linux: CI.
  • Web: check-web; not built there, since a page has no crash to recover from that way.

Follow-ups

  • Hot exit (quit never asks; everything comes back next launch) is this checkpoint plus a setting. Left for its own PR.
  • The crash notice (step 3) reads recoveredFromCheckpoint().

🤖 Generated with Claude Code

foxnne added a commit that referenced this pull request Oct 10, 2026
Part of `plans/DASHBOARD_PLAN.md`, "What performance work lacked": frame
causes. Coordinated with the session doing the profiler; it takes spans.

## What changes

**The profiler says why each frame happened.** fizzy draws only when
something asks for a frame, so the frame-loop bugs are a frame that
never comes (work waiting for the next unrelated event) and frames that
never stop (something asking every frame). A profile showed where a
frame's time went, but not who asked for the frame. Two of those bugs
came up this week, the crash checkpoint (#361) and the restart handover
(#356), and each cost a round of guessing.

While the profiler records, each frame is put down to one of three
causes:
- **its events,** less the pointer `position` event dvui adds to every
frame;
- **else the `dvui.refresh` calls before it,** by file and line;
- **else something else:** a timer or an animation coming due.

`core.profile.report` gains a `.causes` section with the counts and the
places that asked for frames, most first. Idle under the agent's
profiler, for example:

```zig
.causes = .{
    .frames = 252, .by_events = 0, .by_refresh = 251, .by_other = 1,
    .refreshed_from = .{
        .{ .place = "Host.refresh (a plugin):0", .count = 252 },
    },
},
```

That first reading found a bug of ours: the agent's own `fizzy_profile`
keeps fizzy awake while it measures. The fix is in the agent plugin,
next.

**Where the places come from:**
- **fizzy and the built-in plugins:** dvui already records each
refresh's file and line as a debug log line while
`dvui.debug.logRefresh` is on, which is what `FIZZY_LOG_REFRESH` turns
on. fizzy turns it on while the profiler records. The app's log function
hands those lines to `core.profile.interceptLog`, matched by format at
compile time, so no other log call costs anything, and they're counted
instead of printed.
- For release builds, dvui's debug level is compiled in
(`log_scope_levels`), and its other debug lines are dropped in `logFn`
as before. A ReleaseFast run logged none of them.
- **A plugin dylib:** its refresh lines reach the host as text through
`Host.logLine` and are read the same way (`interceptPluginLine`). Only
its Debug builds produce them.
- **`Host.refresh`,** how every plugin asks for a frame: it wakes the
backend directly, so it's recorded at that point as `Host.refresh (a
plugin)`. Which plugin asked, it doesn't say.

**No layout change:** `FrameCauses` is its own allocation, made the
first frame the profiler records and published under its own key, as
`Lookback` is (#353). The profiler's `abi` stays 4. It starts over each
time the profiler starts recording, so it describes what happened while
someone was looking.

## SDK impact

- [x] Core-only or additive: reaches plugins at the next SDK release
(`core.profile`: `FrameCauses`, `frameCauses`, `interceptLog`,
`ReportOptions.causes`). No fingerprint move (`test-sdk-version`
passes).

## Verified

- [x] macOS, Debug and ReleaseFast, a sandbox profiled through the agent
plugin:
- **Idle:** all frames put down to `Host.refresh`, as above. The same in
ReleaseFast, with no dvui debug lines in its log.
- **With `FIZZY_LOG_REFRESH=1`:** the intercept fires for every refresh
record dvui prints (234 in a run), and the lines still print.
- [x] `fizzy-profile-causes-tests` (new, std-only): the three causes;
the busiest place first; a long path keeping its tail; places past the
32 kept counted together; reset.
- [x] `zig build`, `zig build -Doptimize=ReleaseFast`, `zig build test`
(533/534, 1 skipped), `test-integration` (375/375), `check-web`,
`test-sdk-version`.
- [ ] Windows, Linux: CI.
- **Not covered:** a test that holds dvui to the format strings it logs.
The integration tests take their log function from Zig's test runner, so
the intercept can't run there. If dvui changes the format, refresh
places stop being recorded, and frames show up as "other".

## Follow-ups

- **Which plugin called `Host.refresh`:** a source location through it,
which moves the SDK boundary, so it waits for a release that moves it
anyway.
- **A dylib plugin's own `dvui.refresh`** calls are counted only in its
Debug builds, and only once its dvui copy has refresh records on.
Syncing that flag through `dvui_context` is the seam.
- **Agent:** `fizzy_profile` should measure without driving frames, and
a step to compare before and after.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
While a document has unsaved changes, fizzy keeps every open document as a restart keeps them
(KeptDocuments) in <config>/checkpoint/: a second or so after the person types, and at least
every half minute regardless, so an agent's edits are kept too. A clean quit deletes it. At the
next launch a checkpoint still there means the last run ended without quitting (a crash, a
kill, a power cut): its documents open from it, unsaved changes, caret and all, and a toast says
so. State is captured on the UI thread and written on a thread of its own, to a folder that
replaces the last one only once whole.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@foxnne
foxnne force-pushed the app/crash-checkpoint branch from 14e3c12 to dd26181 Compare October 10, 2026 22:40
@foxnne
foxnne merged commit 36cd31c into main Oct 11, 2026
10 checks passed
@foxnne
foxnne deleted the app/crash-checkpoint branch October 11, 2026 11:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant