Skip to content

fix(watch): Keep portal lists fresh after create, update and delete - #1633

Merged
yahyafakhroji merged 2 commits into
mainfrom
fix/watch-reliability
Oct 6, 2026
Merged

yahyafakhroji merged 2 commits into
mainfrom
fix/watch-reliability

Conversation

@yahyafakhroji

@yahyafakhroji yahyafakhroji commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Problem
Lists and detail pages in the portal could stay stale until a hard refresh. The client never refetched after its live stream dropped, the tab was hidden, or the hub missed events. List pages served cached data without refetching, and some mutations left their list to a watch that had already unmounted. On the server, upstream watches could die silently, and a subscribe that landed on the other replica failed for good.

Solution
The client resyncs every watched query after any gap and treats data as fresh only while it's watched. Every resource writes its own list on create, update and delete, with optimistic rows and "Deleting…" until the server confirms. The hub restarts upstreams per user, tells clients when a channel may have missed changes, and can relay subscribes across pods through Redis behind WATCH_RELAY_ENABLED, which is off by default. Tables show pending and changed rows, and a header notice appears only while live updates are delayed.

How to test
bun test and the component specs pass. In the portal, delete a resource from its detail page and go back to the list, or hide the tab while another session changes a resource and return to it: the list updates without a refresh.

Preview
Create secret
1-create-secret-shows-instantly
Reconnect Indicator
4-reconnect-indicator
Two tab live update
3-two-tab-live-update

Closes #1630
Refs #1614

@yahyafakhroji yahyafakhroji added the bug Something isn't working label Oct 5, 2026
@yahyafakhroji yahyafakhroji self-assigned this Oct 5, 2026
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown

🧪 Test Summary

Job Status
Bun Unit Tests ❌ failure
E2E Regression (0) ✅ success
E2E Regression (1) ✅ success
E2E Regression (2) ✅ success
E2E Regression (3) ✅ success
E2E Smoke ✅ success
Unit Tests ✅ success

View workflow run

Need another run? Use Re-run failed jobs: one shard costs a few minutes, the whole workflow about 26.

@yahyafakhroji
yahyafakhroji force-pushed the fix/watch-reliability branch from 0c3255b to 457b67c Compare October 6, 2026 00:46
@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

🧪 Test Summary

Job Status
Bun Unit Tests ❌ failure
E2E Regression (0) ✅ success
E2E Regression (1) ✅ success
E2E Regression (2) ✅ success
E2E Regression (3) ✅ success
E2E Smoke ✅ success
Unit Tests ✅ success

View workflow run

Need another run? Use Re-run failed jobs: one shard costs a few minutes, the whole workflow about 26.

@yahyafakhroji
yahyafakhroji force-pushed the fix/watch-reliability branch from 457b67c to ae51501 Compare October 6, 2026 01:08
@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

🧪 Test Summary

Job Status
Bun Unit Tests ✅ success
E2E Regression (0) ✅ success
E2E Regression (1) ✅ success
E2E Regression (2) ✅ success
E2E Regression (3) ✅ success
E2E Smoke ✅ success
Unit Tests ✅ success

View workflow run

Need another run? Use Re-run failed jobs: one shard costs a few minutes, the whole workflow about 26.

Upstream watches could end while idle and never restart, missed events
were never signalled, and a subscribe that landed on the other replica
failed for good. The hub now restarts and backs off upstreams per user,
sends a resync event whenever a channel may have missed changes, and
can relay subscribes through Redis behind WATCH_RELAY_ENABLED.
Lists stayed stale until a hard refresh: the client never refetched
after its stream dropped or the hub reported a gap, pages served cached
lists without refetching, and some mutations left their list to a watch
that had already unmounted. The client now resyncs every watched query
after any gap, every resource writes its own list on create, update and
delete, and tables show pending, changed and reconnecting states.
@yahyafakhroji
yahyafakhroji force-pushed the fix/watch-reliability branch from ae51501 to d20d91a Compare October 6, 2026 02:27
@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

🧪 Test Summary

Job Status
Bun Unit Tests ✅ success
E2E Regression (0) ✅ success
E2E Regression (1) ✅ success
E2E Regression (2) ✅ success
E2E Regression (3) ✅ success
E2E Smoke ✅ success
Unit Tests ✅ success

View workflow run

Need another run? Use Re-run failed jobs: one shard costs a few minutes, the whole workflow about 26.

@yahyafakhroji
yahyafakhroji merged commit d36b401 into main Oct 6, 2026
18 checks passed
@yahyafakhroji
yahyafakhroji deleted the fix/watch-reliability branch October 6, 2026 08:22
mattdjenkinson added a commit that referenced this pull request Oct 6, 2026
…1640)

**Problem**
The org projects list sometimes showed projects from other
organizations. They flashed in, then disappeared on their own or after a
refresh. Milo's org-scoped project watch wasn't filtered by
organization, so it streamed every org's projects. Before #1633 the
portal never applied ADDED events to the projects list, so this went
unnoticed. Now that watch events write into the list, the replay sent on
each subscribe added other orgs' projects until the next REST refetch
removed them.

**Solution**
The root cause is fixed in Milo in milo-os/milo#828. This PR is a
backstop so the portal never trusts the stream's scoping by itself.
Watch caches can now take an `accepts` check, and the project list watch
drops any project whose `organizationId` isn't the org being viewed. I
deliberately didn't make the watch apply the REST list's Ready filter,
because that would stop newly created projects showing while they
provision. A unit test covers another org's ADDED being ignored and the
org's own projects still being added; it fails without the guard.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Keep portal lists fresh after create, update and delete

2 participants