Skip to content

docs: how to sync the store between machines - #60

Merged
mspinola merged 1 commit into
mainfrom
docs/store-syncing
Jul 26, 2026
Merged

mspinola merged 1 commit into
mainfrom
docs/store-syncing

Conversation

@mspinola

Copy link
Copy Markdown
Owner

Norgate needs Windows, so a research Mac or a Linux dashboard is usually a read-only replica of a store produced elsewhere. Nothing documented how to move the files, and the obvious approaches get two things wrong.

What NOT to sync is the load-bearing part

Two of these are correctness issues rather than savings:

Directory Why exclude
_cache/, _raw/ databento producer-internal and rebuildable. On one real store, 188 MB of 270 MB
citpy/ written by cotmetrics on the consumer, not by cotdata. A mirroring sync deletes it or overwrites locally-derived output with nothing
manifest.json the legacy aggregate. Nothing writes it, and it held both producer halves in ONE file, exactly the shape a file sync resolves last-writer-wins

Excluding them takes the payload from 270 MB to about 82 MB.

Recommendations

One producer writes everything. Two producers writing two stores that a sync later merges means every shared file is last-writer-wins.

Give COT its own task. It does not need Norgate, so it should not inherit NDU's "run only when user is logged on" constraint.

Avoid Dropbox and Google Drive, for three specific reasons rather than general dislike:

  • conflict copies (something (conflicted copy).parquet) land inside the store, where any directory scan treats them as real
  • Files On-Demand / Smart Sync placeholders break read_parquet on the machine doing research
  • both sync continuously, so they will replicate mid-producer-run

Example scripts

examples/windows/sync-store.cmd — robocopy /MIR with the exclusions.

It normalises robocopy's exit codes, which is the gotcha here: robocopy returns 0-7 for success (1 = files copied, 3 = copied plus extras) and 8+ for failure. A wrapper that does not handle that makes Task Scheduler report every successful sync as a failure, and "restart on failure" then loops on it.

examples/mac/pull-store.sh — rsync over SSH for a consumer-pull setup, in two passes so a manifest never lands before the data it describes.

Also documented

What is already safe: every parquet is committed with an atomic os.replace, so a concurrent sync sees either the old file or the new one, never a partial. What a mid-run sync can catch is a partially updated set, which is why the sync belongs after the producer task rather than on its own timer.

And how to verify a sync landed, including reading the lag column correctly now that it measures write time.

Checks

Angle-bracket-free .cmd invariant preserved, bash -n clean on the shell script, doc links resolve, suite 134 passed, ruff clean. Docs only.

Norgate needs Windows, so a research Mac or a Linux dashboard is usually a
read-only replica of a store produced elsewhere. Nothing documented how to move
the files, and the obvious approaches get two things wrong.

WHAT NOT TO SYNC is the load-bearing part, and two of the exclusions are
correctness issues rather than savings:

  _cache/ and _raw/  databento producer-internal and rebuildable. On one real
                     store these are 188MB of 270MB, so excluding them takes the
                     payload from 270MB to about 82MB.
  citpy/             written by COTMETRICS on the consumer, not by cotdata. A
                     mirroring sync would delete it or overwrite locally-derived
                     output with nothing.
  manifest.json      the legacy aggregate. Nothing writes it, and it held both
                     producer halves in ONE file, which is exactly the shape a
                     file sync resolves last-writer-wins.

Recommends one producer writing everything, with COT given its own task so it
does not inherit NDU's interactive-session requirement (it does not need Norgate).

Advises against Dropbox and Google Drive for three specific reasons: conflict
copies land inside the store where any directory scan treats them as real,
on-demand placeholder files break read_parquet on the research machine, and both
sync continuously so they will replicate mid-producer-run.

Two example scripts:

  examples/windows/sync-store.cmd  robocopy /MIR with the exclusions. Normalises
                                   robocopy's exit codes, which use 0-7 for
                                   success and 8+ for failure, so Task Scheduler
                                   does not report every successful sync as an
                                   error and retry-loop on it.
  examples/mac/pull-store.sh       rsync over SSH for a consumer-pull setup, in
                                   two passes so a manifest never lands before
                                   the data it describes.

Also notes what is already safe: every parquet is committed with an atomic
os.replace, so a concurrent sync sees the old file or the new one, never a
partial. What a mid-run sync can catch is a partially updated SET, which is why
the sync belongs after the producer task rather than on its own timer.

Suite 134 passed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@mspinola
mspinola merged commit 6b46238 into main Jul 26, 2026
5 checks passed
@mspinola
mspinola deleted the docs/store-syncing branch July 26, 2026 21:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant