Repository navigation
docs: how to sync the store between machines - #60
Merged
Merged
Conversation
Norgate needs Windows, so a research Mac or a Linux dashboard is usually a
read-only replica of a store produced elsewhere. Nothing documented how to move
the files, and the obvious approaches get two things wrong.
WHAT NOT TO SYNC is the load-bearing part, and two of the exclusions are
correctness issues rather than savings:
_cache/ and _raw/ databento producer-internal and rebuildable. On one real
store these are 188MB of 270MB, so excluding them takes the
payload from 270MB to about 82MB.
citpy/ written by COTMETRICS on the consumer, not by cotdata. A
mirroring sync would delete it or overwrite locally-derived
output with nothing.
manifest.json the legacy aggregate. Nothing writes it, and it held both
producer halves in ONE file, which is exactly the shape a
file sync resolves last-writer-wins.
Recommends one producer writing everything, with COT given its own task so it
does not inherit NDU's interactive-session requirement (it does not need Norgate).
Advises against Dropbox and Google Drive for three specific reasons: conflict
copies land inside the store where any directory scan treats them as real,
on-demand placeholder files break read_parquet on the research machine, and both
sync continuously so they will replicate mid-producer-run.
Two example scripts:
examples/windows/sync-store.cmd robocopy /MIR with the exclusions. Normalises
robocopy's exit codes, which use 0-7 for
success and 8+ for failure, so Task Scheduler
does not report every successful sync as an
error and retry-loop on it.
examples/mac/pull-store.sh rsync over SSH for a consumer-pull setup, in
two passes so a manifest never lands before
the data it describes.
Also notes what is already safe: every parquet is committed with an atomic
os.replace, so a concurrent sync sees the old file or the new one, never a
partial. What a mid-run sync can catch is a partially updated SET, which is why
the sync belongs after the producer task rather than on its own timer.
Suite 134 passed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Norgate needs Windows, so a research Mac or a Linux dashboard is usually a read-only replica of a store produced elsewhere. Nothing documented how to move the files, and the obvious approaches get two things wrong.
What NOT to sync is the load-bearing part
Two of these are correctness issues rather than savings:
_cache/,_raw/citpy/manifest.jsonExcluding them takes the payload from 270 MB to about 82 MB.
Recommendations
One producer writes everything. Two producers writing two stores that a sync later merges means every shared file is last-writer-wins.
Give COT its own task. It does not need Norgate, so it should not inherit NDU's "run only when user is logged on" constraint.
Avoid Dropbox and Google Drive, for three specific reasons rather than general dislike:
something (conflicted copy).parquet) land inside the store, where any directory scan treats them as realread_parqueton the machine doing researchExample scripts
examples/windows/sync-store.cmd—robocopy /MIRwith the exclusions.It normalises robocopy's exit codes, which is the gotcha here: robocopy returns 0-7 for success (1 = files copied, 3 = copied plus extras) and 8+ for failure. A wrapper that does not handle that makes Task Scheduler report every successful sync as a failure, and "restart on failure" then loops on it.
examples/mac/pull-store.sh—rsyncover SSH for a consumer-pull setup, in two passes so a manifest never lands before the data it describes.Also documented
What is already safe: every parquet is committed with an atomic
os.replace, so a concurrent sync sees either the old file or the new one, never a partial. What a mid-run sync can catch is a partially updated set, which is why the sync belongs after the producer task rather than on its own timer.And how to verify a sync landed, including reading the lag column correctly now that it measures write time.
Checks
Angle-bracket-free
.cmdinvariant preserved,bash -nclean on the shell script, doc links resolve, suite 134 passed, ruff clean. Docs only.