Summary
Two concurrent calls to the same cached MapInfra from different threads in one process can both claim and submit the same items. The duplicate writers then expose a second issue on a shared NFS cache: a reader can cache a partially visible np.memmap, and MemmapArray.__load_from_info__ does not retry a non-empty short read before reshaping.
This caused two Slurm arrays to each process the same 106 items, followed by:
ValueError: cannot reshape array of size 117702656 into shape (985040,298)
This appears related to, but distinct from, #304: there was no file is not a database warning in this run. The in-flight SQLite registry remained queryable but its same-PID reentrancy rule allowed two threads to own the same claim.
Environment
- exca
0.5.29
- pinned commit
49afd5d87afc9a3d53549c7cf6f140b44be1409e
- Python 3.12, Linux, Slurm/submitit
- shared NFS cache
- distributed parent job launched with 8 srun ranks
Trigger
A data-preparation routine contains the same cached extractor twice:
- directly as a Slurm-backed extractor, prepared in a
ThreadPoolExecutor;
- nested inside a
ChannelPositions-style extractor, prepared by the main thread.
Both call the same MapInfra with the same cache UID and the same 106 item UIDs concurrently from distinct threads in one PID.
Relevant in-flight behavior:
wait_for_inflight() skips rows whose info.pid == os.getpid();
inflight_session() marks existing same-PID rows as pre_owned;
claim() treats same-PID rows as reentrant ownership.
That is correct for nested calls in one logical session, but also conflates independent concurrent threads.
Evidence: duplicate submission
Both calls submitted all 106 identical items under the same cache UID:
Sent 106 samples for neuralset.extractors.neuro.MegExtractor._get_data,...
into 106 jobs on cluster 'slurm' (eg: 2513754_0)
Sent 106 samples for neuralset.extractors.neuro.MegExtractor._get_data,...
into 106 jobs on cluster 'slurm' (eg: 2513753_0)
Other ranks subsequently saw each item split between those two owners, e.g.:
Waiting for 106 in-flight items (of 106 requested) held by:
2513754 [slurm:PENDING] x55, 2513753 [slurm:PENDING] x51
The cache contains two metadata records for the exact same key in different JSONL files:
#key = Armeni2022Ten/.../sub-001_ses-001_task-compr_meg.ds_0.000_4925.200
shape = [985040, 298]
learnfair0542-3499247-info.jsonl -> data/learnfair0542-3499247.data
learnfair0266-3905239-info.jsonl -> data/learnfair0266-3905239.data
Evidence: stale partial memmap
The failing metadata declares shape (985040, 298), requiring:
- 293,541,920 float32 values
- 1,174,167,680 bytes
At failure time, the cached memmap slice exposed only 117,702,656 float32 values (470,810,624 bytes), producing the reshape exception. After the workers completed, the referenced data file measured 1,174,196,288 bytes, i.e. it was large enough for the declared entry.
MemmapArray.__load_from_info__ currently retries only when data.size == 0:
for _ in range(2):
mm = ctx.cached(cache_key, lambda: np.memmap(...))
data = mm[offset : offset + length]
if data.size:
break
ctx.invalidate(cache_key)
return data.view(dtype=dtype).reshape(shape)
A non-empty but short mapping is therefore retained and reshaped. Because the memmap is resource-cached by filename, later file growth is not visible through that mapping.
Expected
- Independent threads in one PID should not be treated as the same reentrant in-flight session. A claim/session token or
(pid, thread_id) ownership could distinguish them while preserving true nested reentrancy.
MemmapArray should validate the exact requested byte count (data.size == length), invalidate/reopen on a short mapping, and ideally wait/retry while a published file is still becoming visible.
- Data should be durably/atomically visible before its JSONL metadata is published across NFS (or readers should tolerate visibility reordering).
- Duplicate cache-key writers should not make cache selection nondeterministic.
I can help produce a reduced concurrency test if useful.
Summary
Two concurrent calls to the same cached
MapInfrafrom different threads in one process can both claim and submit the same items. The duplicate writers then expose a second issue on a shared NFS cache: a reader can cache a partially visiblenp.memmap, andMemmapArray.__load_from_info__does not retry a non-empty short read before reshaping.This caused two Slurm arrays to each process the same 106 items, followed by:
This appears related to, but distinct from, #304: there was no
file is not a databasewarning in this run. The in-flight SQLite registry remained queryable but its same-PID reentrancy rule allowed two threads to own the same claim.Environment
0.5.2949afd5d87afc9a3d53549c7cf6f140b44be1409eTrigger
A data-preparation routine contains the same cached extractor twice:
ThreadPoolExecutor;ChannelPositions-style extractor, prepared by the main thread.Both call the same
MapInfrawith the same cache UID and the same 106 item UIDs concurrently from distinct threads in one PID.Relevant in-flight behavior:
wait_for_inflight()skips rows whoseinfo.pid == os.getpid();inflight_session()marks existing same-PID rows aspre_owned;claim()treats same-PID rows as reentrant ownership.That is correct for nested calls in one logical session, but also conflates independent concurrent threads.
Evidence: duplicate submission
Both calls submitted all 106 identical items under the same cache UID:
Other ranks subsequently saw each item split between those two owners, e.g.:
The cache contains two metadata records for the exact same key in different JSONL files:
Evidence: stale partial memmap
The failing metadata declares shape
(985040, 298), requiring:At failure time, the cached memmap slice exposed only 117,702,656 float32 values (470,810,624 bytes), producing the reshape exception. After the workers completed, the referenced data file measured 1,174,196,288 bytes, i.e. it was large enough for the declared entry.
MemmapArray.__load_from_info__currently retries only whendata.size == 0:A non-empty but short mapping is therefore retained and reshaped. Because the memmap is resource-cached by filename, later file growth is not visible through that mapping.
Expected
(pid, thread_id)ownership could distinguish them while preserving true nested reentrancy.MemmapArrayshould validate the exact requested byte count (data.size == length), invalidate/reopen on a short mapping, and ideally wait/retry while a published file is still becoming visible.I can help produce a reduced concurrency test if useful.