When systems are grouped (allow_merge_split = True) and mask_calc_with_accumulation_period = False, the system masks contain only a fraction of the LP objects that were actually tracked. The cause appears to be an overwrite in fill_mask_arrays. A one-line change fixes it in my tests.
Setup --
LPT applied to 6-hourly ERA5 precipitation over the Northern Hemisphere midlatitudes (30–70°N, 0.5° grid). Key options: accumulation_hours = 24, thresh = 1.0 (mm/day), filter_stdev = 1, min_points = 10, min_area = 25000, min_lpt_duration_hours = 24, fall_below_threshold_max_hours = 6, mask_coarse_grid_factor = 1, mask_detailed_output = False, mask_calc_with_accumulation_period = False, mask_calc_with_filter_radius = False, mask_n_cores = 8.
The LPO step was run once, and the LPT step was then run twice on the same objects, changing only the merge/split treatment.
he LP objects themselves cover 39.17% of grid point × time steps, so the break_up run reproduces them well. The allow run tracks slightly more unique objects, yet its composite mask covers about a quarter of the object area. With do_lpt_individual_group_masks = True, the union of all 93 group mask files covers only 1.13%, i.e. less than the composite.
Cause --
In fill_mask_arrays, the first parallel fill writes its results back by assignment:
output = p.starmap(fill_matrix, [(mask_arrays[mask_name][t], i, j, 1) for t, i, j in zip(...)])
Assign output to the correct location.
for t0, t in enumerate(dt_indices):
mask_arrays[mask_name][t] = output[t0]
Each worker receives the array for time t as it was before the batch, fills its own points, and returns a copy. When several LP objects of the same system share a time index, the later assignment overwrites the earlier one, and those points are lost.
With break_up_merge_split = True a system has roughly one object per time step, so duplicates are rare and the overwrite is harmless. With allow_merge_split = True a system carries all branches (here ~256 objects per system over lifetimes of at most 120 time steps), so many objects share a time index and most are lost.
The accumulation-period block further down in the same function combines correctly with sparse_max, which is why the problem does not appear when mask_calc_with_accumulation_period = True (as in the MJO example configuration): the same points are re-filled with OR logic.
Fix--
for t0, t in enumerate(dt_indices):
mask_arrays[mask_name][t] = sparse_max(mask_arrays[mask_name][t], output[t0])
When systems are grouped (allow_merge_split = True) and mask_calc_with_accumulation_period = False, the system masks contain only a fraction of the LP objects that were actually tracked. The cause appears to be an overwrite in fill_mask_arrays. A one-line change fixes it in my tests.
Setup --
LPT applied to 6-hourly ERA5 precipitation over the Northern Hemisphere midlatitudes (30–70°N, 0.5° grid). Key options: accumulation_hours = 24, thresh = 1.0 (mm/day), filter_stdev = 1, min_points = 10, min_area = 25000, min_lpt_duration_hours = 24, fall_below_threshold_max_hours = 6, mask_coarse_grid_factor = 1, mask_detailed_output = False, mask_calc_with_accumulation_period = False, mask_calc_with_filter_radius = False, mask_n_cores = 8.
The LPO step was run once, and the LPT step was then run twice on the same objects, changing only the merge/split treatment.
he LP objects themselves cover 39.17% of grid point × time steps, so the break_up run reproduces them well. The allow run tracks slightly more unique objects, yet its composite mask covers about a quarter of the object area. With do_lpt_individual_group_masks = True, the union of all 93 group mask files covers only 1.13%, i.e. less than the composite.
Cause --
In fill_mask_arrays, the first parallel fill writes its results back by assignment:
output = p.starmap(fill_matrix, [(mask_arrays[mask_name][t], i, j, 1) for t, i, j in zip(...)])
Assign output to the correct location.
for t0, t in enumerate(dt_indices):
mask_arrays[mask_name][t] = output[t0]
Each worker receives the array for time t as it was before the batch, fills its own points, and returns a copy. When several LP objects of the same system share a time index, the later assignment overwrites the earlier one, and those points are lost.
With break_up_merge_split = True a system has roughly one object per time step, so duplicates are rare and the overwrite is harmless. With allow_merge_split = True a system carries all branches (here ~256 objects per system over lifetimes of at most 120 time steps), so many objects share a time index and most are lost.
The accumulation-period block further down in the same function combines correctly with sparse_max, which is why the problem does not appear when mask_calc_with_accumulation_period = True (as in the MJO example configuration): the same points are re-filled with OR logic.
Fix--
for t0, t in enumerate(dt_indices):
mask_arrays[mask_name][t] = sparse_max(mask_arrays[mask_name][t], output[t0])