Skip to content

Support drop_last=False: emit the ragged final batch of a finite input #139

Description

@nclack

What

damacy groups pushed samples into fixed-size batches of samples_per_batch, and a batch becomes available to the consumer only once it has accumulated exactly samples_per_batch samples. When the number of samples in a finite input is not a multiple of samples_per_batch, the trailing < samples_per_batch samples never reach that threshold — they are never formed into a batch and never surface. They are silently discarded.

In the vocabulary of a standard batched dataloader (e.g. PyTorch's DataLoader), this is drop_last=True, and it is currently the only available behavior. There is no drop_last=False: no way to ask that the final partial batch be emitted (as a smaller, truncated batch) so that every pushed sample is accounted for.

Why it matters

For a finite input — evaluation or inference over a fixed dataset — every sample is expected to be processed exactly once. Silently dropping the ragged tail loses up to samples_per_batch - 1 samples with no signal, which corrupts counts/metrics and is surprising. Batched dataloaders expose drop_last precisely because both behaviors are legitimate: training often prefers drop_last=True (uniform batch shapes), while eval/inference usually needs drop_last=False (no sample left behind).

For a streaming / unbounded input there is no "final" batch, so the distinction does not arise; such consumers are unaffected.

The gap

The behavior is fixed at drop_last=True. A caller consuming a finite input has no way to retrieve its ragged final batch.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions