Skip to content

Batch transfer sample file and help texts do not match what the batch file parser requires #418

Description

@samuelvkwong

Summary

The shipped batch transfer sample file has no Pseudonym column, so uploading it unchanged is rejected for every user without batch_transfer.can_transfer_unpseudonymized (the default). The help texts that send batch query users on to batch transfer do not state the two preconditions the parser and serializer enforce: a pseudonym per row, and the 500-row non-staff limit. This is the documentation half of the batch query -> batch transfer round trip discussed in the September 2026 ADIT/RADIS brainstorming; the feature half is a separate issue.

Current behaviour

  • adit/static/samples/batch_transfer_sample.xlsx has 12 columns (PatientID, PatientName, BirthDate, StudyDate, StudyTime, StudyDescription, ModalitiesInStudy, NumberOfStudyRelatedInstances, AccessionNumber, StudyInstanceUID, SeriesInstanceUID, SeriesDescription) and no Pseudonym column.
  • adit/batch_transfer/serializers.py:70-73 rejects a row without a pseudonym unless the user has the permission; adit/batch_transfer/forms.py:126-128 passes that permission in. So the sample fails validation for default users.
  • The parser reads only PatientID, StudyInstanceUID, SeriesInstanceUID and Pseudonym (adit/batch_transfer/parsers.py:9-14); every other column is ignored (adit/core/parsers.py:48-52).
  • The batch query exporter writes Pseudonym only when at least one query task had one (adit/batch_query/utils/exporters.py:13-19, :31-35). adit/static/samples/batch_query_sample.xlsx has no Pseudonym column either, so a user who follows both samples never gets the column in the export and has to add it by hand. The batch query help does list Pseudonym as an optional input column (adit/batch_query/templates/batch_query/_batch_query_help.html:73-76, docs/user-docs/user-guide.md:91) but does not say it is needed for the transfer step.
  • adit/batch_transfer/templates/batch_transfer/_batch_transfer_help.html:23-25, _batch_query_help.html:84-88 and docs/user-docs/user-guide.md:101 say the exported file "can be used for the batch transfer" without either precondition. The pseudonym requirement appears only in the column list (_batch_transfer_help.html:49-53, user-guide.md:111).
  • MAX_BATCH_QUERY_SIZE = 1000 counts query tasks (adit/settings/base.py:510); MAX_BATCH_TRANSFER_SIZE = 500 counts non-empty rows before grouping by study (:514, adit/core/parsers.py:62-63). One query task can return many studies or series, so an export can easily exceed 500 rows and the user has to split the file. The form calls the limit "tasks" (adit/batch_transfer/forms.py:92-95, :135-138), which is not what is counted.
  • The only test of the export is a 200 smoke test on a job with no results (adit/batch_query/tests/test_views.py:115-122).

Requested behaviour

  1. Add a Pseudonym column to both sample files.
  2. Next to the "can be used for the batch transfer" sentence in the three help locations, state that a Pseudonym column is required unless the user may transfer unpseudonymized, and that the non-staff limit is 500 rows so a large export may need splitting.
  3. Add an exporter round-trip test: write_results on a job with BatchQueryResultFactory rows (adit/batch_query/factories.py:41), read the file back, assert the header, and assert BatchTransferFileParser.parse accepts it.

Open questions

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingdocumentationImprovements or additions to documentation

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions