You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
conda→UV migration: the postprocessor launcher requires conda, and two things now assert that requirement #525
Filed so the migration does not trip over a check that is correct today and will be wrong the moment it lands. Context: reports/conda_to_uv_migration_investigation.md.
The situation
The platform has two environment mechanisms in the same delivery, and they are used by different legs of the same run:
A rented pod can therefore satisfy the training leg and not the delivery leg. Since the postprocessor is the last step, that failure lands after every GPU hour is spent — on the 300-lesson roster that is roughly 16 of them.
What now asserts the conda requirement
Both added 2026-09-30 in #524, and both must change when the migration lands:
tools/podrun/pod_run_fao_delivery.sh, in preflight:
command -v conda >/dev/null 2>&1|| note "conda — the un_fao postprocessor launcher requires it (tools/launcher/postprocessor.sh:72). A uv venv is not enough."
This is a good check today — it converts a 16-GPU-hour-late failure into a seconds-long refusal — and a wrong one after migration, where it would refuse a perfectly capable pod.
tests/test_fao_delivery_runner.py::test_it_checks_conda_because_the_postprocessor_needs_it, which asserts that check exists. It will need deleting or inverting, not adjusting.
docs/runpod_run_guide.md Phase 4b also tells the operator that conda is one of the two most likely causes of a late failure. That prose goes stale at the same moment.
Why this is worth an issue rather than a comment
The check and its test were written because the requirement was invisible and cost a run. After migration the requirement disappears, and a check for a vanished requirement is worse than no check: it refuses correct machines and it tells an operator to install something they do not need. The failure mode inverts, so it cannot be left to be noticed.
What the migration should do here
Decide the target for the postprocessor prefix.envs/views-postprocessing is a conda prefix shared by both postprocessors (C-116), and that sharing is load-bearing: it is why pins have to be declared identically in un_fao and un_crafd requirements. Whatever replaces it needs the same property or the C-116 reasoning has to be revisited, not just the mechanism.
Remove the preflight check and its test in the same change that removes the conda dependency, so the tree never has a check for a requirement that no longer exists.
Update docs/runpod_run_guide.md Phase 4b, which currently names conda as a likely late failure.
Re-check tools/launcher/postprocessor.sh's other conda-coupled behaviour, not only the four calls above — notably the #385 pin verification, which reads direct_url.json out of sysconfig.get_paths()['purelib']afterconda activate. That path resolution is what makes the check see the right prefix, and it has to keep doing so.
reports/conda_to_uv_migration_investigation.md is the standing investigation; this issue is a downstream consequence for it to absorb rather than a competing plan.
Filed so the migration does not trip over a check that is correct today and will be wrong the moment it lands. Context:
reports/conda_to_uv_migration_investigation.md.The situation
The platform has two environment mechanisms in the same delivery, and they are used by different legs of the same run:
uv venv --python 3.11 /workspace/venvtools/podrun/pod_run_model.shun_fao/un_crafdpostprocessorsconda shell.bash hook,conda create --prefix,conda activatetools/launcher/postprocessor.sh:72,81,113,117A rented pod can therefore satisfy the training leg and not the delivery leg. Since the postprocessor is the last step, that failure lands after every GPU hour is spent — on the 300-lesson roster that is roughly 16 of them.
What now asserts the conda requirement
Both added 2026-09-30 in #524, and both must change when the migration lands:
tools/podrun/pod_run_fao_delivery.sh, in preflight:This is a good check today — it converts a 16-GPU-hour-late failure into a seconds-long refusal — and a wrong one after migration, where it would refuse a perfectly capable pod.
tests/test_fao_delivery_runner.py::test_it_checks_conda_because_the_postprocessor_needs_it, which asserts that check exists. It will need deleting or inverting, not adjusting.docs/runpod_run_guide.mdPhase 4b also tells the operator that conda is one of the two most likely causes of a late failure. That prose goes stale at the same moment.Why this is worth an issue rather than a comment
The check and its test were written because the requirement was invisible and cost a run. After migration the requirement disappears, and a check for a vanished requirement is worse than no check: it refuses correct machines and it tells an operator to install something they do not need. The failure mode inverts, so it cannot be left to be noticed.
What the migration should do here
envs/views-postprocessingis a conda prefix shared by both postprocessors (C-116), and that sharing is load-bearing: it is why pins have to be declared identically inun_faoandun_crafdrequirements. Whatever replaces it needs the same property or the C-116 reasoning has to be revisited, not just the mechanism.docs/runpod_run_guide.mdPhase 4b, which currently names conda as a likely late failure.tools/launcher/postprocessor.sh's other conda-coupled behaviour, not only the four calls above — notably the#385pin verification, which readsdirect_url.jsonout ofsysconfig.get_paths()['purelib']afterconda activate. That path resolution is what makes the check see the right prefix, and it has to keep doing so.Related
views-datafactoryfloors, currently deferred intests/test_requirements_hygiene.DEFERRED_PACKAGES).reports/conda_to_uv_migration_investigation.mdis the standing investigation; this issue is a downstream consequence for it to absorb rather than a competing plan.