Conversation
collect_outputs(output_id=...) assigns self._outputs_dict on the way out
even in single-output mode. When output_id names something the runnable
doesn't define, the loop matches nothing, so that assignment replaces the
whole cache with {}.
Outputs are collected eagerly while Galaxy is still up (the output_collectors
in BaseEngine._run_test_cases); BaseEngine.test() then rebuilds structured
data from the cache after Galaxy has shut down. Wiping the cache forces those
lookups back to the API, and planemo dies with a raw bioblend.ConnectionError
and no test report at all - instead of reporting "Expected output [x] not
found in results."
Fixes galaxyproject#1625
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
jmchilton
marked this pull request as ready for review
September 24, 2026 20:03
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Updated: Why is this PR message soooo long 😿 - it just fixes #1625 - that is the PR description 😆 . -John
Fixes #1625 —
planemo testcrashed with a rawbioblend.ConnectionErrorand wrote notest report when a workflow test YAML named an output the workflow doesn't define.
Root cause
collect_outputs()has two modes. Full mode buildsself._outputs_dict; single-outputmode (
output_id=..., used byget_output) returns just that one value. But the finalself._outputs_dict = outputs_dictatplanemo/galaxy/activity.py:710runs in bothmodes. In single-output mode the loop
continues past every non-matching output, so whenoutput_idnames an output the runnable doesn't define, nothing is ever added and thatassignment replaces the whole cache with
{}.That cache is load-bearing. Outputs are collected eagerly while Galaxy is still running —
the
output_collectorsadded in 5ff19c8 ("Ignore missing invocation outputs whenfetching outputs") call
structured_test_datafrom insideGalaxyEngine._run'sensure_runnables_servedblock.BaseEngine.test()then rebuilds the same structured dataafter
_runhas returned and Galaxy has shut down, expecting every lookup to hit the cache.So a single typo'd output label does this:
{}the API, and
bioblendraises — out ofstructured_data, uncaught, no reportThe "Expected output [x] not found in results." message the reporter wanted already exists
at
planemo/runnable.py:606; the crash just happened first.Introduced by de8b2a9 ("Download outputs only if needed in test assertions"), which added
single-output mode without guarding the trailing assignment.
Fix
Return
Nonefrom single-output mode instead of falling through to the assignment.get_outputalready caches thatNoneunder the requested id itself.The same path also covers an output the workflow does define but didn't create (optional
outputs,
output_srcfalsy) — that previously wiped the cache too.Test
tests/test_cmd_test.py::CmdTestTestCase::test_workflow_test_undefined_output— a realintegration test, runs Galaxy end to end. New fixtures
tests/data/wf20_undefined_output.yml(one
catstep, one output) and its-test.yml, which expectswf_output_1and thenwf_output_typo.Ordering matters: the undefined label has to come after a real one, so the real output is
looked up again in pass 2 against the wiped cache. With the undefined label first, the cache
is empty anyway and nothing goes wrong.
The test asserts on the JSON report rather than the exit code alone, because an uncaught
exception also exits 1 — a crash is distinguished by writing no report at all.
Red-to-green (managed Galaxy, local):
FileNotFoundError: .../tool_test_output.json— planemo died withbioblend.ConnectionErroratactivity.py:735→collect_outputs→_get_metadata→show_datasetfailure,output_problems == ["Expected output [wf_output_typo] not found in results."]Reproduced the reporter's traceback exactly before fixing — same call chain,
show_datasetwhere theirs hit
show_dataset_collection. Log showed 2 collections before shutdown, 1after, which is the wipe.
Neighbouring workflow tests re-run green:
test_workflow_test_simple_yaml,test_workflow_test_output_sanitization,test_workflow_with_identical_output_names,test_workflow_with_optional_input_output_not_provided(this one exercises thenot-created-output path through the same code). 5 passed, 1 skipped (dockerized).
black / isort / flake8 clean; mypy unchanged (pre-existing errors only).
Worth a second opinion
structured_datais computed twice per test case — once eagerly while Galaxy is up, onceagain after it has shut down — and correctness depends entirely on the cache surviving in
between. This restores that invariant rather than removing the fragility. Collecting once,
or failing loudly when a post-shutdown lookup misses the cache, would be sturdier.
🤖 Generated with Claude Code