fix(metrics): handle period-only text in ROUGE evaluation - #10089
Open
Excelius-Wang wants to merge 1 commit into
Open
fix(metrics): handle period-only text in ROUGE evaluation#10089Excelius-Wang wants to merge 1 commit into
Excelius-Wang wants to merge 1 commit into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
NLG evaluation can abort when a generated response contains only periods. Such responses pass the existing nonempty-token check, but rouge removes periods as sentence delimiters and raises
ValueError: Hypothesis is empty.A period-only reference can similarly raiseReference is empty.Give ROUGE metrics a zero contribution when either joined token sequence consists only of periods. Keep these samples in the existing MeanMetric averages and preserve the original tokens for BLEU. Ordinary text still uses the existing Rouge implementation.
Validation: eight CPU regression tests pass, including mixed-batch weighting, period-only references, BLEU preservation, existing whitespace/period behavior and unrelated-error propagation. A separate probe with real jieba/rouge/nltk also verifies the serialized NlgMetrics entry point and reproduces the upstream failures. An additional 307-pair public-text compatibility check preserves 304 existing results and fixes failures on three original SQuAD period annotations. These are text/annotation comparisons, not model-generated benchmark predictions. Applicable pre-commit hooks pass. No full model training run was performed.