Skip to content

Tests don't handle non-UTF-8 locales well #7217

Description

@aitap

This grew out of a response to a comment.

We have approximately 24 locale-dependent tests, and they do slightly different things:

  • Some of them (e.g. 1590.12) are there to test behaviour in a UTF-8 environment.
  • Others (like 1590.14) assume that a non-UTF-8 locale uses the Windows-1252 encoding; I started a branch to address that.
  • 2253.* try to establish a locale in which the tests should pass.
    • Strictly speaking, it's not a UTF-8 code page that the Japanese tests depend on, it's identical(enc2native('\u4E00'), '\u4E00'); they should also pass on a pre-UCRT R version on Japanese Windows, or GNU/Linux with ja_JP.EUC-JP).
  • 2266.* instead skips the tests, without trying to set a UTF-8 (or Latin-1/Latin-3/Latin-4/ISO-8859-15) locale.

Maybe we should be skipping tests more proactively?

  • Want to test behaviour when the system encoding is UTF-8? Skip if !l10n_info()$`UTF-8`.
  • Want to test printing of non-ASCII characters? Skip if !identical(enc2native('\uXXXX'), '\uXXXX').
    • Or spend extra effort to decide what should the code be doing with non-representable characters (truncate enc2native(toprint)?), then enc2native(output) before comparing it with capture.output().
  • Want to test behaviour on Windows-1252 specifically? Skip if !identical(l10n_info()$codepage, 1252L).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    encodingissues related to Encodingtests

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions