Skip to content

gh-158445: Reject invalid UCS4 in PyUnicode_FromKindAndData() - #158447

Closed
vstinner wants to merge 1 commit into
python:mainfrom
vstinner:from_ucs4
Closed

vstinner wants to merge 1 commit into
python:mainfrom
vstinner:from_ucs4

Conversation

@vstinner

@vstinner vstinner commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

PyUnicode_FromKindAndData(PyUnicode_4BYTE_KIND) and PyUnicodeWriter_WriteUCS4() now raise an exception if a character is not in the [U+0000; U+10ffff] range, instead of creating an invalid str object.

  • Add _testinternalcapi._Py_MAX_UNICODE.
  • Add unicode_invalid_character() helper function.

PyUnicode_FromKindAndData(PyUnicode_4BYTE_KIND) and
PyUnicodeWriter_WriteUCS4() now raise an exception if a character is
not in the [U+0000; U+10ffff] range, instead of creating an invalid
str object.

* Add _testinternalcapi._Py_MAX_UNICODE.
* Add unicode_invalid_character() helper function.
@read-the-docs-community

Copy link
Copy Markdown

Documentation build overview

📚 cpython-previews | 🛠️ Build #34835623 | 📁 Comparing da5cb33 against main (115c297)

  🔍 Preview build  

2 files changed
± c-api/unicode.html
± whatsnew/changelog.html

@vstinner

Copy link
Copy Markdown
Member Author

@serhiy-storchaka: Would you mind to review this change? See the issue for the rationale.

The change makes the two functions a little bit slower, but also makes them safer. It should not be possible to create an invalid string in Python.

In the wild, I mostly saw invalid characters when debugging CPython. For example, PyUnicode_New(size, 0x10ffff) creates a UCS-4 buffer filled with the byte pattern 0xff which creates invalid characters \Uffffffff on purpose: to detect usage of uninitialized characters.

The other case where I saw invalid characters was on Solaris with wchar_t* strings (Py_UCS4 strings in practice). The _Py_DecodeNonUnicodeWchar() function was added to fix these characters.

@serhiy-storchaka

Copy link
Copy Markdown
Member

No, I do not think it is worth to slow down this function. If you need an additional check -- use the UTF32 decoder.

@vstinner

Copy link
Copy Markdown
Member Author

@serhiy-storchaka:

No, I do not think it is worth to slow down this function. If you need an additional check -- use the UTF32 decoder.

It's a little bit surprising that only 2 functions of the C API ignores invalid characters. But you have a point with performance.

I wrote PR gh-158502 to document the undefined behavior, only detect invalid characters in debug mode (raise SystemError), and add tests on the behavior in release and debug mode.

@vstinner

Copy link
Copy Markdown
Member Author

Rejected in favor of #158502.

@vstinner vstinner closed this Sep 30, 2026
@vstinner
vstinner deleted the from_ucs4 branch September 30, 2026 14:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants