Conversation
PyUnicode_FromKindAndData(PyUnicode_4BYTE_KIND) and PyUnicodeWriter_WriteUCS4() now raise an exception if a character is not in the [U+0000; U+10ffff] range, instead of creating an invalid str object. * Add _testinternalcapi._Py_MAX_UNICODE. * Add unicode_invalid_character() helper function.
Documentation build overview
|
|
@serhiy-storchaka: Would you mind to review this change? See the issue for the rationale. The change makes the two functions a little bit slower, but also makes them safer. It should not be possible to create an invalid string in Python. In the wild, I mostly saw invalid characters when debugging CPython. For example, PyUnicode_New(size, 0x10ffff) creates a UCS-4 buffer filled with the byte pattern The other case where I saw invalid characters was on Solaris with |
|
No, I do not think it is worth to slow down this function. If you need an additional check -- use the UTF32 decoder. |
It's a little bit surprising that only 2 functions of the C API ignores invalid characters. But you have a point with performance. I wrote PR gh-158502 to document the undefined behavior, only detect invalid characters in debug mode (raise |
|
Rejected in favor of #158502. |
PyUnicode_FromKindAndData(PyUnicode_4BYTE_KIND) and PyUnicodeWriter_WriteUCS4() now raise an exception if a character is not in the [U+0000; U+10ffff] range, instead of creating an invalid str object.