gh-158451: Fix python-gdb.py for invalid Unicode strings - #158475
Conversation
python-gdb.py is now able to format an invalid Unicode string: render invalid characters as "\Uhhhhhhhh" instead of raising an exception. The PyUnicodeWriter API fills a UCS-4 buffer with 0xff byte pattern to detect usage of uninitialized characters. It produces invalid characters '\Uffffffff'. PyBytesObjectPtr now iterates on bytes (list of integers), instead of creating a temporary Unicode string. Add more bytes tests in test_gdb pretty printer. Remove the unused 'ch2' variable.
|
!buildbot AMD64 Ubuntu |
|
🤖 New build scheduled with the buildbot fleet by @vstinner for commit 1fada1c 🤖 Results will be shown at: https://buildbot.python.org/all/#/grid?branch=refs%2Fpull%2F158475%2Fmerge The command will test the builders whose names match following regular expression: The builders matched are:
|
|
!buildbot AMD64 Fedora Stable PR |
|
🤖 New build scheduled with the buildbot fleet by @vstinner for commit 1fada1c 🤖 Results will be shown at: https://buildbot.python.org/all/#/grid?branch=refs%2Fpull%2F158475%2Fmerge The command will test the builders whose names match following regular expression: The builders matched are:
|
|
Thanks @vstinner for the PR 🌮🎉.. I'm working now to backport this PR to: 3.14. |
|
Thanks @vstinner for the PR 🌮🎉.. I'm working now to backport this PR to: 3.15. |
|
GH-158495 is a backport of this pull request to the 3.14 branch. |
|
GH-158496 is a backport of this pull request to the 3.15 branch. |
…58475) (#158495) gh-158451: Fix python-gdb.py for invalid Unicode strings (GH-158475) python-gdb.py is now able to format an invalid Unicode string: render invalid characters as "\Uhhhhhhhh" instead of raising an exception. The PyUnicodeWriter API fills a UCS-4 buffer with 0xff byte pattern to detect usage of uninitialized characters. It produces invalid characters '\Uffffffff'. PyBytesObjectPtr now iterates on bytes (list of integers), instead of creating a temporary Unicode string. Add more bytes tests in test_gdb pretty printer. Remove the unused 'ch2' variable. (cherry picked from commit 7eada7c) Co-authored-by: Victor Stinner <vstinner@python.org>
python-gdb.py is now able to format an invalid Unicode string: render invalid characters as "\Uhhhhhhhh" instead of raising an exception.
The PyUnicodeWriter API fills a UCS-4 buffer with 0xff byte pattern to detect usage of uninitialized characters. It produces invalid characters '\Uffffffff'.
PyBytesObjectPtr now iterates on bytes (list of integers), instead of creating a temporary Unicode string. Add more bytes tests in test_gdb pretty printer.
Remove the unused 'ch2' variable.