gh-151464: exclude '<>' token from tokenize output - #154854
Conversation
Was: ``` $ echo '1 <> 2' | python -m tokenize 1,0-1,1: NUMBER '1' 1,2-1,4: OP '<>' 1,5-1,6: NUMBER '2' 1,6-1,7: NEWLINE '\n' 2,0-2,0: ENDMARKER '' ``` Now (regardless on ``__future__.barry_as_FLUFL`` import): ``` $ echo '1 <> 2' | ./python -m tokenize 1,0-1,1: NUMBER '1' 1,2-1,3: OP '<' 1,3-1,4: OP '>' 1,5-1,6: NUMBER '2' 1,6-1,7: NEWLINE '\n' 2,0-2,0: ENDMARKER '' ``` in accordance with the Grammar: https://docs.python.org/3.14/reference/lexical_analysis.html#operators-and-delimiters Also adds a custom error message for ``<>`` ("not equal" in Pascal and Python 2).
Aniketsy
left a comment
There was a problem hiding this comment.
Thanks for this work :)
I've just gone through these changes and changes looks good as per my understanding ofcourse i'm not expert 😀 , and it was really interesting to going through these changes, at first I found its bit tricky to understand some portion.
I'm excited to see the review process in this from experts and try to understand how it goes, also if you feel this comment as noise please feel free to mark as off-topic.
|
A Python core developer has requested some changes be made to your pull request before we can consider merging it. If you could please address their requests along with any other requests in other reviews from core developers that would be appreciated. Once you have made the requested changes, please leave a comment on this pull request containing the phrase |
I have made the requested changes; please review again |
|
Thanks for making the requested changes! @pablogsal: please review the changes made to this pull request. |
Documentation build overview
398 files changed ·
|
|
@pablogsal, does it make sense for you at all? I doubt I can reduce patch further. |
# Conflicts: # Grammar/python.gram # Parser/parser.c # Parser/pegen.c
|
Thanks @skirpichev for the PR, and @pablogsal for merging it 🌮🎉.. I'm working now to backport this PR to: 3.13, 3.14, 3.15. |
|
Sorry, @skirpichev and @pablogsal, I could not cleanly backport this to |
|
Sorry, @skirpichev and @pablogsal, I could not cleanly backport this to |
|
Sorry, @skirpichev and @pablogsal, I could not cleanly backport this to |
|
GH-158189 is a backport of this pull request to the 3.15 branch. |
|
GH-158190 is a backport of this pull request to the 3.14 branch. |
|
GH-158191 is a backport of this pull request to the 3.13 branch. |
…onGH-154854) * pythongh-151464: exclude '<>' token from tokenize output Was: ``` $ echo '1 <> 2' | python -m tokenize 1,0-1,1: NUMBER '1' 1,2-1,4: OP '<>' 1,5-1,6: NUMBER '2' 1,6-1,7: NEWLINE '\n' 2,0-2,0: ENDMARKER '' ``` Now (regardless on ``__future__.barry_as_FLUFL`` import): ``` $ echo '1 <> 2' | ./python -m tokenize 1,0-1,1: NUMBER '1' 1,2-1,3: OP '<' 1,3-1,4: OP '>' 1,5-1,6: NUMBER '2' 1,6-1,7: NEWLINE '\n' 2,0-2,0: ENDMARKER '' ``` in accordance with the Grammar: https://docs.python.org/3.14/reference/lexical_analysis.html#operators-and-delimiters Also adds a custom error message for ``<>`` ("not equal" in Pascal and Python 2). * +1 * address review: lowercase and move invalid rule * address review: news * address review: move test_guido_as_bdfl_ineq_tokens() * address review: revert _PyTokenizer_From* changes * + revert unrelated change * address review: remove whatsnew entry (cherry picked from commit 198bc76)
…onGH-154854) * pythongh-151464: exclude '<>' token from tokenize output Was: ``` $ echo '1 <> 2' | python -m tokenize 1,0-1,1: NUMBER '1' 1,2-1,4: OP '<>' 1,5-1,6: NUMBER '2' 1,6-1,7: NEWLINE '\n' 2,0-2,0: ENDMARKER '' ``` Now (regardless on ``__future__.barry_as_FLUFL`` import): ``` $ echo '1 <> 2' | ./python -m tokenize 1,0-1,1: NUMBER '1' 1,2-1,3: OP '<' 1,3-1,4: OP '>' 1,5-1,6: NUMBER '2' 1,6-1,7: NEWLINE '\n' 2,0-2,0: ENDMARKER '' ``` in accordance with the Grammar: https://docs.python.org/3.14/reference/lexical_analysis.html#operators-and-delimiters Also adds a custom error message for ``<>`` ("not equal" in Pascal and Python 2). * +1 * address review: lowercase and move invalid rule * address review: news * address review: move test_guido_as_bdfl_ineq_tokens() * address review: revert _PyTokenizer_From* changes * + revert unrelated change * address review: remove whatsnew entry (cherry picked from commit 198bc76)
…onGH-154854) * pythongh-151464: exclude '<>' token from tokenize output Was: ``` $ echo '1 <> 2' | python -m tokenize 1,0-1,1: NUMBER '1' 1,2-1,4: OP '<>' 1,5-1,6: NUMBER '2' 1,6-1,7: NEWLINE '\n' 2,0-2,0: ENDMARKER '' ``` Now (regardless on ``__future__.barry_as_FLUFL`` import): ``` $ echo '1 <> 2' | ./python -m tokenize 1,0-1,1: NUMBER '1' 1,2-1,3: OP '<' 1,3-1,4: OP '>' 1,5-1,6: NUMBER '2' 1,6-1,7: NEWLINE '\n' 2,0-2,0: ENDMARKER '' ``` in accordance with the Grammar: https://docs.python.org/3.14/reference/lexical_analysis.html#operators-and-delimiters Also adds a custom error message for ``<>`` ("not equal" in Pascal and Python 2). * +1 * address review: lowercase and move invalid rule * address review: news * address review: move test_guido_as_bdfl_ineq_tokens() * address review: revert _PyTokenizer_From* changes * + revert unrelated change * address review: remove whatsnew entry (cherry picked from commit 198bc76)
…#158191) * gh-151464: exclude '<>' token from tokenize output Was: ``` $ echo '1 <> 2' | python -m tokenize 1,0-1,1: NUMBER '1' 1,2-1,4: OP '<>' 1,5-1,6: NUMBER '2' 1,6-1,7: NEWLINE '\n' 2,0-2,0: ENDMARKER '' ``` Now (regardless on ``__future__.barry_as_FLUFL`` import): ``` $ echo '1 <> 2' | ./python -m tokenize 1,0-1,1: NUMBER '1' 1,2-1,3: OP '<' 1,3-1,4: OP '>' 1,5-1,6: NUMBER '2' 1,6-1,7: NEWLINE '\n' 2,0-2,0: ENDMARKER '' ``` in accordance with the Grammar: https://docs.python.org/3.14/reference/lexical_analysis.html#operators-and-delimiters Also adds a custom error message for ``<>`` ("not equal" in Pascal and Python 2). * +1 * address review: lowercase and move invalid rule * address review: news * address review: move test_guido_as_bdfl_ineq_tokens() * address review: revert _PyTokenizer_From* changes * + revert unrelated change * address review: remove whatsnew entry (cherry picked from commit 198bc76) Co-authored-by: Sergey B Kirpichev <skirpichev@gmail.com>
…#158189) * gh-151464: exclude '<>' token from tokenize output Was: ``` $ echo '1 <> 2' | python -m tokenize 1,0-1,1: NUMBER '1' 1,2-1,4: OP '<>' 1,5-1,6: NUMBER '2' 1,6-1,7: NEWLINE '\n' 2,0-2,0: ENDMARKER '' ``` Now (regardless on ``__future__.barry_as_FLUFL`` import): ``` $ echo '1 <> 2' | ./python -m tokenize 1,0-1,1: NUMBER '1' 1,2-1,3: OP '<' 1,3-1,4: OP '>' 1,5-1,6: NUMBER '2' 1,6-1,7: NEWLINE '\n' 2,0-2,0: ENDMARKER '' ``` in accordance with the Grammar: https://docs.python.org/3.14/reference/lexical_analysis.html#operators-and-delimiters Also adds a custom error message for ``<>`` ("not equal" in Pascal and Python 2). * +1 * address review: lowercase and move invalid rule * address review: news * address review: move test_guido_as_bdfl_ineq_tokens() * address review: revert _PyTokenizer_From* changes * + revert unrelated change * address review: remove whatsnew entry (cherry picked from commit 198bc76) Co-authored-by: Sergey B Kirpichev <skirpichev@gmail.com>
…onGH-154854) * pythongh-151464: exclude '<>' token from tokenize output Was: ``` $ echo '1 <> 2' | python -m tokenize 1,0-1,1: NUMBER '1' 1,2-1,4: OP '<>' 1,5-1,6: NUMBER '2' 1,6-1,7: NEWLINE '\n' 2,0-2,0: ENDMARKER '' ``` Now (regardless on ``__future__.barry_as_FLUFL`` import): ``` $ echo '1 <> 2' | ./python -m tokenize 1,0-1,1: NUMBER '1' 1,2-1,3: OP '<' 1,3-1,4: OP '>' 1,5-1,6: NUMBER '2' 1,6-1,7: NEWLINE '\n' 2,0-2,0: ENDMARKER '' ``` in accordance with the Grammar: https://docs.python.org/3.14/reference/lexical_analysis.html#operators-and-delimiters Also adds a custom error message for ``<>`` ("not equal" in Pascal and Python 2). * +1 * address review: lowercase and move invalid rule * address review: news * address review: move test_guido_as_bdfl_ineq_tokens() * address review: revert _PyTokenizer_From* changes * + revert unrelated change * address review: remove whatsnew entry (cherry picked from commit 198bc76)
…#158190) * gh-151464: exclude '<>' token from tokenize output Was: ``` $ echo '1 <> 2' | python -m tokenize 1,0-1,1: NUMBER '1' 1,2-1,4: OP '<>' 1,5-1,6: NUMBER '2' 1,6-1,7: NEWLINE '\n' 2,0-2,0: ENDMARKER '' ``` Now (regardless on ``__future__.barry_as_FLUFL`` import): ``` $ echo '1 <> 2' | ./python -m tokenize 1,0-1,1: NUMBER '1' 1,2-1,3: OP '<' 1,3-1,4: OP '>' 1,5-1,6: NUMBER '2' 1,6-1,7: NEWLINE '\n' 2,0-2,0: ENDMARKER '' ``` in accordance with the Grammar: https://docs.python.org/3.14/reference/lexical_analysis.html#operators-and-delimiters Also adds a custom error message for ``<>`` ("not equal" in Pascal and Python 2). * +1 * address review: lowercase and move invalid rule * address review: news * address review: move test_guido_as_bdfl_ineq_tokens() * address review: revert _PyTokenizer_From* changes * + revert unrelated change * address review: remove whatsnew entry (cherry picked from commit 198bc76) Co-authored-by: Sergey B Kirpichev <skirpichev@gmail.com>
|
Great work on the tokenizer fix, @skirpichev! Thanks a lot! ❤️ |
Was:
Now (regardless on
__future__.barry_as_FLUFLimport):in accordance with the Grammar:
https://docs.python.org/3.14/reference/lexical_analysis.html#operators-and-delimiters
Also adds a custom error message for
<>("not equal" in Pascal and Python 2).