gh-153568: Don't materialize parser token text that is never read - #153576
Merged
Merged
Conversation
maurycy
reviewed
Jul 11, 2026
Contributor
|
It seems that 1e18605, suggested by me, broke a lot! I think it means that |
pablogsal
force-pushed
the
gh-153568-token-text
branch
2 times, most recently
from
August 27, 2026 11:39
382cb01 to
f78606c
Compare
Only tokens whose text is actually consumed get a bytes object; operators and structural tokens no longer allocate one.
Co-authored-by: Maurycy Pawłowski-Wieroński <maurycy@maurycy.com>
pablogsal
force-pushed
the
gh-153568-token-text
branch
from
September 14, 2026 15:34
f78606c to
3a71cad
Compare
pablogsal
enabled auto-merge (squash)
September 14, 2026 20:27
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The parser created a bytes object for every token, but only identifiers, keywords, numbers, strings and type comments ever have their text read. Operators and structural tokens now skip the allocation entirely.
Benchmark (parsing 8 of the largest stdlib files, 1.3 MB, 20 times per run with
_PyParser_ASTFromString— parser only, no AST-to-Python conversion; pyperf, interleaved runs):