Repository navigation
Include column offsets for bytecode instructions #88116
Description
Activity
If we could include column offsets from the AST nodes for every bytecode instructions we could leverage these to offer much better tracebacks and a lot more information to debuggers and similar tools. For instance, in an expression such as:
z['aaaaaa']['bbbbbb']['cccccc']['dddddd].sddfsdf.sdfsdf
where one of these elements is None, we could tell exactly what of these pieces is the actual None and we could do some cool highlighting on tracebacks.
Similarly, coverage tools and debuggers could also make this distinction and therefore could offer reports with more granularity.
The cost is not 0: it would be two integers per bytecode instruction, but I think it may be worth the effort.
I'm going to prepare a PEP since the discussion regarding if the two integers per bytecode are worth enough is going to be eternal.
The additional cost will not only be the line number table, but we need to store the line for exceptions that are reraised after cleanup.
Adding a column will mean more stack consumption.Yup, but I still think is worth the cost, giving that debugging improvements are usually extremely popular among users.
Specific examples of current messages and proposed improvements would help focus discussion.
If you are willing to only handle code lines up to 256 chars, only 2 bytes should be needed. (0,0) or (255,255) could mean 'somewhere beyond the 256th char'.
Specific examples of current messages and proposed improvements would help focus discussion.
Yeah, I am proposing going from:
>>> x['aaaaaa']['bbbbbb']['cccccc']['dddddd'].sddfsdf.sdfsdf Traceback (most recent call last): File "<stdin>", line 1, in <module> TypeError: 'NoneType' object is not subscriptable
to
>>> x['aaaaaa']['bbbbbb']['cccccc']['dddddd'].sddfsdf.sdfsdf ^^^^^^^^^^ Traceback (most recent call last): File "<stdin>", line 1, in <module> TypeError: 'NoneType' object is not subscriptable
Basically, to highlight in all exceptions the range in the displayed line where the error ocurred. For instance:
>>> foo(a, b/z+2, c, 132432 /x, d /y) ^^^^^^^^^ Traceback (most recent call last): File "<stdin>", line 1, in <module> ZeroDivisionError: division by zero
Marking what expression evaluated to None would be extremely helpful. So would marking the 0 denominator when there is more than one candidate: "e = a/b + c/d". It should be easy to revise IDLE Shell's print_exception to tag the span. In some cases, the code that goes from a traceback line to a line in a file and marks it could do the same.
What would you do when the expression is not the last line?
try:
x/y
...
except Exception as e:
...
raise eThe except and raise might even be in a separate module.
I look forward to the PEP and discussion.So would marking the 0 denominator when there is more than one candidate: "e = a/b + c/d".
No, it will mark the offset of the bytecode that was getting executed when the exception was raised. Is just a way to mark what is raising the particular exception
What would you do when the expression is not the last line?
There is some logic needed to re-raise exceptions. But basically it boils down to comparate the line numbers to decide if you want to propagate or not the offsets.
35 remaining items
As awkward of a beast as a tuple subclass emulating the previous fixed size namedtuple while adding additional attributes seems, it is very practical.
Users don't care about the specifics. They just need existing code to keep working and will want access to new values when updating this code or writing new code. Which means it should still be a tuple subclass for typing reasons known as
inspect.Tracebackandinspect.FrameInfo, support the existing indices and unpacking semantics, and support named field access. That this would be adding additional fields that indexing does not provide is a mere curiosity that most users would not notice as modern code should be field based anyways.Nobody should care if the implementation behind the returned instances is unusual. The behavior will be what they expect. It avoids the need to expand into multiple different return types via an arg flag or the creation of parallel APIs.
This is a good practical example of why we should not design APIs to return tuples if they could conceivably change in the future. And why people unpacking on API calls that return tuples might find it wise to always
[:slice]the return value as a style idiom.
All this said... we could also just add another field to the namedtuple. Some code inspection among existing API users is required, but I expect most users either immediately unpack upon return value assignment or use field names. Mix and matching of indexing and named access on a single return value is probably rare. The easy modification to existing code using unpacking to deal with this kind of API change upon version upgrade is to append a
[:5]or similar to the API call. We have had other tuple returning APIs add fields in the past where this was the workaround though I've forgotten what they were off the top of my head.This would be the more disruptive option. It is nice to avoid this if possible as it delays people's ability to update to and even test their code on 3.11. So I still lean towards the hybrid "half namedtuple" class.
Reacted by Jelle Zijlstra- added a commit that references this issue
on Apr 23, 2022 @markshannon I'm noticing that the implementation for varints and signed varints that we use has two encodings for 0 (both uval==0 and uval==1 decode to 0), and no way to encode
-2**31. Is this a problem? What prevents us to need to encode-2**31?Also, weirdly there are two ways to encode 0 and no ways to encode -INT_MIN. @markshannon is this expected?
- added a commit that references this issue
on Nov 13, 2022 - added a commit that references this issue
on Oct 26, 2023
Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.
Show more details
GitHub fields:
bugs.python.org fields: