Bug report
When the input starts with a BOM, the utf-8-sig codec reports UnicodeDecodeError offsets relative to the input without the BOM, i.e. error.start is off-by-3:
data = b"\xef\xbb\xbf\xff" # Invalid byte at position 3
try:
data.decode("utf-8-sig")
except UnicodeDecodeError as e:
print(e) # 'utf-8' codec can't decode byte 0xff in position 0: invalid start byte
The input is being sliced in utf_8_sig.py to remove the BOM, but raised exceptions are not modified to correct the offsets and object attributes. This occurs in a few places, e.g.
|
def decode(input, errors='strict'): |
|
prefix = 0 |
|
if input[:3] == codecs.BOM_UTF8: |
|
input = input[3:] |
|
prefix = 3 |
|
(output, consumed) = codecs.utf_8_decode(input, errors, True) |
|
return (output, consumed+prefix) |
The utf-16 and utf-32 codecs also remove the BOM, but it is included in the offsets and object attributes in their errors, so this looks like an oversight.
CPython versions tested on:
3.14, 3.15rc2
Operating systems tested on:
Linux
Linked PRs
Bug report
When the input starts with a BOM, the utf-8-sig codec reports
UnicodeDecodeErroroffsets relative to the input without the BOM, i.e.error.startis off-by-3:The input is being sliced in
utf_8_sig.pyto remove the BOM, but raised exceptions are not modified to correct the offsets and object attributes. This occurs in a few places, e.g.cpython/Lib/encodings/utf_8_sig.py
Lines 18 to 24 in 5a22a62
The utf-16 and utf-32 codecs also remove the BOM, but it is included in the offsets and object attributes in their errors, so this looks like an oversight.
CPython versions tested on:
3.14, 3.15rc2
Operating systems tested on:
Linux
Linked PRs