Skip to content

utf-8-sig: UnicodeDecodeError offsets do not count the BOM #159071

Description

@alicederyn

Bug report

When the input starts with a BOM, the utf-8-sig codec reports UnicodeDecodeError offsets relative to the input without the BOM, i.e. error.start is off-by-3:

data = b"\xef\xbb\xbf\xff"  # Invalid byte at position 3
try:
    data.decode("utf-8-sig")
except UnicodeDecodeError as e:
    print(e)  # 'utf-8' codec can't decode byte 0xff in position 0: invalid start byte

The input is being sliced in utf_8_sig.py to remove the BOM, but raised exceptions are not modified to correct the offsets and object attributes. This occurs in a few places, e.g.

def decode(input, errors='strict'):
prefix = 0
if input[:3] == codecs.BOM_UTF8:
input = input[3:]
prefix = 3
(output, consumed) = codecs.utf_8_decode(input, errors, True)
return (output, consumed+prefix)

The utf-16 and utf-32 codecs also remove the BOM, but it is included in the offsets and object attributes in their errors, so this looks like an oversight.

CPython versions tested on:

3.14, 3.15rc2

Operating systems tested on:

Linux

Linked PRs

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    interpreter-core(Objects, Python, Grammar, and Parser dirs)type-bugAn unexpected behavior, bug, or error

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions