Skip to content

gh-159071: Correct UnicodeDecodeError offsets in utf-8-sig codec when input starts with BOM - #159081

Open
devpryan7792 wants to merge 1 commit into
python:mainfrom
devpryan7792:fix/gh-159071-utf-8-sig-offsets
Open

devpryan7792 wants to merge 1 commit into
python:mainfrom
devpryan7792:fix/gh-159071-utf-8-sig-offsets

Conversation

@devpryan7792

@devpryan7792 devpryan7792 commented Oct 9, 2026 •

Copy link
Copy Markdown

When the input starts with a UTF-8 BOM (codecs.BOM_UTF8), encodings.utf_8_sig was slicing the input (input = input[3:]) before calling codecs.utf_8_decode. As a result, any raised UnicodeDecodeError reported offsets and object attributes relative to the sliced input without the BOM (i.e. start was off by 3, and object omitted the initial BOM).

This aligns utf-8-sig with other BOM-stripping codecs (utf-16 and utf-32), which preserve the BOM in the exception's object, start, and end attributes:

  1. Updated decode(), IncrementalDecoder._buffer_decode(), and StreamReader.decode() in Lib/encodings/utf_8_sig.py to pass the original input to codecs.utf_8_decode and strip the leading \ufeff from the decoded output.
  2. Added unit test test_decode_error_offsets_with_bom to UTF8SigTest in Lib/test/test_codecs.py.
  3. Added NEWS blurb entry.

Fixes #159071.

…c when input starts with BOM

Signed-off-by: Pradyumn jha <devpryan7792@users.noreply.github.com>
@python-cla-bot

python-cla-bot Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

All commit authors signed the Contributor License Agreement.

CLA signed

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

utf-8-sig: UnicodeDecodeError offsets do not count the BOM

1 participant