Skip to content

CSV sniffing detects wrong field delimiter #97611

Description

@Dunedan

Bug report

csv.Sniffer().sniff() detects the wrong field delimiter if the possible valid delimiters contain \t and the provided data contains one line starting with the combination of double quotes followed by a tab and has the same combination of double quotes and a tab at least once more afterwards in the data. In that case the delimiter is always reported as \t instead of the correct one.

Configuring the preferred separator as suggested in #80678 (comment) doesn't have any effect on this behavior.

Here is an example to reproduce it:

import csv
from pprint import pprint

data = """
Surname;First Name;Year of birth
"\tDoe;Jane"\t;1971
"""

sniffer = csv.Sniffer()
dialect = sniffer.sniff(data, delimiters=",;\t")
pprint(dict(dialect.__dict__))

Expected output:

{'__doc__': None,
 '__module__': 'csv',
 '_name': 'sniffed',
 'delimiter': ';',
 'doublequote': False,
 'lineterminator': '\r\n',
 'quotechar': '"',
 'quoting': 0,
 'skipinitialspace': False}

Actual output:

{'__doc__': None,
 '__module__': 'csv',
 '_name': 'sniffed',
 'delimiter': '\t',
 'doublequote': False,
 'lineterminator': '\r\n',
 'quotechar': '"',
 'quoting': 0,
 'skipinitialspace': False}

Your environment

  • CPython versions tested on: 3.7.14, 3.8.14, 3.9.14, 3.10.7
  • Operating system and architecture: Debian/sid on amd64

Metadata

Metadata

Assignees

Labels

3.13bugs and security fixes3.14bugs and security fixes3.15pre-release feature fixes, bugs and security fixesstdlibStandard Library Python modules in the Lib/ directorytype-bugAn unexpected behavior, bug, or error

Projects

Status
Done

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions