Skip to content

gh-159092: Optimize comparisons in list operations - #159093

Open
hetaozdh wants to merge 1 commit into
python:mainfrom
hetaozdh:list-fast-eq
Open

hetaozdh wants to merge 1 commit into
python:mainfrom
hetaozdh:list-fast-eq

Conversation

@hetaozdh

@hetaozdh hetaozdh commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

List scans (in, index, count, remove, and element comparisons in list.__eq__()) call PyObject_RichCompareBool(). For exact builtin type pairs such as int, float or str, we can avoid the generic rich-comparison dispatch and use a specialized comparison instead, following the approach already used by list.sort().

Benchmarks

All numbers below are the best of 7 repeated runs (except for pyperformance cases) on two builds of the same tree differing only in Objects/listobject.c, using a free-threading build (--disable-gil, Clang 17, -O3, macOS arm64). Lists contain 20,000 elements unless stated otherwise. “Miss” means the value is not in the list, so the entire list is scanned.

Full scans and other entry points

Microbenchmarks report ns/element for 20,000-element lists:

Operation Main PR Speedup
x in list[int], miss 9.11 3.21 2.84×
list.index(x), miss 9.46 3.52 2.69×
list.count(x), miss 9.67 3.65 2.65×
x in list[str], miss 8.76 4.70 1.86×
list.remove(x), match at end 8.36 2.09 4.00×
[int] == [int] 9.20 3.78 2.43×
[str] == [str] 10.52 8.45 1.24×
1,000 custom objects, miss 65.96 65.24 ~1.0×

The corresponding total times for full scans were:

Workload Main PR Speedup
20k ints, miss 181.8 µs 63.8 µs 2.85×
20k ints, hit at end 182.0 µs 63.7 µs 2.86×
20k floats, miss 192.4 µs 58.3 µs 3.30×
20k strs, miss 150.0 µs 77.6 µs 1.93×
20k custom objects, miss 1.351 ms 1.344 ms ~1.0×

Small lists and mixed/custom-element lists

Workload Main PR Result
1 int, miss 40.1 ns 34.8 ns 1.15× faster
1 int, hit 33.0 ns 33.5 ns ~1.0×
3 ints, miss 58.4 ns 39.9 ns 1.46× faster
3 ints, hit at end 50.8 ns 39.0 ns 1.30× faster
1 custom object, miss 122.1 ns 121.9 ns ~1.0×
3 custom objects, miss 289.3 ns 286.2 ns ~1.0×
Mixed list [1, 'a', 2.0, None, (1, 2)] * 4000, miss 237.9 µs 237.9 µs ~1.0×
Same mixed list, hit on second element 39.8 ns 40.0 ns ~1.0×
20k objects with Python __eq__, miss 1.351 ms 1.344 ms ~1.0×
3 objects with Python __eq__, miss 122.1 ns 121.9 ns ~1.0×

pyperformance

  • bm_hexiom: approximately 1.10× faster
  • bm_meteor_contest, bm_go, bm_deltablue, and bm_comprehensions: no significant difference.

Correctness

The fast path preserves the identity shortcut of PyObject_RichCompareBool(), including NaN identity semantics. Other type combinations retain the existing comparison path, preserving rich-comparison dispatch and recursive-comparison protection.

Tested with test_list, test_sort, test_operator and test_compare, plus differential checks covering subclasses, NaNs, custom __eq__ implementations, large integers, string representations and mutation during comparison. The existing test_deopt_from_append_list environment failure also occurs on unmodified main.

@bedevere-app bedevere-app Bot added the type-feature A feature request or enhancement label Oct 9, 2026
@hetaozdh
hetaozdh force-pushed the list-fast-eq branch 3 times, most recently from 11eac6e to 0b26e33 Compare October 9, 2026 23:08
@picnixz

picnixz commented Oct 10, 2026

Copy link
Copy Markdown
Member

What about lists with mixed elements? and lists of very small sizes? and non-matches?

@corona10

Copy link
Copy Markdown
Member

My comment: #159092 (comment)

@hetaozdh

Copy link
Copy Markdown
Contributor Author

What about lists with mixed elements? and lists of very small sizes? and non-matches?

Thanks for pointing this out! I’ve added benchmarks for those. The results show no measurable regression for mixed-element lists or custom objects. It improves performance for small lists of integers as well as full scans that don’t find a match.

There is a small regression (around 1–2%) for some types that don’t use the fast path, such as tuple, bool, and bytes. This is likely because the three additional type checks performed for each element when none of the fast paths applies. We could potentially reduce this overhead by checking the target type before entering the loop and selecting the appropriate comparison path upfront, though mixed-element lists would still need to handle different element types.

@hetaozdh

Copy link
Copy Markdown
Contributor Author

My comment: #159092 (comment)

The reason why I did not use specialization is that I think the key difference is that the overhead here is in the comparison for each element, rather than the bytecode-level dispatch.

list.index(), count(), and remove() are C methods. So specialization cannot optimize their inner loops. list.__contains__() has the same issue, and only list == list goes through COMPARE_OP. So specialization cannot cover all entry points.

PyObject_RichCompareBool() is called directly from the list's C-level loop, so the per-element comparison does not go through specialization. Specializing CONTAINS_OP could reduce some per-call dispatch overhead, but the cost addressed here is incurred for every element scanned. These are different levels of overhead, and I don't think specialization is the right place for this optimization.

Use specialized comparisons for exact int, float and str objects to avoid unnecessary rich-comparison dispatch in list operations.
@picnixz

picnixz commented Oct 10, 2026

Copy link
Copy Markdown
Member
  1. Please do not force push, we cannot review what changed otherwise.
  2. What about non-gil builds with PGO/LTO? your benchmarks are only for FT AFAIU.

@picnixz

picnixz commented Oct 10, 2026

Copy link
Copy Markdown
Member

I am not against the change but this adds some maintennce burden and complicate the code. The speedupa are attractive enough but I wonder whether the operations are commom enough though.

Finding a string in a list of string is relevant. Likewise for ints. But floats... not really IMO.

@hetaozdh

Copy link
Copy Markdown
Contributor Author
  1. Please do not force push, we cannot review what changed otherwise.
  2. What about non-gil builds with PGO/LTO? your benchmarks are only for FT AFAIU.

Sorry for the force push. I’m running the benchmarks on a GIL-enabled build with PGO/LTO and will update the results once they’re ready.

@hetaozdh

Copy link
Copy Markdown
Contributor Author

I am not against the change but this adds some maintennce burden and complicate the code. The speedupa are attractive enough but I wonder whether the operations are commom enough though.

Finding a string in a list of string is relevant. Likewise for ints. But floats... not really IMO.

I’m also considering adding tests and assertions to verify that the fast paths preserve the semantics of PyObject_RichCompareBool(), which may help with the maintenance burden.

I agree that finding strings and ints in lists is a more compelling use case than finding floats. I kept the float fast path because it gives the largest speedup in the microbenchmarks, and the implementation is small. However, I think it’s reasonable to leave floats out for now and keep only the int and str fast paths.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting review type-feature A feature request or enhancement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants