Repository navigation
Conversation
11eac6e to
0b26e33
Compare
|
What about lists with mixed elements? and lists of very small sizes? and non-matches? |
|
My comment: #159092 (comment) |
Thanks for pointing this out! I’ve added benchmarks for those. The results show no measurable regression for mixed-element lists or custom objects. It improves performance for small lists of integers as well as full scans that don’t find a match. There is a small regression (around 1–2%) for some types that don’t use the fast path, such as tuple, bool, and bytes. This is likely because the three additional type checks performed for each element when none of the fast paths applies. We could potentially reduce this overhead by checking the target type before entering the loop and selecting the appropriate comparison path upfront, though mixed-element lists would still need to handle different element types. |
The reason why I did not use specialization is that I think the key difference is that the overhead here is in the comparison for each element, rather than the bytecode-level dispatch.
|
Use specialized comparisons for exact int, float and str objects to avoid unnecessary rich-comparison dispatch in list operations.
0b26e33 to
8dfd0f0
Compare
|
|
I am not against the change but this adds some maintennce burden and complicate the code. The speedupa are attractive enough but I wonder whether the operations are commom enough though. Finding a string in a list of string is relevant. Likewise for ints. But floats... not really IMO. |
Sorry for the force push. I’m running the benchmarks on a GIL-enabled build with PGO/LTO and will update the results once they’re ready. |
I’m also considering adding tests and assertions to verify that the fast paths preserve the semantics of I agree that finding strings and ints in lists is a more compelling use case than finding floats. I kept the float fast path because it gives the largest speedup in the microbenchmarks, and the implementation is small. However, I think it’s reasonable to leave floats out for now and keep only the int and str fast paths. |
List scans (
in,index,count,remove, and element comparisons inlist.__eq__()) callPyObject_RichCompareBool(). For exact builtin type pairs such asint,floatorstr, we can avoid the generic rich-comparison dispatch and use a specialized comparison instead, following the approach already used bylist.sort().Benchmarks
All numbers below are the best of 7 repeated runs (except for pyperformance cases) on two builds of the same tree differing only in
Objects/listobject.c, using a free-threading build (--disable-gil, Clang 17,-O3, macOS arm64). Lists contain 20,000 elements unless stated otherwise. “Miss” means the value is not in the list, so the entire list is scanned.Full scans and other entry points
Microbenchmarks report ns/element for 20,000-element lists:
x in list[int], misslist.index(x), misslist.count(x), missx in list[str], misslist.remove(x), match at end[int] == [int][str] == [str]The corresponding total times for full scans were:
Small lists and mixed/custom-element lists
[1, 'a', 2.0, None, (1, 2)] * 4000, miss__eq__, miss__eq__, misspyperformance
bm_hexiom: approximately 1.10× fasterbm_meteor_contest,bm_go,bm_deltablue, andbm_comprehensions: no significant difference.Correctness
The fast path preserves the identity shortcut of
PyObject_RichCompareBool(), including NaN identity semantics. Other type combinations retain the existing comparison path, preserving rich-comparison dispatch and recursive-comparison protection.Tested with
test_list,test_sort,test_operatorandtest_compare, plus differential checks covering subclasses, NaNs, custom__eq__implementations, large integers, string representations and mutation during comparison. The existingtest_deopt_from_append_listenvironment failure also occurs on unmodified main.list.__contains__/index/count/remove#159092