Skip to content

Free-threaded 3.15: a functools.partial shared between threads no longer scales (it slows down as you add threads) #159157

Description

@mattsta

Bug report

I was migrating some of my "massively threaded" 3.14t code to 3.15t and my performance agent saw everything collapse under concurrency after the python version change.

Turns out a new 3.15-only fix (?) for a partial application is now over-locking everything. For my own use I just migrated away from using functools.partial everywhere now, but ... FYI i guess.

agent report:

=========================

Free-threaded 3.15: a functools.partial shared between threads no longer scales (it slows down as you add threads)

Summary

On the free-threaded build, CPython 3.15 calls to a functools.partial stop
scaling across threads when the partial (or just its target function) is shared
between threads. On a 10-core Apple M1 Max, 24 threads calling one module-level
partial(f, 1) reach 0.20x the single-thread rate, so total throughput drops
when threads are added. 3.14.3t reaches 6.37x on the same machine. The
slowdown shows up at 2 threads (0.34x). The cause is the use-after-free fix for
gh-154189 (PR gh-154508, 3.15 backport gh-154652, first shipped in 3.15.0rc1). It
made partial_vectorcall take and drop strong references to pto->fn,
pto->args and pto->kw on every call. In the free-threaded build an
incref/decref from a thread that does not own the object is an atomic
read-modify-write on that object's ob_ref_shared. Every calling thread now writes
the same three cache lines (the target function, the bound-args tuple, and the
keywords dict, which partial allocates even when empty) six times per call. In
3.14 the call path takes no references on these objects: it reads them borrowed,
and the callee frame uses stackrefs that skip deferred-refcount and immortal
objects. Giving each thread its own partial does not help when the target
function is shared (0.24x). Making the partial, its function and its args tuple
immortal also does not help, because the per-partial empty keywords dict is still
refcounted (0.20x). Scaling comes back only when all three (fn, args, keywords)
are immortal (5.43x). The source of partial_vectorcall is the same in v3.15.0
(tagged 2026-10-09) and main. The 3.14 branch did not take the backport.

Environment

Interpreter A 3.14.3 free-threading build (main, Mar 21 2026, 17:04:12) [Clang 22.1.1 ] (pyenv source build)
Interpreter B 3.15.0rc3 free-threading build (main, Oct 3 2026, 00:55:19) [Clang 22.1.3 ] (python-build-standalone, installed by uv)
GIL disabled in both (sys._is_gil_enabled() == False)
Platform macOS-14.8.2-arm64-arm-64bit-Mach-O
CPU Apple M1 Max, 10 cores (8 performance + 2 efficiency), os.cpu_count() == 10, 128-byte cache line
Load other work was running on the host, and probes ran under nice -n 10

24 threads on 10 cores is deliberate oversubscription (it matches a 24-worker
server). The no-shared-state control loop therefore tops out at about 5.5-6.7x on
this machine, not 24x.

Results

The minimal repro is below. Each cell is aggregate calls/s over 1 s per
measurement. Speedup = calls/s at N threads ÷ calls/s at 1 thread, for the same
shape.

shape 3.14.3t 1T 3.14.3t 24T 3.14.3t speedup 3.15.0rc3t 1T 3.15.0rc3t 24T 3.15.0rc3t speedup
control (a function private to each thread) 31.3 M 208.8 M 6.67x 38.4 M 216.5 M 5.63x
shared: one partial(f, 1) for all threads 19.1 M 121.8 M 6.37x 18.6 M 3.7 M 0.20x
own: one partial(f, 1) per thread, f shared 19.1 M 103.2 M 5.40x 19.9 M 4.7 M 0.24x

Thread sweep (0.5 s per point):

threads 3.14.3t shared 3.15.0rc3t shared 3.15.0rc3t own (shared f) 3.15.0rc3t control
1 1.00x (19.1 M) 1.00x (17.9 M) 1.00x (20.4 M) 1.00x (38.8 M)
2 1.57x 0.34x 0.48x 1.89x
4 3.49x 0.30x 0.29x 2.85x
8 5.40x 0.20x 0.20x 3.42x
24 4.95x 0.21x 0.22x 4.50x

On 3.15, 1-thread cost is unchanged. An uncontended atomic is cheap. The cost is
all in contention.

Variants that isolate the mechanism

These come from probe_variants.py (Appendix B), with 1 and 24 threads at 1 s
each. "shared" means one object built by the main thread and called by all
threads. "own" means each worker builds its own object from the same recipe.
Every thread passes its own argument object. Pinning is done with the exported
_Py_SetImmortal through ctypes.

variant what is shared / pinned 3.14.3t shared 3.14.3t own 3.15.0rc3t shared 3.15.0rc3t own
partial(f, 1)(2) 6.21x 6.22x 0.17x 0.19x
own partial over a per-thread function nothing shared in "own" 6.08x 6.16x 0.18x 6.42x
partial(max, 1)(2) (builtin target) max shared in both shapes 6.45x 6.26x 0.19x 0.20x
partial subclass (f, 1)(2) 6.24x 6.29x 0.17x 0.17x
operator.call(p, 2) 5.97x 6.02x 0.19x 0.18x
pin fn only p.func immortal 5.53x — 0.21x 5.72x¹
pin keywords dict only p.keywords immortal 5.54x — 0.17x 0.19x
pin partial + fn + args keywords dict not pinned 5.46x — 0.20x 5.55x¹
pin partial + fn + args + keywords dict everything the call touches 5.77x — 5.43x 5.50x¹
partial(fk, a=1)(2) (stored keyword) 1.00x 6.76x 0.45x 0.29x
partial(fk, a=1), everything pinned 0.83x 6.15x¹ 0.67x 5.65x¹

¹ In the pinning rows, the "own" partials are built after the shared run has
already made f (or fk) immortal, so they show that a private partial over an
immortal target scales.

What the table shows:

  • The partial type, its vectorcall path and the partial object's own refcount are
    not the problem. Own partials over own functions scale on 3.15 (6.42x).
  • The problem is a write to every object the call touches: the target function,
    the args tuple and the keywords dict. All three have to be immortal before 3.15
    scales again. Pinning any one or two of them still leaves 0.17-0.21x.
  • Keyword-bound partials serialise on both versions even with everything
    immortal (0.83x / 0.67x). That is a separate, older bottleneck: a per-call lock
    on the shared keywords dict. See Keyword partials.

Minimal reproducer

Standard library only. Save it as repro_partial_shared.py and run it with a
free-threaded interpreter:

python3.14t repro_partial_shared.py            # 1 and 24 threads, 1 s each
python3.15t repro_partial_shared.py            # 1 and 24 threads, 1 s each
python3.15t repro_partial_shared.py 1 2 4 8 24 0.5   # sweep, 0.5 s per point
"""Free-threaded scaling of ONE shared functools.partial vs one partial per thread.

Standard library only. Run on a free-threaded build (python3.14t / python3.15t):

    python3.15t repro_partial_shared.py            # 1 and 24 threads, 1 s each
    python3.15t repro_partial_shared.py 1 8 24 2.0 # thread counts..., seconds

Each thread calls its target in a tight loop for SECONDS and counts calls.
"speedup" = aggregate calls/s at N threads / calls/s at 1 thread.
"""

import functools
import os
import platform
import sys
import threading
import time


def f(a, b):
    return b


SHARED = functools.partial(f, 1)  # one object, called by every thread


def control_target(b):  # the no-shared-state control: each thread defines its own below
    return b


def worker(mode, seconds, barrier, rates, slot):
    if mode == "shared":
        target = SHARED
    elif mode == "own":
        target = functools.partial(f, 1)  # one partial per thread
    else:  # "control": a function object private to this thread
        def target(b):
            return b
    clock = time.perf_counter
    barrier.wait()
    calls = 0
    start = clock()
    end = start + seconds
    while True:
        for _ in range(1000):
            target(2)
        calls += 1000
        now = clock()
        if now >= end:
            break
    rates[slot] = calls / (now - start)


def measure(mode, threads, seconds):
    rates = [0.0] * threads
    barrier = threading.Barrier(threads)
    ts = [
        threading.Thread(target=worker, args=(mode, seconds, barrier, rates, i))
        for i in range(threads)
    ]
    for t in ts:
        t.start()
    for t in ts:
        t.join()
    return sum(rates)


def main():
    nums = [a for a in sys.argv[1:] if "." not in a]
    secs = [a for a in sys.argv[1:] if "." in a]
    counts = [int(a) for a in nums] or [1, 24]
    seconds = float(secs[0]) if secs else 1.0
    gil = sys._is_gil_enabled()
    print(sys.version.replace("\n", " "))
    print(platform.platform(), "| cpus:", os.cpu_count(), "| GIL enabled:", gil)
    base = {}
    print(f"{'mode':8} {'threads':>7} {'calls/s':>14} {'speedup':>8}")
    for mode in ("control", "shared", "own"):
        for n in counts:
            rate = measure(mode, n, seconds)
            if n == counts[0]:
                base[mode] = rate
            speed = rate / base[mode]
            print(f"{mode:8} {n:7d} {rate:14,.0f} {speed:7.2f}x")


if __name__ == "__main__":
    main()

Output on the machine above:

3.14.3 free-threading build (main, Mar 21 2026, 17:04:12) [Clang 22.1.1 ]
macOS-14.8.2-arm64-arm-64bit-Mach-O | cpus: 10 | GIL enabled: False
mode     threads        calls/s  speedup
control        1     31,330,215    1.00x
control       24    208,830,981    6.67x
shared         1     19,134,731    1.00x
shared        24    121,813,373    6.37x
own            1     19,124,673    1.00x
own           24    103,228,171    5.40x

3.15.0rc3 free-threading build (main, Oct  3 2026, 00:55:19) [Clang 22.1.3 ]
macOS-14.8.2-arm64-arm-64bit-Mach-O | cpus: 10 | GIL enabled: False
mode     threads        calls/s  speedup
control        1     38,416,566    1.00x
control       24    216,453,219    5.63x
shared         1     18,561,162    1.00x
shared        24      3,729,330    0.20x
own            1     19,941,575    1.00x
own           24      4,743,980    0.24x

The probe keeps the shared object in a local variable, so only the call is
measured. It does not measure a LOAD_GLOBAL of a non-deferred object, which
carries its own refcount traffic.

Root cause

The change

gh-154189, "use-after-free:
functools partial", was reported 2026-07-19 against main (free-threading debug
TSan build). It is a single-threaded re-entrancy bug. partial_vectorcall
held a raw PyObject ** into pto->args. In the keyword-merge path, a call
keyword's __hash__ could call partial.__setstate__ on the same partial, which
freed the old args tuple while the C frame was still using it.

The fix is PR gh-154508, merged to
main 2026-07-24 as
1a742d4,
and backported to 3.15 by gh-154652
(6246342,
2026-07-27). The NEWS entry reads:

Fixed a potential use-after-free when calling functools.partial. Now, when
invoking a partial() object, the stored function, positional arguments, and
keyword arguments are preserved for the duration of the call in case of
reentrancy.

Which releases have it:

  • 3.15: first in v3.15.0rc1. v3.15.0b4 does not contain the commit; rc1 does.
    Present in rc2, rc3 and v3.15.0 final (tagged 2026-10-09; the
    partial_vectorcall source is byte-identical to rc3) and in main.
  • 3.14: not backported. The automatic cherry-pick conflicted, and the author and
    reviewer agreed a backport was not worth the effort ("I don't think PR needs to
    be backported to 3.14." / "I agree; it doesn't look worth the effort."). v3.14.8
    has the old code.

The issue and PR do not discuss free-threaded performance, and the PR includes no
benchmark.

Before: v3.14.3 Modules/_functoolsmodule.c, partial_vectorcall (positional path)

    /* pto->kw is mutable, so need to check every time */
    if (PyDict_GET_SIZE(pto->kw)) {
        return partial_vectorcall_fallback(tstate, pto, args, nargsf, kwnames);
    }
    ...
    PyObject **pto_args = _PyTuple_ITEMS(pto->args);
    Py_ssize_t pto_nargs = PyTuple_GET_SIZE(pto->args);
    ...
        if (pto_nargs == 1 && (nargsf & PY_VECTORCALL_ARGUMENTS_OFFSET)) {
            PyObject **newargs = (PyObject **)args - 1;
            PyObject *tmp = newargs[0];
            newargs[0] = pto_args[0];
            PyObject *ret = _PyObject_VectorcallTstate(tstate, pto->fn, newargs,
                                                       nargs + 1, kwnames);
            newargs[0] = tmp;
            return ret;
        }

All reads are borrowed. The call path performs no refcount writes on pto->fn,
pto->args or pto->kw.

After: v3.15.0rc3 / v3.15.0 / main, partial_vectorcall

    PyObject *result = NULL;
    PyObject *partial_function = Py_NewRef(pto->fn);
    PyObject *partial_args = Py_NewRef(pto->args);
    PyObject *partial_keywords = Py_NewRef(pto->kw);

    PyObject **pto_args = _PyTuple_ITEMS(partial_args);
    Py_ssize_t pto_nargs = PyTuple_GET_SIZE(partial_args);
    Py_ssize_t pto_nkwds = PyDict_GET_SIZE(partial_keywords);
    ...
            result = _PyObject_VectorcallTstate(tstate, partial_function, newargs,
                                                nargs + 1, kwnames);
            newargs[0] = tmp;
            goto done;
    ...
 done:
    Py_DECREF(partial_function);
    Py_DECREF(partial_args);
    Py_DECREF(partial_keywords);
    return result;

The three Py_NewRef / Py_DECREF pairs run on every call, including the
positional-only fast paths, where no user code runs between reading the fields and
making the call. Keyword merging is the only place the re-entrancy in gh-154189
can happen.

Mechanism

In the free-threaded build each object has an owning thread (ob_tid), a
non-atomic ob_ref_local that only the owner writes, and an atomic
ob_ref_shared that every other thread writes (Include/refcount.h, 3.15):

static inline Py_ALWAYS_INLINE void Py_INCREF(PyObject *op)
{
    uint32_t local = _Py_atomic_load_uint32_relaxed(&op->ob_ref_local);
    uint32_t new_local = local + 1;
    if (new_local == 0) {
        // local is equal to _Py_IMMORTAL_REFCNT_LOCAL: do nothing
        return;
    }
    if (_Py_IsOwnedByCurrentThread(op)) {
        _Py_atomic_store_uint32_relaxed(&op->ob_ref_local, new_local);
    }
    else {
        _Py_atomic_add_ssize(&op->ob_ref_shared, (1 << _Py_REF_SHARED_SHIFT));
    }

Py_DECREF from a non-owner goes to _Py_DecRefShared (Objects/object.c). That
function runs a compare-and-swap loop on ob_ref_shared, and the loop retries
under contention.

A long-lived partial and its fn, args and kw are owned by the thread that
created them, usually the importing or main thread. Every worker thread is
therefore a non-owner. So each call on 3.15 does 3 atomic adds and 3 CAS loops on
three objects that every thread touches. Each read-modify-write needs exclusive
ownership of the cache line, so the line holding each object header moves between
cores on every call. The header also holds ob_type, ob_tid and
ob_ref_local, which every caller reads, so those reads miss too. Throughput is
then limited by how fast the coherence fabric can move those lines, not by the
number of cores. The aggregate settles around 3-5 M calls/s at any thread count:
0.34x at 2 threads, about 0.2x from 8 threads up.

Why own partials over a shared function still collapse. fn is shared, so
Py_NewRef(pto->fn) writes the same line from every thread. The per-thread args
tuple and keywords dict are private, so the slowdown is a little smaller (0.24x
vs 0.20x), but it is the same effect. Own partials over per-thread functions
touch only private objects and scale (6.42x).

Why 3.14 did not have this. The callee side never writes the shared refcounts.
_PyEval_Vector takes its references with PyStackRef_FromPyObjectNew:

    for (size_t i = 0; i < argcount; i++) {
        arguments[i] = PyStackRef_FromPyObjectNew(args[i]);
    }
    ...
    _PyInterpreterFrame *frame = _PyEvalFramePushAndInit(
        tstate, PyStackRef_FromPyObjectNew(func), locals,

In the free-threaded build that is a no-op for immortal objects (small ints) and
for deferred-refcount objects. Top-level functions and methods are created with
deferred refcounting (Objects/funcobject.c: _PyObject_SetDeferredRefcount when
the code is not CO_NESTED, or is CO_METHOD). Py_NewRef does not check the
deferred bit, so on 3.15 the same deferred f gets an atomic increment on every
call. That is why a plain shared f(…) call scales on both versions (6.0x /
6.7x) while partial(f, 1) does not on 3.15.

Why making partial + fn + args immortal did not help. partial_new always
stores a dict in pto->kw. With no keywords it is a fresh, empty dict per partial,
owned by the creating thread. 3.15 increfs and decrefs that dict on every call even
though it is empty. Pinning the partial object itself does nothing, because the
call path never touches the partial's own refcount. Only pinning fn, args and
the keywords dict
brings scaling back (5.43x). The same reasoning covers a
partial of a builtin (max): PyCFunction objects are neither immortal nor
deferred, and the 3.15 path increfs them.

Is it intended?

The correctness fix is intended: a real use-after-free reachable from pure Python.
The free-threaded cost looks unintended. Nothing in the issue, the PR review or the
NEWS entry mentions it. The PR has no benchmark, and the 3.14 backport was skipped
for effort, not because of any trade-off. The fix also guards less than its cost
suggests:

No open issue tracks the performance regression. Searches of the cpython tracker
for partial plus free-threading, scaling or contention since 2026-07 found only the
crash reports above.

Suggested upstream fixes

__setstate__ is the only writer of fn, args and kw. It is called by pickle
and copy on a freshly built object, almost never on a live, shared one. The fix
should make that rare writer pay and keep the per-call reader free of shared
writes. In rough order of preference:

  1. Take the per-call snapshot as GC-visible C stack references instead of
    Py_NewRef.
    CPython already has the machinery: _PyCStackRef with
    _PyThreadState_PushCStackRefNew / _PyThreadState_PopCStackRef
    (Include/internal/pycore_stackref.h). typeobject.c's call_method uses it
    for exactly this shape: it holds the looked-up __call__ function for the
    duration of a call without touching that shared function's refcount. That is
    why a per-thread instance of a class with __call__ scales on 3.15 even though
    every thread calls the same __call__ function. PyStackRef_FromPyObjectNew
    costs nothing for deferred and immortal objects, and the GC sees C stack refs,
    so a deferred object cannot be collected under the call. For fn this is
    enough on its own when fn is a top-level function (already deferred). For
    args and kw, partial_new / partial_setstate could enable deferred
    refcounting on the tuple and dict they store, and stop allocating a fresh empty
    dict when there are no keywords (keep NULL, or share one immortal empty
    sentinel internally). Result: correctness against same-thread re-entrancy is
    preserved, and the call does no shared writes.
  2. Keep the positional fast paths borrowed, as in 3.14, and take strong
    references only on the keyword-merge path.
    That path is where user code
    (__hash__ / __eq__ of call keywords, PyDict_Copy, PyDict_SetItem) runs
    before the call. This is the narrowest fix for the reported reproducer. It
    should be paired with (3) to cover a callee that re-enters __setstate__ on its
    own partial.
  3. Make the displaced state outlive any call in progress, in the writer. In
    __setstate__, retire the old fn, args and kw with
    _PyObject_XDecRefDelayed (QSBR) in the free-threaded build. That handles
    cross-thread readers, the Free-threaded CPython 3.14: concurrent functools.partial.__setstate__ crashes with repr(), calls, or another __setstate__ #157841 shape. It does not fully cover same-thread
    re-entrancy, because the calling thread can report a quiescent state while its
    call is still running (for example, when it detaches for blocking I/O).
    Alternatively, keep the displaced objects alive on the partial until
    tp_dealloc. That is simple, costs nothing per call, and only matters for the
    rare object that is re-stated.
  4. For Free-threaded CPython 3.14: concurrent functools.partial.__setstate__ crashes with repr(), calls, or another __setstate__ #157841, do not add a per-call lock on the partial. Lock in the
    writer, and publish a coherent state the reader can load without writing a
    shared line: an immutable state record holding fn, args, kw and
    phcount, read with an acquire load into a C stack ref and retired via QSBR.
    An atomic pointer plus _Py_TryIncrefCompare on a non-deferred state object
    would bring back the same contention this report measures.

Keyword partials: a separate, older bottleneck

A partial with stored keywords serialises when shared, on both versions, even
with everything immortal: 0.83x on 3.14.3t, 0.67x on 3.15.0rc3t.

Both are per-call acquisitions of the same per-object mutex. The lock exists
because p.keywords returns the live dict and user code can mutate it. A
contention-free design would snapshot the keywords into immutable
(kwnames tuple, values) at construction and in __setstate__. That is a
behaviour change for code that mutates p.keywords in place, which is probably
rare and undocumented.

Workarounds for users (until fixed)

Measured on 3.15.0rc3t, shared object called by 24 threads:

replacement for a shared partial(f, x) 3.15.0rc3t 3.14.3t notes
top-level def p(b): return f(CONST, b) 6.71x 5.96x Best with no C code, when the bound value is a constant or immortal. Module-level functions are deferred-refcounted.
closure over the bound value, cells pinned (immortal) 6.33x 6.01x Pin with PyUnstable_SetImmortal (3.15 public unstable C API) or _Py_SetImmortal; leaks by design.
closure over the bound value, cells not pinned 0.18x 0.14x COPY_FREE_VARS increfs each shared cell per call, on both versions.
instance of a slotted class with __call__, instance pinned 5.55x 6.05x
instance of a slotted class with __call__, not pinned 0.21x 0.31x self is increfed into the callee frame per call, on both versions.
partial with fn, args and keywords all pinned 5.43x 5.77x Pinning only partial + fn + args is not enough (0.20x).
one partial per thread over a per-thread function 6.42x 6.16x One partial per thread over a shared function still collapses (0.19x).

General rule on free-threaded CPython: a long-lived callable shared by many
threads scales only if every object the call path increfs is immortal or deferred
and reached through stackrefs. Bound self, closure cells and, on 3.15, partial
fields are increfed per call.

Appendix A: other common callables (shared vs own, 1 → 24 threads)

callable 3.14.3t shared 3.14.3t own 3.15.0rc3t shared 3.15.0rc3t own verdict
plain top-level function 5.96x 6.23x 6.71x 6.07x scales on both
operator.itemgetter(1)(tup) 6.11x 6.13x 5.85x 5.87x scales on both
operator.attrgetter("a")(obj) 6.16x 6.06x 5.82x 5.92x scales on both
operator.methodcaller("count", 1)(tup) 0.31x 0.30x 5.15x 5.23x serialises on 3.14 even per thread; fixed in 3.15
operator.methodcaller("meth", b=1)(obj) 0.25x 0.28x 5.43x 5.70x same as above
bound builtin method {}.get 5.58x 5.46x 5.36x 5.70x scales on both
bound Python method obj.meth, not pinned 0.17x 5.96x 0.14x 5.82x shared serialises on both (incref of self)
bound Python method, instance + method pinned 5.93x 5.76x 5.17x 5.99x scales
slotted instance __call__, not pinned 0.31x 6.72x 0.21x 6.20x shared serialises on both (incref of self)
closure, cells not pinned 0.14x 5.94x 0.18x 6.36x shared serialises on both (cell incref)
functools.partial (positional) 6.21x 6.22x 0.17x 0.19x regressed in 3.15
functools.partial (keyword) 1.00x 6.76x 0.45x 0.29x shared serialises on both (dict lock); 3.15 adds the fn incref

Only functools.partial regressed between 3.14 and 3.15. methodcaller improved
in 3.15. The 3.14 collapse even with per-thread objects is probably the per-call
reference to the shared method descriptor found by attribute lookup; that was not
investigated further. The bound-method, slotted __call__ and closure collapses
are older and the same on both versions. Each comes from a per-call incref of a
shared, non-immortal, non-deferred object (self or a cell), and pinning that
object fixes it.

Appendix B: the variant probe

Run it as python3.15t probe_variants.py <variant ...> (LIST prints the names).
Pinning variants permanently immortalise the target function, so run each one in a
fresh process, as was done for the tables above.

"""Variant probe: which callables stop scaling when one instance is shared by N threads.

Standard library only (ctypes is used solely to call the exported
`_Py_SetImmortal` so a variant can pin an object). Free-threaded build required.

    python3.15t probe_variants.py LIST                 # list variant names
    python3.15t probe_variants.py partial_pos closure  # run some variants
    python3.15t probe_variants.py --seconds 0.5 --threads 1,24 all

For every variant two shapes are measured at each thread count:
  shared  one object built once by the main thread, called by every thread
  own     each thread builds its own object (by the same recipe) before the start
Every thread passes its own argument object, so only the callable is shared.
speedup = aggregate calls/s at N threads / calls/s at 1 thread (same shape).
"""

import ctypes
import functools
import json
import operator
import os
import sys
import threading
import time

_set_immortal = ctypes.pythonapi._Py_SetImmortal
_set_immortal.argtypes = [ctypes.py_object]
_set_immortal.restype = None


def pin(*objs):
    for o in objs:
        _set_immortal(o)


def f(a, b):
    return b


def fk(b, a=0):
    return b


def g(b):
    return b


class Slotted:
    __slots__ = ("a",)

    def __init__(self, a):
        self.a = a

    def __call__(self, b):
        return b

    def meth(self, b):
        return b


class PartialSub(functools.partial):
    pass


def make_closure(a):
    def inner(b):
        return b if a else a
    return inner


# Each variant: name -> (factory, argument factory, description).
# factory(shared: bool) returns the callable; for "shared" it runs once in the
# main thread, for "own" once per worker thread.
def _partial_pinned(which):
    def factory(shared):
        p = functools.partial(f, 1)
        if shared:
            objs = {"all": (p, p.func, p.args, p.keywords),
                    "fn_args": (p, p.func, p.args),
                    "kw": (p.keywords,),
                    "fn": (p.func,),
                    "args": (p.args,)}[which]
            pin(*objs)
        return p
    return factory


def _partial_kw_pinned(shared):
    p = functools.partial(fk, a=1)
    if shared:
        pin(p, p.func, p.args, p.keywords)
    return p


def _partial_own_fn(shared):
    if shared:
        return functools.partial(f, 1)
    def own_f(a, b):
        return b
    return functools.partial(own_f, 1)


def _closure_pinned(shared):
    c = make_closure(1)
    if shared:
        pin(*c.__closure__)
    return c


def _bound_py(shared):
    return Slotted(1).meth


def _slotted_pinned(shared):
    s = Slotted(1)
    if shared:
        pin(s)
    return s


def _bound_py_pinned(shared):
    s = Slotted(1)
    m = s.meth
    if shared:
        pin(s, m)
    return m


def _bound_c(shared):
    return {}.get  # dict.get bound to a dict; call d.get(2) -> None


VARIANTS = {
    "func": (lambda s: g, lambda: 2, "plain module-level Python function g(b)"),
    "partial_pos": (lambda s: functools.partial(f, 1), lambda: 2, "partial(f, 1)(2)"),
    "partial_ownfn": (_partial_own_fn, lambda: 2,
                      "own = partial over a per-thread function (shared = partial_pos)"),
    "partial_pin_all": (_partial_pinned("all"), lambda: 2,
                        "shared partial; partial, func, args tuple, keywords dict immortal"),
    "partial_pin_fn_args": (_partial_pinned("fn_args"), lambda: 2,
                            "shared partial; partial, func, args immortal (keywords dict NOT)"),
    "partial_pin_kw": (_partial_pinned("kw"), lambda: 2,
                       "shared partial; only the empty keywords dict immortal"),
    "partial_pin_fn": (_partial_pinned("fn"), lambda: 2,
                       "shared partial; only func immortal"),
    "partial_kw": (lambda s: functools.partial(fk, a=1), lambda: 2, "partial(fk, a=1)(2)"),
    "partial_kw_pin_all": (_partial_kw_pinned, lambda: 2,
                           "shared partial(fk, a=1); partial, func, args, keywords immortal"),
    "partial_builtin": (lambda s: functools.partial(max, 1), lambda: 2,
                        "partial(max, 1)(2) -- builtin target"),
    "partial_sub": (lambda s: PartialSub(f, 1), lambda: 2, "partial subclass(f, 1)(2)"),
    "partial_opcall": (lambda s: functools.partial(f, 1), lambda: 2,
                       "operator.call(p, 2) [loop uses operator.call]"),
    "closure": (lambda s: make_closure(1), lambda: 2, "closure over a cell, not pinned"),
    "closure_pin": (_closure_pinned, lambda: 2, "closure, cells immortal when shared"),
    "slotted_call": (lambda s: Slotted(1), lambda: 2, "slotted instance with __call__"),
    "itemgetter": (lambda s: operator.itemgetter(1), lambda: (1, 2, 3), "itemgetter(1)(tup)"),
    "attrgetter": (lambda s: operator.attrgetter("a"), lambda: Slotted(2), "attrgetter('a')(obj)"),
    "methodcaller": (lambda s: operator.methodcaller("count", 1), lambda: (1, 2, 3),
                     "methodcaller('count', 1)(tup)"),
    "methodcaller_kw": (lambda s: operator.methodcaller("meth", b=1), lambda: Slotted(2),
                        "methodcaller('meth', b=1)(obj)"),
    "bound_py": (_bound_py, lambda: 2, "bound method of a Python class instance"),
    "slotted_call_pin": (_slotted_pinned, lambda: 2, "slotted __call__, instance immortal when shared"),
    "bound_py_pin": (_bound_py_pinned, lambda: 2, "bound Python method, instance + method immortal when shared"),
    "bound_c": (_bound_c, lambda: 2, "bound builtin method {}.get"),
}


def worker(name, shared_obj, seconds, barrier, rates, slot):
    factory, argf, _ = VARIANTS[name]
    target = shared_obj if shared_obj is not None else factory(False)
    arg = argf()
    clock = time.perf_counter
    opcall = operator.call
    barrier.wait()
    calls = 0
    start = clock()
    end = start + seconds
    if name == "partial_opcall":
        while True:
            for _ in range(1000):
                opcall(target, arg)
            calls += 1000
            now = clock()
            if now >= end:
                break
    else:
        while True:
            for _ in range(1000):
                target(arg)
            calls += 1000
            now = clock()
            if now >= end:
                break
    rates[slot] = calls / (now - start)


def measure(name, shared, threads, seconds):
    factory = VARIANTS[name][0]
    shared_obj = factory(True) if shared else None
    rates = [0.0] * threads
    barrier = threading.Barrier(threads)
    ts = [threading.Thread(target=worker, args=(name, shared_obj, seconds, barrier, rates, i))
          for i in range(threads)]
    for t in ts:
        t.start()
    for t in ts:
        t.join()
    return sum(rates)


def main():
    args = sys.argv[1:]
    seconds, threads = 1.0, [1, 24]
    if "--seconds" in args:
        i = args.index("--seconds")
        seconds = float(args[i + 1])
        del args[i:i + 2]
    if "--threads" in args:
        i = args.index("--threads")
        threads = [int(x) for x in args[i + 1].split(",")]
        del args[i:i + 2]
    if args == ["LIST"]:
        for k, v in VARIANTS.items():
            print(f"{k:20} {v[2]}")
        return
    names = list(VARIANTS) if args == ["all"] else args
    ver = f"{sys.version_info.major}.{sys.version_info.minor}"
    print(sys.version.replace("\n", " "), "| cpus:", os.cpu_count(),
          "| GIL:", sys._is_gil_enabled())
    print(f"{'variant':20} {'shape':6} " + " ".join(f"{n:>3}T calls/s" for n in threads)
          + "  speedup")
    out_path = os.path.join(os.path.dirname(os.path.abspath(__file__)), f"results_{ver}.jsonl")
    with open(out_path, "a") as out:
        for name in names:
            for shape in ("shared", "own"):
                rates = [measure(name, shape == "shared", n, seconds) for n in threads]
                speed = rates[-1] / rates[0]
                print(f"{name:20} {shape:6} " + " ".join(f"{r:14,.0f}" for r in rates)
                      + f"  {speed:6.2f}x", flush=True)
                out.write(json.dumps({"py": ver, "variant": name, "shape": shape,
                                      "threads": threads, "rates": rates,
                                      "speedup": speed, "seconds": seconds}) + "\n")


if __name__ == "__main__":
    main()

How we found it

A free-threaded web framework runs 24 worker threads that share module-level
callables. Its per-workload scale-out ledger (1/8/24 threads against a
no-shared-state control) showed every path that called a long-lived
functools.partial dropping from "scales" on 3.14 to below 1x on 3.15. The
framework's workaround was to replace every long-lived shared partial with a
closure whose cells are pinned at construction.

=========================

CPython versions tested on:

3.15

Operating systems tested on:

macOS

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    type-bugAn unexpected behavior, bug, or error

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions