You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Free-threaded 3.15: a functools.partial shared between threads no longer scales (it slows down as you add threads) #159157
I was migrating some of my "massively threaded" 3.14t code to 3.15t and my performance agent saw everything collapse under concurrency after the python version change.
Turns out a new 3.15-only fix (?) for a partial application is now over-locking everything. For my own use I just migrated away from using functools.partial everywhere now, but ... FYI i guess.
agent report:
=========================
Free-threaded 3.15: a functools.partial shared between threads no longer scales (it slows down as you add threads)
Summary
On the free-threaded build, CPython 3.15 calls to a functools.partial stop
scaling across threads when the partial (or just its target function) is shared
between threads. On a 10-core Apple M1 Max, 24 threads calling one module-level partial(f, 1) reach 0.20x the single-thread rate, so total throughput drops
when threads are added. 3.14.3t reaches 6.37x on the same machine. The
slowdown shows up at 2 threads (0.34x). The cause is the use-after-free fix for gh-154189 (PR gh-154508, 3.15 backport gh-154652, first shipped in 3.15.0rc1). It
made partial_vectorcall take and drop strong references to pto->fn, pto->args and pto->kw on every call. In the free-threaded build an
incref/decref from a thread that does not own the object is an atomic
read-modify-write on that object's ob_ref_shared. Every calling thread now writes
the same three cache lines (the target function, the bound-args tuple, and the
keywords dict, which partial allocates even when empty) six times per call. In
3.14 the call path takes no references on these objects: it reads them borrowed,
and the callee frame uses stackrefs that skip deferred-refcount and immortal
objects. Giving each thread its own partial does not help when the target
function is shared (0.24x). Making the partial, its function and its args tuple
immortal also does not help, because the per-partial empty keywords dict is still
refcounted (0.20x). Scaling comes back only when all three (fn, args, keywords)
are immortal (5.43x). The source of partial_vectorcall is the same in v3.15.0
(tagged 2026-10-09) and main. The 3.14 branch did not take the backport.
3.15.0rc3 free-threading build (main, Oct 3 2026, 00:55:19) [Clang 22.1.3 ] (python-build-standalone, installed by uv)
GIL
disabled in both (sys._is_gil_enabled() == False)
Platform
macOS-14.8.2-arm64-arm-64bit-Mach-O
CPU
Apple M1 Max, 10 cores (8 performance + 2 efficiency), os.cpu_count() == 10, 128-byte cache line
Load
other work was running on the host, and probes ran under nice -n 10
24 threads on 10 cores is deliberate oversubscription (it matches a 24-worker
server). The no-shared-state control loop therefore tops out at about 5.5-6.7x on
this machine, not 24x.
Results
The minimal repro is below. Each cell is aggregate calls/s over 1 s per
measurement. Speedup = calls/s at N threads ÷ calls/s at 1 thread, for the same
shape.
shape
3.14.3t 1T
3.14.3t 24T
3.14.3t speedup
3.15.0rc3t 1T
3.15.0rc3t 24T
3.15.0rc3t speedup
control (a function private to each thread)
31.3 M
208.8 M
6.67x
38.4 M
216.5 M
5.63x
shared: one partial(f, 1) for all threads
19.1 M
121.8 M
6.37x
18.6 M
3.7 M
0.20x
own: one partial(f, 1) per thread, f shared
19.1 M
103.2 M
5.40x
19.9 M
4.7 M
0.24x
Thread sweep (0.5 s per point):
threads
3.14.3t shared
3.15.0rc3t shared
3.15.0rc3t own (shared f)
3.15.0rc3t control
1
1.00x (19.1 M)
1.00x (17.9 M)
1.00x (20.4 M)
1.00x (38.8 M)
2
1.57x
0.34x
0.48x
1.89x
4
3.49x
0.30x
0.29x
2.85x
8
5.40x
0.20x
0.20x
3.42x
24
4.95x
0.21x
0.22x
4.50x
On 3.15, 1-thread cost is unchanged. An uncontended atomic is cheap. The cost is
all in contention.
Variants that isolate the mechanism
These come from probe_variants.py (Appendix B), with 1 and 24 threads at 1 s
each. "shared" means one object built by the main thread and called by all
threads. "own" means each worker builds its own object from the same recipe.
Every thread passes its own argument object. Pinning is done with the exported _Py_SetImmortal through ctypes.
variant
what is shared / pinned
3.14.3t shared
3.14.3t own
3.15.0rc3t shared
3.15.0rc3t own
partial(f, 1)(2)
6.21x
6.22x
0.17x
0.19x
own partial over a per-thread function
nothing shared in "own"
6.08x
6.16x
0.18x
6.42x
partial(max, 1)(2) (builtin target)
max shared in both shapes
6.45x
6.26x
0.19x
0.20x
partial subclass (f, 1)(2)
6.24x
6.29x
0.17x
0.17x
operator.call(p, 2)
5.97x
6.02x
0.19x
0.18x
pin fn only
p.func immortal
5.53x
—
0.21x
5.72x¹
pin keywords dict only
p.keywords immortal
5.54x
—
0.17x
0.19x
pin partial + fn + args
keywords dict not pinned
5.46x
—
0.20x
5.55x¹
pin partial + fn + args + keywords dict
everything the call touches
5.77x
—
5.43x
5.50x¹
partial(fk, a=1)(2) (stored keyword)
1.00x
6.76x
0.45x
0.29x
partial(fk, a=1), everything pinned
0.83x
6.15x¹
0.67x
5.65x¹
¹ In the pinning rows, the "own" partials are built after the shared run has
already made f (or fk) immortal, so they show that a private partial over an
immortal target scales.
What the table shows:
The partial type, its vectorcall path and the partial object's own refcount are
not the problem. Own partials over own functions scale on 3.15 (6.42x).
The problem is a write to every object the call touches: the target function,
the args tuple and the keywords dict. All three have to be immortal before 3.15
scales again. Pinning any one or two of them still leaves 0.17-0.21x.
Keyword-bound partials serialise on both versions even with everything
immortal (0.83x / 0.67x). That is a separate, older bottleneck: a per-call lock
on the shared keywords dict. See Keyword partials.
Minimal reproducer
Standard library only. Save it as repro_partial_shared.py and run it with a
free-threaded interpreter:
python3.14t repro_partial_shared.py # 1 and 24 threads, 1 s each
python3.15t repro_partial_shared.py # 1 and 24 threads, 1 s each
python3.15t repro_partial_shared.py 1 2 4 8 24 0.5 # sweep, 0.5 s per point
"""Free-threaded scaling of ONE shared functools.partial vs one partial per thread.Standard library only. Run on a free-threaded build (python3.14t / python3.15t): python3.15t repro_partial_shared.py # 1 and 24 threads, 1 s each python3.15t repro_partial_shared.py 1 8 24 2.0 # thread counts..., secondsEach thread calls its target in a tight loop for SECONDS and counts calls."speedup" = aggregate calls/s at N threads / calls/s at 1 thread."""importfunctoolsimportosimportplatformimportsysimportthreadingimporttimedeff(a, b):
returnbSHARED=functools.partial(f, 1) # one object, called by every threaddefcontrol_target(b): # the no-shared-state control: each thread defines its own belowreturnbdefworker(mode, seconds, barrier, rates, slot):
ifmode=="shared":
target=SHAREDelifmode=="own":
target=functools.partial(f, 1) # one partial per threadelse: # "control": a function object private to this threaddeftarget(b):
returnbclock=time.perf_counterbarrier.wait()
calls=0start=clock()
end=start+secondswhileTrue:
for_inrange(1000):
target(2)
calls+=1000now=clock()
ifnow>=end:
breakrates[slot] =calls/ (now-start)
defmeasure(mode, threads, seconds):
rates= [0.0] *threadsbarrier=threading.Barrier(threads)
ts= [
threading.Thread(target=worker, args=(mode, seconds, barrier, rates, i))
foriinrange(threads)
]
fortints:
t.start()
fortints:
t.join()
returnsum(rates)
defmain():
nums= [aforainsys.argv[1:] if"."notina]
secs= [aforainsys.argv[1:] if"."ina]
counts= [int(a) forainnums] or [1, 24]
seconds=float(secs[0]) ifsecselse1.0gil=sys._is_gil_enabled()
print(sys.version.replace("\n", " "))
print(platform.platform(), "| cpus:", os.cpu_count(), "| GIL enabled:", gil)
base= {}
print(f"{'mode':8}{'threads':>7}{'calls/s':>14}{'speedup':>8}")
formodein ("control", "shared", "own"):
fornincounts:
rate=measure(mode, n, seconds)
ifn==counts[0]:
base[mode] =ratespeed=rate/base[mode]
print(f"{mode:8}{n:7d}{rate:14,.0f}{speed:7.2f}x")
if__name__=="__main__":
main()
Output on the machine above:
3.14.3 free-threading build (main, Mar 21 2026, 17:04:12) [Clang 22.1.1 ]
macOS-14.8.2-arm64-arm-64bit-Mach-O | cpus: 10 | GIL enabled: False
mode threads calls/s speedup
control 1 31,330,215 1.00x
control 24 208,830,981 6.67x
shared 1 19,134,731 1.00x
shared 24 121,813,373 6.37x
own 1 19,124,673 1.00x
own 24 103,228,171 5.40x
3.15.0rc3 free-threading build (main, Oct 3 2026, 00:55:19) [Clang 22.1.3 ]
macOS-14.8.2-arm64-arm-64bit-Mach-O | cpus: 10 | GIL enabled: False
mode threads calls/s speedup
control 1 38,416,566 1.00x
control 24 216,453,219 5.63x
shared 1 18,561,162 1.00x
shared 24 3,729,330 0.20x
own 1 19,941,575 1.00x
own 24 4,743,980 0.24x
The probe keeps the shared object in a local variable, so only the call is
measured. It does not measure a LOAD_GLOBAL of a non-deferred object, which
carries its own refcount traffic.
Root cause
The change
gh-154189, "use-after-free:
functools partial", was reported 2026-07-19 against main (free-threading debug
TSan build). It is a single-threaded re-entrancy bug. partial_vectorcall
held a raw PyObject ** into pto->args. In the keyword-merge path, a call
keyword's __hash__ could call partial.__setstate__ on the same partial, which
freed the old args tuple while the C frame was still using it.
The fix is PR gh-154508, merged to main 2026-07-24 as 1a742d4,
and backported to 3.15 by gh-154652
(6246342,
2026-07-27). The NEWS entry reads:
Fixed a potential use-after-free when calling functools.partial. Now, when
invoking a partial() object, the stored function, positional arguments, and
keyword arguments are preserved for the duration of the call in case of
reentrancy.
Which releases have it:
3.15: first in v3.15.0rc1. v3.15.0b4 does not contain the commit; rc1 does.
Present in rc2, rc3 and v3.15.0 final (tagged 2026-10-09; the partial_vectorcall source is byte-identical to rc3) and in main.
3.14: not backported. The automatic cherry-pick conflicted, and the author and
reviewer agreed a backport was not worth the effort ("I don't think PR needs to
be backported to 3.14." / "I agree; it doesn't look worth the effort."). v3.14.8
has the old code.
The issue and PR do not discuss free-threaded performance, and the PR includes no
benchmark.
The three Py_NewRef / Py_DECREF pairs run on every call, including the
positional-only fast paths, where no user code runs between reading the fields and
making the call. Keyword merging is the only place the re-entrancy in gh-154189
can happen.
Mechanism
In the free-threaded build each object has an owning thread (ob_tid), a
non-atomic ob_ref_local that only the owner writes, and an atomic ob_ref_shared that every other thread writes (Include/refcount.h, 3.15):
staticinlinePy_ALWAYS_INLINEvoidPy_INCREF(PyObject*op)
{
uint32_tlocal=_Py_atomic_load_uint32_relaxed(&op->ob_ref_local);
uint32_tnew_local=local+1;
if (new_local==0) {
// local is equal to _Py_IMMORTAL_REFCNT_LOCAL: do nothingreturn;
}
if (_Py_IsOwnedByCurrentThread(op)) {
_Py_atomic_store_uint32_relaxed(&op->ob_ref_local, new_local);
}
else {
_Py_atomic_add_ssize(&op->ob_ref_shared, (1 << _Py_REF_SHARED_SHIFT));
}
Py_DECREF from a non-owner goes to _Py_DecRefShared (Objects/object.c). That
function runs a compare-and-swap loop on ob_ref_shared, and the loop retries
under contention.
A long-lived partial and its fn, args and kw are owned by the thread that
created them, usually the importing or main thread. Every worker thread is
therefore a non-owner. So each call on 3.15 does 3 atomic adds and 3 CAS loops on
three objects that every thread touches. Each read-modify-write needs exclusive
ownership of the cache line, so the line holding each object header moves between
cores on every call. The header also holds ob_type, ob_tid and ob_ref_local, which every caller reads, so those reads miss too. Throughput is
then limited by how fast the coherence fabric can move those lines, not by the
number of cores. The aggregate settles around 3-5 M calls/s at any thread count:
0.34x at 2 threads, about 0.2x from 8 threads up.
Why own partials over a shared function still collapse.fn is shared, so Py_NewRef(pto->fn) writes the same line from every thread. The per-thread args
tuple and keywords dict are private, so the slowdown is a little smaller (0.24x
vs 0.20x), but it is the same effect. Own partials over per-thread functions
touch only private objects and scale (6.42x).
Why 3.14 did not have this. The callee side never writes the shared refcounts. _PyEval_Vector takes its references with PyStackRef_FromPyObjectNew:
In the free-threaded build that is a no-op for immortal objects (small ints) and
for deferred-refcount objects. Top-level functions and methods are created with
deferred refcounting (Objects/funcobject.c: _PyObject_SetDeferredRefcount when
the code is not CO_NESTED, or is CO_METHOD). Py_NewRef does not check the
deferred bit, so on 3.15 the same deferred f gets an atomic increment on every
call. That is why a plain shared f(…) call scales on both versions (6.0x /
6.7x) while partial(f, 1) does not on 3.15.
Why making partial + fn + args immortal did not help.partial_new always
stores a dict in pto->kw. With no keywords it is a fresh, empty dict per partial,
owned by the creating thread. 3.15 increfs and decrefs that dict on every call even
though it is empty. Pinning the partial object itself does nothing, because the
call path never touches the partial's own refcount. Only pinning fn, args and
the keywords dict brings scaling back (5.43x). The same reasoning covers a
partial of a builtin (max): PyCFunction objects are neither immortal nor
deferred, and the 3.15 path increfs them.
Is it intended?
The correctness fix is intended: a real use-after-free reachable from pure Python.
The free-threaded cost looks unintended. Nothing in the issue, the PR review or the
NEWS entry mentions it. The PR has no benchmark, and the 3.14 backport was skipped
for effort, not because of any trade-off. The fix also guards less than its cost
suggests:
No open issue tracks the performance regression. Searches of the cpython tracker
for partial plus free-threading, scaling or contention since 2026-07 found only the
crash reports above.
Suggested upstream fixes
__setstate__ is the only writer of fn, args and kw. It is called by pickle
and copy on a freshly built object, almost never on a live, shared one. The fix
should make that rare writer pay and keep the per-call reader free of shared
writes. In rough order of preference:
Take the per-call snapshot as GC-visible C stack references instead of Py_NewRef. CPython already has the machinery: _PyCStackRef with _PyThreadState_PushCStackRefNew / _PyThreadState_PopCStackRef
(Include/internal/pycore_stackref.h). typeobject.c's call_method uses it
for exactly this shape: it holds the looked-up __call__ function for the
duration of a call without touching that shared function's refcount. That is
why a per-thread instance of a class with __call__ scales on 3.15 even though
every thread calls the same __call__ function. PyStackRef_FromPyObjectNew
costs nothing for deferred and immortal objects, and the GC sees C stack refs,
so a deferred object cannot be collected under the call. For fn this is
enough on its own when fn is a top-level function (already deferred). For args and kw, partial_new / partial_setstate could enable deferred
refcounting on the tuple and dict they store, and stop allocating a fresh empty
dict when there are no keywords (keep NULL, or share one immortal empty
sentinel internally). Result: correctness against same-thread re-entrancy is
preserved, and the call does no shared writes.
Keep the positional fast paths borrowed, as in 3.14, and take strong
references only on the keyword-merge path. That path is where user code
(__hash__ / __eq__ of call keywords, PyDict_Copy, PyDict_SetItem) runs
before the call. This is the narrowest fix for the reported reproducer. It
should be paired with (3) to cover a callee that re-enters __setstate__ on its
own partial.
Make the displaced state outlive any call in progress, in the writer. In __setstate__, retire the old fn, args and kw with _PyObject_XDecRefDelayed (QSBR) in the free-threaded build. That handles
cross-thread readers, the Free-threaded CPython 3.14: concurrent functools.partial.__setstate__ crashes with repr(), calls, or another __setstate__ #157841 shape. It does not fully cover same-thread
re-entrancy, because the calling thread can report a quiescent state while its
call is still running (for example, when it detaches for blocking I/O).
Alternatively, keep the displaced objects alive on the partial until tp_dealloc. That is simple, costs nothing per call, and only matters for the
rare object that is re-stated.
For Free-threaded CPython 3.14: concurrent functools.partial.__setstate__ crashes with repr(), calls, or another __setstate__ #157841, do not add a per-call lock on the partial. Lock in the
writer, and publish a coherent state the reader can load without writing a
shared line: an immutable state record holding fn, args, kw and phcount, read with an acquire load into a C stack ref and retired via QSBR.
An atomic pointer plus _Py_TryIncrefCompare on a non-deferred state object
would bring back the same contention this report measures.
Keyword partials: a separate, older bottleneck
A partial with stored keywords serialises when shared, on both versions, even
with everything immortal: 0.83x on 3.14.3t, 0.67x on 3.15.0rc3t.
In 3.14 the call falls back to partial_call, which does PyDict_Copy(pto->kw). That copy locks the shared source dict.
Both are per-call acquisitions of the same per-object mutex. The lock exists
because p.keywords returns the live dict and user code can mutate it. A
contention-free design would snapshot the keywords into immutable (kwnames tuple, values) at construction and in __setstate__. That is a
behaviour change for code that mutates p.keywords in place, which is probably
rare and undocumented.
Workarounds for users (until fixed)
Measured on 3.15.0rc3t, shared object called by 24 threads:
replacement for a shared partial(f, x)
3.15.0rc3t
3.14.3t
notes
top-level def p(b): return f(CONST, b)
6.71x
5.96x
Best with no C code, when the bound value is a constant or immortal. Module-level functions are deferred-refcounted.
closure over the bound value, cells pinned (immortal)
6.33x
6.01x
Pin with PyUnstable_SetImmortal (3.15 public unstable C API) or _Py_SetImmortal; leaks by design.
closure over the bound value, cells not pinned
0.18x
0.14x
COPY_FREE_VARS increfs each shared cell per call, on both versions.
instance of a slotted class with __call__, instance pinned
5.55x
6.05x
instance of a slotted class with __call__, not pinned
0.21x
0.31x
self is increfed into the callee frame per call, on both versions.
partial with fn, argsandkeywords all pinned
5.43x
5.77x
Pinning only partial + fn + args is not enough (0.20x).
one partial per thread over a per-thread function
6.42x
6.16x
One partial per thread over a shared function still collapses (0.19x).
General rule on free-threaded CPython: a long-lived callable shared by many
threads scales only if every object the call path increfs is immortal or deferred
and reached through stackrefs. Bound self, closure cells and, on 3.15, partial
fields are increfed per call.
Appendix A: other common callables (shared vs own, 1 → 24 threads)
callable
3.14.3t shared
3.14.3t own
3.15.0rc3t shared
3.15.0rc3t own
verdict
plain top-level function
5.96x
6.23x
6.71x
6.07x
scales on both
operator.itemgetter(1)(tup)
6.11x
6.13x
5.85x
5.87x
scales on both
operator.attrgetter("a")(obj)
6.16x
6.06x
5.82x
5.92x
scales on both
operator.methodcaller("count", 1)(tup)
0.31x
0.30x
5.15x
5.23x
serialises on 3.14 even per thread; fixed in 3.15
operator.methodcaller("meth", b=1)(obj)
0.25x
0.28x
5.43x
5.70x
same as above
bound builtin method {}.get
5.58x
5.46x
5.36x
5.70x
scales on both
bound Python method obj.meth, not pinned
0.17x
5.96x
0.14x
5.82x
shared serialises on both (incref of self)
bound Python method, instance + method pinned
5.93x
5.76x
5.17x
5.99x
scales
slotted instance __call__, not pinned
0.31x
6.72x
0.21x
6.20x
shared serialises on both (incref of self)
closure, cells not pinned
0.14x
5.94x
0.18x
6.36x
shared serialises on both (cell incref)
functools.partial (positional)
6.21x
6.22x
0.17x
0.19x
regressed in 3.15
functools.partial (keyword)
1.00x
6.76x
0.45x
0.29x
shared serialises on both (dict lock); 3.15 adds the fn incref
Only functools.partial regressed between 3.14 and 3.15. methodcaller improved
in 3.15. The 3.14 collapse even with per-thread objects is probably the per-call
reference to the shared method descriptor found by attribute lookup; that was not
investigated further. The bound-method, slotted __call__ and closure collapses
are older and the same on both versions. Each comes from a per-call incref of a
shared, non-immortal, non-deferred object (self or a cell), and pinning that
object fixes it.
Appendix B: the variant probe
Run it as python3.15t probe_variants.py <variant ...> (LIST prints the names).
Pinning variants permanently immortalise the target function, so run each one in a
fresh process, as was done for the tables above.
"""Variant probe: which callables stop scaling when one instance is shared by N threads.Standard library only (ctypes is used solely to call the exported`_Py_SetImmortal` so a variant can pin an object). Free-threaded build required. python3.15t probe_variants.py LIST # list variant names python3.15t probe_variants.py partial_pos closure # run some variants python3.15t probe_variants.py --seconds 0.5 --threads 1,24 allFor every variant two shapes are measured at each thread count: shared one object built once by the main thread, called by every thread own each thread builds its own object (by the same recipe) before the startEvery thread passes its own argument object, so only the callable is shared.speedup = aggregate calls/s at N threads / calls/s at 1 thread (same shape)."""importctypesimportfunctoolsimportjsonimportoperatorimportosimportsysimportthreadingimporttime_set_immortal=ctypes.pythonapi._Py_SetImmortal_set_immortal.argtypes= [ctypes.py_object]
_set_immortal.restype=Nonedefpin(*objs):
foroinobjs:
_set_immortal(o)
deff(a, b):
returnbdeffk(b, a=0):
returnbdefg(b):
returnbclassSlotted:
__slots__= ("a",)
def__init__(self, a):
self.a=adef__call__(self, b):
returnbdefmeth(self, b):
returnbclassPartialSub(functools.partial):
passdefmake_closure(a):
definner(b):
returnbifaelseareturninner# Each variant: name -> (factory, argument factory, description).# factory(shared: bool) returns the callable; for "shared" it runs once in the# main thread, for "own" once per worker thread.def_partial_pinned(which):
deffactory(shared):
p=functools.partial(f, 1)
ifshared:
objs= {"all": (p, p.func, p.args, p.keywords),
"fn_args": (p, p.func, p.args),
"kw": (p.keywords,),
"fn": (p.func,),
"args": (p.args,)}[which]
pin(*objs)
returnpreturnfactorydef_partial_kw_pinned(shared):
p=functools.partial(fk, a=1)
ifshared:
pin(p, p.func, p.args, p.keywords)
returnpdef_partial_own_fn(shared):
ifshared:
returnfunctools.partial(f, 1)
defown_f(a, b):
returnbreturnfunctools.partial(own_f, 1)
def_closure_pinned(shared):
c=make_closure(1)
ifshared:
pin(*c.__closure__)
returncdef_bound_py(shared):
returnSlotted(1).methdef_slotted_pinned(shared):
s=Slotted(1)
ifshared:
pin(s)
returnsdef_bound_py_pinned(shared):
s=Slotted(1)
m=s.methifshared:
pin(s, m)
returnmdef_bound_c(shared):
return {}.get# dict.get bound to a dict; call d.get(2) -> NoneVARIANTS= {
"func": (lambdas: g, lambda: 2, "plain module-level Python function g(b)"),
"partial_pos": (lambdas: functools.partial(f, 1), lambda: 2, "partial(f, 1)(2)"),
"partial_ownfn": (_partial_own_fn, lambda: 2,
"own = partial over a per-thread function (shared = partial_pos)"),
"partial_pin_all": (_partial_pinned("all"), lambda: 2,
"shared partial; partial, func, args tuple, keywords dict immortal"),
"partial_pin_fn_args": (_partial_pinned("fn_args"), lambda: 2,
"shared partial; partial, func, args immortal (keywords dict NOT)"),
"partial_pin_kw": (_partial_pinned("kw"), lambda: 2,
"shared partial; only the empty keywords dict immortal"),
"partial_pin_fn": (_partial_pinned("fn"), lambda: 2,
"shared partial; only func immortal"),
"partial_kw": (lambdas: functools.partial(fk, a=1), lambda: 2, "partial(fk, a=1)(2)"),
"partial_kw_pin_all": (_partial_kw_pinned, lambda: 2,
"shared partial(fk, a=1); partial, func, args, keywords immortal"),
"partial_builtin": (lambdas: functools.partial(max, 1), lambda: 2,
"partial(max, 1)(2) -- builtin target"),
"partial_sub": (lambdas: PartialSub(f, 1), lambda: 2, "partial subclass(f, 1)(2)"),
"partial_opcall": (lambdas: functools.partial(f, 1), lambda: 2,
"operator.call(p, 2) [loop uses operator.call]"),
"closure": (lambdas: make_closure(1), lambda: 2, "closure over a cell, not pinned"),
"closure_pin": (_closure_pinned, lambda: 2, "closure, cells immortal when shared"),
"slotted_call": (lambdas: Slotted(1), lambda: 2, "slotted instance with __call__"),
"itemgetter": (lambdas: operator.itemgetter(1), lambda: (1, 2, 3), "itemgetter(1)(tup)"),
"attrgetter": (lambdas: operator.attrgetter("a"), lambda: Slotted(2), "attrgetter('a')(obj)"),
"methodcaller": (lambdas: operator.methodcaller("count", 1), lambda: (1, 2, 3),
"methodcaller('count', 1)(tup)"),
"methodcaller_kw": (lambdas: operator.methodcaller("meth", b=1), lambda: Slotted(2),
"methodcaller('meth', b=1)(obj)"),
"bound_py": (_bound_py, lambda: 2, "bound method of a Python class instance"),
"slotted_call_pin": (_slotted_pinned, lambda: 2, "slotted __call__, instance immortal when shared"),
"bound_py_pin": (_bound_py_pinned, lambda: 2, "bound Python method, instance + method immortal when shared"),
"bound_c": (_bound_c, lambda: 2, "bound builtin method {}.get"),
}
defworker(name, shared_obj, seconds, barrier, rates, slot):
factory, argf, _=VARIANTS[name]
target=shared_objifshared_objisnotNoneelsefactory(False)
arg=argf()
clock=time.perf_counteropcall=operator.callbarrier.wait()
calls=0start=clock()
end=start+secondsifname=="partial_opcall":
whileTrue:
for_inrange(1000):
opcall(target, arg)
calls+=1000now=clock()
ifnow>=end:
breakelse:
whileTrue:
for_inrange(1000):
target(arg)
calls+=1000now=clock()
ifnow>=end:
breakrates[slot] =calls/ (now-start)
defmeasure(name, shared, threads, seconds):
factory=VARIANTS[name][0]
shared_obj=factory(True) ifsharedelseNonerates= [0.0] *threadsbarrier=threading.Barrier(threads)
ts= [threading.Thread(target=worker, args=(name, shared_obj, seconds, barrier, rates, i))
foriinrange(threads)]
fortints:
t.start()
fortints:
t.join()
returnsum(rates)
defmain():
args=sys.argv[1:]
seconds, threads=1.0, [1, 24]
if"--seconds"inargs:
i=args.index("--seconds")
seconds=float(args[i+1])
delargs[i:i+2]
if"--threads"inargs:
i=args.index("--threads")
threads= [int(x) forxinargs[i+1].split(",")]
delargs[i:i+2]
ifargs== ["LIST"]:
fork, vinVARIANTS.items():
print(f"{k:20}{v[2]}")
returnnames=list(VARIANTS) ifargs== ["all"] elseargsver=f"{sys.version_info.major}.{sys.version_info.minor}"print(sys.version.replace("\n", " "), "| cpus:", os.cpu_count(),
"| GIL:", sys._is_gil_enabled())
print(f"{'variant':20}{'shape':6} "+" ".join(f"{n:>3}T calls/s"forninthreads)
+" speedup")
out_path=os.path.join(os.path.dirname(os.path.abspath(__file__)), f"results_{ver}.jsonl")
withopen(out_path, "a") asout:
fornameinnames:
forshapein ("shared", "own"):
rates= [measure(name, shape=="shared", n, seconds) forninthreads]
speed=rates[-1] /rates[0]
print(f"{name:20}{shape:6} "+" ".join(f"{r:14,.0f}"forrinrates)
+f" {speed:6.2f}x", flush=True)
out.write(json.dumps({"py": ver, "variant": name, "shape": shape,
"threads": threads, "rates": rates,
"speedup": speed, "seconds": seconds}) +"\n")
if__name__=="__main__":
main()
How we found it
A free-threaded web framework runs 24 worker threads that share module-level
callables. Its per-workload scale-out ledger (1/8/24 threads against a
no-shared-state control) showed every path that called a long-lived functools.partial dropping from "scales" on 3.14 to below 1x on 3.15. The
framework's workaround was to replace every long-lived shared partial with a
closure whose cells are pinned at construction.
Bug report
I was migrating some of my "massively threaded" 3.14t code to 3.15t and my performance agent saw everything collapse under concurrency after the python version change.
Turns out a new 3.15-only fix (?) for a
partialapplication is now over-locking everything. For my own use I just migrated away from usingfunctools.partialeverywhere now, but ... FYI i guess.agent report:
=========================
Free-threaded 3.15: a
functools.partialshared between threads no longer scales (it slows down as you add threads)Summary
On the free-threaded build, CPython 3.15 calls to a
functools.partialstopscaling across threads when the partial (or just its target function) is shared
between threads. On a 10-core Apple M1 Max, 24 threads calling one module-level
partial(f, 1)reach 0.20x the single-thread rate, so total throughput dropswhen threads are added. 3.14.3t reaches 6.37x on the same machine. The
slowdown shows up at 2 threads (0.34x). The cause is the use-after-free fix for
gh-154189 (PR gh-154508, 3.15 backport gh-154652, first shipped in 3.15.0rc1). It
made
partial_vectorcalltake and drop strong references topto->fn,pto->argsandpto->kwon every call. In the free-threaded build anincref/decref from a thread that does not own the object is an atomic
read-modify-write on that object's
ob_ref_shared. Every calling thread now writesthe same three cache lines (the target function, the bound-args tuple, and the
keywords dict, which partial allocates even when empty) six times per call. In
3.14 the call path takes no references on these objects: it reads them borrowed,
and the callee frame uses stackrefs that skip deferred-refcount and immortal
objects. Giving each thread its own partial does not help when the target
function is shared (0.24x). Making the partial, its function and its args tuple
immortal also does not help, because the per-partial empty keywords dict is still
refcounted (0.20x). Scaling comes back only when all three (fn, args, keywords)
are immortal (5.43x). The source of
partial_vectorcallis the same in v3.15.0(tagged 2026-10-09) and
main. The 3.14 branch did not take the backport.Environment
3.14.3 free-threading build (main, Mar 21 2026, 17:04:12) [Clang 22.1.1 ](pyenv source build)3.15.0rc3 free-threading build (main, Oct 3 2026, 00:55:19) [Clang 22.1.3 ](python-build-standalone, installed by uv)sys._is_gil_enabled() == False)macOS-14.8.2-arm64-arm-64bit-Mach-Oos.cpu_count() == 10, 128-byte cache linenice -n 1024 threads on 10 cores is deliberate oversubscription (it matches a 24-worker
server). The no-shared-state control loop therefore tops out at about 5.5-6.7x on
this machine, not 24x.
Results
The minimal repro is below. Each cell is aggregate calls/s over 1 s per
measurement. Speedup = calls/s at N threads ÷ calls/s at 1 thread, for the same
shape.
partial(f, 1)for all threadspartial(f, 1)per thread,fsharedThread sweep (0.5 s per point):
f)On 3.15, 1-thread cost is unchanged. An uncontended atomic is cheap. The cost is
all in contention.
Variants that isolate the mechanism
These come from
probe_variants.py(Appendix B), with 1 and 24 threads at 1 seach. "shared" means one object built by the main thread and called by all
threads. "own" means each worker builds its own object from the same recipe.
Every thread passes its own argument object. Pinning is done with the exported
_Py_SetImmortalthroughctypes.partial(f, 1)(2)partial(max, 1)(2)(builtin target)maxshared in both shapes(f, 1)(2)operator.call(p, 2)p.funcimmortalp.keywordsimmortalpartial(fk, a=1)(2)(stored keyword)partial(fk, a=1), everything pinned¹ In the pinning rows, the "own" partials are built after the shared run has
already made
f(orfk) immortal, so they show that a private partial over animmortal target scales.
What the table shows:
not the problem. Own partials over own functions scale on 3.15 (6.42x).
the args tuple and the keywords dict. All three have to be immortal before 3.15
scales again. Pinning any one or two of them still leaves 0.17-0.21x.
immortal (0.83x / 0.67x). That is a separate, older bottleneck: a per-call lock
on the shared keywords dict. See Keyword partials.
Minimal reproducer
Standard library only. Save it as
repro_partial_shared.pyand run it with afree-threaded interpreter:
Output on the machine above:
The probe keeps the shared object in a local variable, so only the call is
measured. It does not measure a
LOAD_GLOBALof a non-deferred object, whichcarries its own refcount traffic.
Root cause
The change
gh-154189, "use-after-free:
functools partial", was reported 2026-07-19 against
main(free-threading debugTSan build). It is a single-threaded re-entrancy bug.
partial_vectorcallheld a raw
PyObject **intopto->args. In the keyword-merge path, a callkeyword's
__hash__could callpartial.__setstate__on the same partial, whichfreed the old args tuple while the C frame was still using it.
The fix is PR gh-154508, merged to
main2026-07-24 as1a742d4,and backported to 3.15 by gh-154652
(
6246342,2026-07-27). The NEWS entry reads:
Which releases have it:
Present in rc2, rc3 and v3.15.0 final (tagged 2026-10-09; the
partial_vectorcallsource is byte-identical to rc3) and inmain.reviewer agreed a backport was not worth the effort ("I don't think PR needs to
be backported to 3.14." / "I agree; it doesn't look worth the effort."). v3.14.8
has the old code.
The issue and PR do not discuss free-threaded performance, and the PR includes no
benchmark.
Before: v3.14.3
Modules/_functoolsmodule.c,partial_vectorcall(positional path)All reads are borrowed. The call path performs no refcount writes on
pto->fn,pto->argsorpto->kw.After: v3.15.0rc3 / v3.15.0 / main,
partial_vectorcallThe three
Py_NewRef/Py_DECREFpairs run on every call, including thepositional-only fast paths, where no user code runs between reading the fields and
making the call. Keyword merging is the only place the re-entrancy in gh-154189
can happen.
Mechanism
In the free-threaded build each object has an owning thread (
ob_tid), anon-atomic
ob_ref_localthat only the owner writes, and an atomicob_ref_sharedthat every other thread writes (Include/refcount.h, 3.15):Py_DECREFfrom a non-owner goes to_Py_DecRefShared(Objects/object.c). Thatfunction runs a compare-and-swap loop on
ob_ref_shared, and the loop retriesunder contention.
A long-lived partial and its
fn,argsandkware owned by the thread thatcreated them, usually the importing or main thread. Every worker thread is
therefore a non-owner. So each call on 3.15 does 3 atomic adds and 3 CAS loops on
three objects that every thread touches. Each read-modify-write needs exclusive
ownership of the cache line, so the line holding each object header moves between
cores on every call. The header also holds
ob_type,ob_tidandob_ref_local, which every caller reads, so those reads miss too. Throughput isthen limited by how fast the coherence fabric can move those lines, not by the
number of cores. The aggregate settles around 3-5 M calls/s at any thread count:
0.34x at 2 threads, about 0.2x from 8 threads up.
Why own partials over a shared function still collapse.
fnis shared, soPy_NewRef(pto->fn)writes the same line from every thread. The per-thread argstuple and keywords dict are private, so the slowdown is a little smaller (0.24x
vs 0.20x), but it is the same effect. Own partials over per-thread functions
touch only private objects and scale (6.42x).
Why 3.14 did not have this. The callee side never writes the shared refcounts.
_PyEval_Vectortakes its references withPyStackRef_FromPyObjectNew:In the free-threaded build that is a no-op for immortal objects (small ints) and
for deferred-refcount objects. Top-level functions and methods are created with
deferred refcounting (
Objects/funcobject.c:_PyObject_SetDeferredRefcountwhenthe code is not
CO_NESTED, or isCO_METHOD).Py_NewRefdoes not check thedeferred bit, so on 3.15 the same deferred
fgets an atomic increment on everycall. That is why a plain shared
f(…)call scales on both versions (6.0x /6.7x) while
partial(f, 1)does not on 3.15.Why making partial + fn + args immortal did not help.
partial_newalwaysstores a dict in
pto->kw. With no keywords it is a fresh, empty dict per partial,owned by the creating thread. 3.15 increfs and decrefs that dict on every call even
though it is empty. Pinning the partial object itself does nothing, because the
call path never touches the partial's own refcount. Only pinning fn, args and
the keywords dict brings scaling back (5.43x). The same reasoning covers a
partial of a builtin (
max):PyCFunctionobjects are neither immortal nordeferred, and the 3.15 path increfs them.
Is it intended?
The correctness fix is intended: a real use-after-free reachable from pure Python.
The free-threaded cost looks unintended. Nothing in the issue, the PR review or the
NEWS entry mentions it. The PR has no benchmark, and the 3.14 backport was skipped
for effort, not because of any trade-off. The fix also guards less than its cost
suggests:
__setstate__safe.Py_NewRef(pto->args)loads the field and increments itlater, so another thread can free the tuple in between. Free-threaded CPython 3.14: concurrent functools.partial.__setstate__ crashes with repr(), calls, or another __setstate__ #157841 (open,
"concurrent functools.partial.setstate crashes with repr(), calls, or
another setstate") reports call+
__setstate__still crashing onmainwiththis fix in place.
the CLA was not signed), took
Py_BEGIN_CRITICAL_SECTION(self)inpartial_vectorcallandpartial_callon every call. That design would replacethree contended atomics with a contended per-object mutex. For a shared partial
it would be at least as serialising. The existing per-call critical section on
the keywords dict already holds keyword partials to 0.67x with everything
immortal, below.
No open issue tracks the performance regression. Searches of the cpython tracker
for partial plus free-threading, scaling or contention since 2026-07 found only the
crash reports above.
Suggested upstream fixes
__setstate__is the only writer offn,argsandkw. It is called by pickleand
copyon a freshly built object, almost never on a live, shared one. The fixshould make that rare writer pay and keep the per-call reader free of shared
writes. In rough order of preference:
Py_NewRef. CPython already has the machinery:_PyCStackRefwith_PyThreadState_PushCStackRefNew/_PyThreadState_PopCStackRef(
Include/internal/pycore_stackref.h).typeobject.c'scall_methoduses itfor exactly this shape: it holds the looked-up
__call__function for theduration of a call without touching that shared function's refcount. That is
why a per-thread instance of a class with
__call__scales on 3.15 even thoughevery thread calls the same
__call__function.PyStackRef_FromPyObjectNewcosts nothing for deferred and immortal objects, and the GC sees C stack refs,
so a deferred object cannot be collected under the call. For
fnthis isenough on its own when
fnis a top-level function (already deferred). Forargsandkw,partial_new/partial_setstatecould enable deferredrefcounting on the tuple and dict they store, and stop allocating a fresh empty
dict when there are no keywords (keep
NULL, or share one immortal emptysentinel internally). Result: correctness against same-thread re-entrancy is
preserved, and the call does no shared writes.
references only on the keyword-merge path. That path is where user code
(
__hash__/__eq__of call keywords,PyDict_Copy,PyDict_SetItem) runsbefore the call. This is the narrowest fix for the reported reproducer. It
should be paired with (3) to cover a callee that re-enters
__setstate__on itsown partial.
__setstate__, retire the oldfn,argsandkwwith_PyObject_XDecRefDelayed(QSBR) in the free-threaded build. That handlescross-thread readers, the Free-threaded CPython 3.14: concurrent functools.partial.__setstate__ crashes with repr(), calls, or another __setstate__ #157841 shape. It does not fully cover same-thread
re-entrancy, because the calling thread can report a quiescent state while its
call is still running (for example, when it detaches for blocking I/O).
Alternatively, keep the displaced objects alive on the partial until
tp_dealloc. That is simple, costs nothing per call, and only matters for therare object that is re-stated.
writer, and publish a coherent state the reader can load without writing a
shared line: an immutable state record holding
fn,args,kwandphcount, read with an acquire load into a C stack ref and retired via QSBR.An atomic pointer plus
_Py_TryIncrefCompareon a non-deferred state objectwould bring back the same contention this report measures.
Keyword partials: a separate, older bottleneck
A partial with stored keywords serialises when shared, on both versions, even
with everything immortal: 0.83x on 3.14.3t, 0.67x on 3.15.0rc3t.
partial_call, which doesPyDict_Copy(pto->kw). That copy locks the shared source dict.pto->kwinsidePy_BEGIN_CRITICAL_SECTION(keyword_dict)on every call. That critical sectionwas added for
PyDict_Nextsafety by Missing critical sections forPyDict_Nextcalls in_functoolsmodule.c#145446.Both are per-call acquisitions of the same per-object mutex. The lock exists
because
p.keywordsreturns the live dict and user code can mutate it. Acontention-free design would snapshot the keywords into immutable
(kwnames tuple, values)at construction and in__setstate__. That is abehaviour change for code that mutates
p.keywordsin place, which is probablyrare and undocumented.
Workarounds for users (until fixed)
Measured on 3.15.0rc3t, shared object called by 24 threads:
partial(f, x)def p(b): return f(CONST, b)PyUnstable_SetImmortal(3.15 public unstable C API) or_Py_SetImmortal; leaks by design.COPY_FREE_VARSincrefs each shared cell per call, on both versions.__call__, instance pinned__call__, not pinnedselfis increfed into the callee frame per call, on both versions.fn,argsandkeywordsall pinnedGeneral rule on free-threaded CPython: a long-lived callable shared by many
threads scales only if every object the call path increfs is immortal or deferred
and reached through stackrefs. Bound
self, closure cells and, on 3.15, partialfields are increfed per call.
Appendix A: other common callables (shared vs own, 1 → 24 threads)
operator.itemgetter(1)(tup)operator.attrgetter("a")(obj)operator.methodcaller("count", 1)(tup)operator.methodcaller("meth", b=1)(obj){}.getobj.meth, not pinnedself)__call__, not pinnedself)functools.partial(positional)functools.partial(keyword)Only
functools.partialregressed between 3.14 and 3.15.methodcallerimprovedin 3.15. The 3.14 collapse even with per-thread objects is probably the per-call
reference to the shared method descriptor found by attribute lookup; that was not
investigated further. The bound-method, slotted
__call__and closure collapsesare older and the same on both versions. Each comes from a per-call incref of a
shared, non-immortal, non-deferred object (
selfor a cell), and pinning thatobject fixes it.
Appendix B: the variant probe
Run it as
python3.15t probe_variants.py <variant ...>(LISTprints the names).Pinning variants permanently immortalise the target function, so run each one in a
fresh process, as was done for the tables above.
How we found it
A free-threaded web framework runs 24 worker threads that share module-level
callables. Its per-workload scale-out ledger (1/8/24 threads against a
no-shared-state control) showed every path that called a long-lived
functools.partialdropping from "scales" on 3.14 to below 1x on 3.15. Theframework's workaround was to replace every long-lived shared partial with a
closure whose cells are pinned at construction.
=========================
CPython versions tested on:
3.15
Operating systems tested on:
macOS