In bytearray.decode it's implementation is protected with a critical section lock
skip_optional_pos:
Py_BEGIN_CRITICAL_SECTION(self);
return_value = bytearray_decode_impl((PyByteArrayObject *)self, encoding, errors);
Py_END_CRITICAL_SECTION();
However inside we call PyUnicode_FromEncodedObject -> PyObject_GetBuffer -> bytearray_getbuffer
Where we fetch the lock again of itself.
static int
bytearray_getbuffer(PyObject *self, Py_buffer *view, int flags)
{
int ret;
Py_BEGIN_CRITICAL_SECTION(self);
ret = bytearray_getbuffer_lock_held(self, view, flags);
Py_END_CRITICAL_SECTION();
return ret;
}
Calling getbuffer is required when the bytearray has been inherited however in the standard path we can look to skip the second getbuffer and it's associated lock.
Below is come micro benchmarks on Macos and x86, I will also run pybenchmark
Single thread, ns per call (lower is better)
| Call |
Size |
arm64 main |
arm64 patched |
Speedup |
x86-64 main |
x86-64 patched |
Speedup |
ba.decode('utf-8') |
16 B |
65.0 |
32.5 |
2.00x |
122.7 |
96.4 |
1.27x |
|
256 B |
71.8 |
37.8 |
1.90x |
141.8 |
112.2 |
1.26x |
|
4 KiB |
242.9 |
205.8 |
1.18x |
357.6 |
324.5 |
1.10x |
|
64 KiB |
2,568 |
2,494 |
1.03x |
4,127 |
4,111 |
1.00x |
ba.decode('latin-1') |
16 B |
61.3 |
32.1 |
1.91x |
123.7 |
96.2 |
1.29x |
|
256 B |
72.3 |
38.4 |
1.88x |
147.3 |
119.8 |
1.23x |
|
4 KiB |
251.1 |
203.7 |
1.23x |
490.7 |
460.1 |
1.07x |
|
64 KiB |
2,462 |
2,440 |
1.01x |
6,329 |
6,306 |
1.00x |
bytes.decode('utf-8') (control) |
16 B |
32.8 |
31.9 |
1.03x |
91.4 |
91.3 |
1.00x |
str(ba, 'utf-8') (control) |
16 B |
37.7 |
37.5 |
1.01x |
122.4 |
121.2 |
1.01x |
One bytearray shared by N threads, ba.decode('utf-8'), million decodes/s across all threads (higher is better)
| Size |
Threads |
arm64 main |
arm64 patched |
Speedup |
x86-64 main |
x86-64 patched |
Speedup |
| 256 B |
1 |
13.59 |
26.68 |
1.96x |
6.55 |
8.50 |
1.30x |
|
2 |
8.98 |
16.22 |
1.81x |
5.06 |
7.15 |
1.41x |
|
4 |
7.63 |
14.56 |
1.91x |
3.64 |
5.01 |
1.38x |
|
6 |
7.30 |
14.67 |
2.01x |
2.39 |
3.24 |
1.35x |
| 64 KiB |
1 |
0.38 |
0.39 |
1.02x |
0.19 |
0.19 |
1.00x |
|
6 |
0.25 |
0.25 |
1.01x |
0.12 |
0.12 |
1.00x |
In
bytearray.decodeit's implementation is protected with a critical section lockHowever inside we call
PyUnicode_FromEncodedObject->PyObject_GetBuffer->bytearray_getbufferWhere we fetch the lock again of itself.
Calling getbuffer is required when the bytearray has been inherited however in the standard path we can look to skip the second getbuffer and it's associated lock.
Below is come micro benchmarks on Macos and x86, I will also run pybenchmark
Single thread, ns per call (lower is better)
ba.decode('utf-8')ba.decode('latin-1')bytes.decode('utf-8')(control)str(ba, 'utf-8')(control)One bytearray shared by N threads,
ba.decode('utf-8'), million decodes/s across all threads (higher is better)