Bug report
Bug description
Since gh-127257 (GH-127361; backported to 3.12 in GH-127905 and to 3.13 in GH-127812), PySSL_SetError maps an ERR_LIB_SYS entry on the OpenSSL error queue to OSError:
if (ERR_GET_LIB(e) == ERR_LIB_SYS) {
// A system error is being reported; reason is set to errno
errno = ERR_GET_REASON(e);
return PyErr_SetFromErrno(PyExc_OSError);
}
The same block appears in both the SSL_ERROR_SYSCALL and the SSL_ERROR_SSL case. Both return early, before the ERR_clear_error() that every other exit of PySSL_SetError goes through (and that the neighbouring e == 0 branch of SSL_ERROR_SYSCALL calls explicitly). The entry therefore stays on the thread's error queue.
OpenSSL's SSL_get_error() inspects that queue before anything else — the man page says "The current thread's error queue must be empty before the TLS/SSL I/O operation is attempted, or SSL_get_error() will not work reliably" — so the next SSL_read_ex()/SSL_write_ex() on that thread that does not complete immediately (would-block, EOF, …), on any connection, is reported as SSL_ERROR_SYSCALL/SSL_ERROR_SSL; PySSL_SetError takes the same early return, raises the same stale OSError, and again does not clear. The thread stays poisoned for the life of the process.
Nothing in the ssl module recovers from this. What sometimes does is a coincidence: OpenSSL's state_machine() calls ERR_clear_error() on entry, so a post-handshake message (e.g. a TLS 1.3 NewSessionTicket) processed inside a later read cleans up by accident. That is why this is hard to see in small tests and easy to hit in a long-running client that handshakes connections on one thread and reads them on another.
Reproducer
Loopback only; needs the openssl CLI for a throwaway certificate.
"""Two TLS connections on loopback. The peer resets the first; sending on it raises ConnectionResetError,
which is correct. The second connection is healthy and idle, so recv() on it should wait for the timeout.
Instead it raises ConnectionResetError at once, and keeps doing so on every call.
Session tickets are switched off on purpose: TLS 1.3 sends them after the handshake and the client processes
them inside its next read, which runs OpenSSL's state machine, whose entry point calls ERR_clear_error() and
would hide the bug by coincidence."""
import os, socket, ssl, struct, subprocess, sys, tempfile, threading, time
print(sys.version.split()[0], ssl.OPENSSL_VERSION, sys.platform)
d = tempfile.mkdtemp()
crt, key = os.path.join(d, "cert.pem"), os.path.join(d, "key.pem")
subprocess.run(["openssl", "req", "-x509", "-newkey", "ec", "-pkeyopt", "ec_paramgen_curve:prime256v1", "-nodes",
"-keyout", key, "-out", crt, "-days", "1", "-subj", "/CN=localhost"], check=True, capture_output=True)
server_ctx = ssl.SSLContext(ssl.PROTOCOL_TLS_SERVER)
server_ctx.load_cert_chain(crt, key)
server_ctx.num_tickets = 0
client_ctx = ssl.create_default_context()
client_ctx.check_hostname, client_ctx.verify_mode = False, ssl.CERT_NONE
listener = socket.socket()
listener.bind(("127.0.0.1", 0))
listener.listen(5)
held = []
def serve():
while True:
conn, _ = listener.accept()
tls = server_ctx.wrap_socket(conn, server_side=True)
tls.settimeout(10)
if tls.recv(16) == b"reset":
raw = socket.socket(fileno=tls.detach())
raw.setsockopt(socket.SOL_SOCKET, socket.SO_LINGER, struct.pack("ii", 1, 0))
raw.close() # RST, no close_notify
else:
held.append(tls) # stays open and silent
threading.Thread(target=serve, daemon=True).start()
def connect(first_message):
c = client_ctx.wrap_socket(socket.create_connection(listener.getsockname()), server_hostname="localhost")
c.settimeout(1)
c.send(first_message)
return c
def show(label, op):
t = time.time()
try:
r = op()
print(f"{label}: returned {r!r} after {time.time() - t:.2f}s")
except Exception as e:
print(f"{label}: {type(e).__name__}({getattr(e, 'errno', None)}) after {time.time() - t:.2f}s")
healthy = connect(b"idle")
show("healthy.recv() before ", lambda: healthy.recv(100)) # TimeoutError after 1s: correct
broken = connect(b"reset")
time.sleep(0.2)
show("broken.send() after peer RST", lambda: broken.send(b"x")) # ConnectionResetError: correct
for i in range(3):
show(f"healthy.recv() after, try {i + 1}", lambda: healthy.recv(100)) # expected: TimeoutError after 1s
Python 3.12.13, OpenSSL 3.5.5, Linux aarch64 (the python:3.12-slim image):
healthy.recv() before : TimeoutError(None) after 1.00s
broken.send() after peer RST: ConnectionResetError(104) after 0.00s
healthy.recv() after, try 1: ConnectionResetError(104) after 0.00s
healthy.recv() after, try 2: ConnectionResetError(104) after 0.00s
healthy.recv() after, try 3: ConnectionResetError(104) after 0.00s
The ConnectionResetError from broken.send() is correct. The three from healthy.recv() are not: that connection is open and idle, and recv() should have waited 1 s and raised TimeoutError, as it did before.
Same script on Python 3.12.8, OpenSSL 3.5.7, macOS (predates the backport):
healthy.recv() before : TimeoutError(None) after 1.00s
broken.send() after peer RST: SSLError(5) after 0.00s
healthy.recv() after, try 1: TimeoutError(None) after 1.00s
healthy.recv() after, try 2: TimeoutError(None) after 1.00s
healthy.recv() after, try 3: TimeoutError(None) after 1.00s
(The SSLError for the reset connection is the older, less precise report that GH-127361 improved on; the healthy connection is unaffected because that code path ends in fill_and_set_sslerror() + ERR_clear_error().)
How it was found
slack_sdk's builtin Socket Mode client reads every WebSocket connection it will ever have on one long-lived thread and handshakes new ones on another. One TLS connection reset by a network gateway put the receive thread's queue in this state. From then on every new connection was established fine, delivered one record, and failed with BrokenPipeError on the first would-block read — reconnect, repeat every 10 seconds, for three days, with the process otherwise healthy. strace showed no failing syscall at all in those cycles (the read returned EAGAIN), and reading the thread's queue through ctypes showed two entries: SSL_R_UNEXPECTED_EOF_WHILE_READING from the reset and ERR_LIB_SYS/EPIPE from OpenSSL's attempt to send the alert (tls_retry_write_records), so SSL_get_error() answered SSL_ERROR_SSL and the early return in that case was the one taken. Calling ERR_clear_error() from the receive thread recovered the client immediately.
Suggested fix
Call ERR_clear_error() before the two return PyErr_SetFromErrno(PyExc_OSError); in the ERR_LIB_SYS branches, as the e == 0 branch of SSL_ERROR_SYSCALL already does. (The broader request to clear the queue before every I/O call is gh-81891 / bpo-37710; this report is about the specific regression from GH-127361.)
CPython versions tested on
3.12.13 (affected); 3.12.8 (not affected). The same two early returns are on main today.
Operating systems tested on
Linux (aarch64, Debian-based python:3.12-slim); macOS for the unaffected comparison.
Linked PRs
Bug report
Bug description
Since gh-127257 (GH-127361; backported to 3.12 in GH-127905 and to 3.13 in GH-127812),
PySSL_SetErrormaps anERR_LIB_SYSentry on the OpenSSL error queue toOSError:The same block appears in both the
SSL_ERROR_SYSCALLand theSSL_ERROR_SSLcase. Both return early, before theERR_clear_error()that every other exit ofPySSL_SetErrorgoes through (and that the neighbouringe == 0branch ofSSL_ERROR_SYSCALLcalls explicitly). The entry therefore stays on the thread's error queue.OpenSSL's
SSL_get_error()inspects that queue before anything else — the man page says "The current thread's error queue must be empty before the TLS/SSL I/O operation is attempted, or SSL_get_error() will not work reliably" — so the nextSSL_read_ex()/SSL_write_ex()on that thread that does not complete immediately (would-block, EOF, …), on any connection, is reported asSSL_ERROR_SYSCALL/SSL_ERROR_SSL;PySSL_SetErrortakes the same early return, raises the same staleOSError, and again does not clear. The thread stays poisoned for the life of the process.Nothing in the ssl module recovers from this. What sometimes does is a coincidence: OpenSSL's
state_machine()callsERR_clear_error()on entry, so a post-handshake message (e.g. a TLS 1.3 NewSessionTicket) processed inside a later read cleans up by accident. That is why this is hard to see in small tests and easy to hit in a long-running client that handshakes connections on one thread and reads them on another.Reproducer
Loopback only; needs the
opensslCLI for a throwaway certificate.Python 3.12.13, OpenSSL 3.5.5, Linux aarch64 (the
python:3.12-slimimage):The
ConnectionResetErrorfrombroken.send()is correct. The three fromhealthy.recv()are not: that connection is open and idle, andrecv()should have waited 1 s and raisedTimeoutError, as it did before.Same script on Python 3.12.8, OpenSSL 3.5.7, macOS (predates the backport):
(The
SSLErrorfor the reset connection is the older, less precise report that GH-127361 improved on; the healthy connection is unaffected because that code path ends infill_and_set_sslerror()+ERR_clear_error().)How it was found
slack_sdk's builtin Socket Mode client reads every WebSocket connection it will ever have on one long-lived thread and handshakes new ones on another. One TLS connection reset by a network gateway put the receive thread's queue in this state. From then on every new connection was established fine, delivered one record, and failed with
BrokenPipeErroron the first would-block read — reconnect, repeat every 10 seconds, for three days, with the process otherwise healthy.straceshowed no failing syscall at all in those cycles (the read returnedEAGAIN), and reading the thread's queue through ctypes showed two entries:SSL_R_UNEXPECTED_EOF_WHILE_READINGfrom the reset andERR_LIB_SYS/EPIPEfrom OpenSSL's attempt to send the alert (tls_retry_write_records), soSSL_get_error()answeredSSL_ERROR_SSLand the early return in that case was the one taken. CallingERR_clear_error()from the receive thread recovered the client immediately.Suggested fix
Call
ERR_clear_error()before the tworeturn PyErr_SetFromErrno(PyExc_OSError);in theERR_LIB_SYSbranches, as thee == 0branch ofSSL_ERROR_SYSCALLalready does. (The broader request to clear the queue before every I/O call is gh-81891 / bpo-37710; this report is about the specific regression from GH-127361.)CPython versions tested on
3.12.13 (affected); 3.12.8 (not affected). The same two early returns are on
maintoday.Operating systems tested on
Linux (aarch64, Debian-based
python:3.12-slim); macOS for the unaffected comparison.Linked PRs