Skip to content

ssl: PySSL_SetError leaves the ERR_LIB_SYS entry on the OpenSSL error queue, so every later SSL I/O on that thread fails with the same OSError #158313

Description

@liaoruoxing

Bug report

Bug description

Since gh-127257 (GH-127361; backported to 3.12 in GH-127905 and to 3.13 in GH-127812), PySSL_SetError maps an ERR_LIB_SYS entry on the OpenSSL error queue to OSError:

if (ERR_GET_LIB(e) == ERR_LIB_SYS) {
    // A system error is being reported; reason is set to errno
    errno = ERR_GET_REASON(e);
    return PyErr_SetFromErrno(PyExc_OSError);
}

The same block appears in both the SSL_ERROR_SYSCALL and the SSL_ERROR_SSL case. Both return early, before the ERR_clear_error() that every other exit of PySSL_SetError goes through (and that the neighbouring e == 0 branch of SSL_ERROR_SYSCALL calls explicitly). The entry therefore stays on the thread's error queue.

OpenSSL's SSL_get_error() inspects that queue before anything else — the man page says "The current thread's error queue must be empty before the TLS/SSL I/O operation is attempted, or SSL_get_error() will not work reliably" — so the next SSL_read_ex()/SSL_write_ex() on that thread that does not complete immediately (would-block, EOF, …), on any connection, is reported as SSL_ERROR_SYSCALL/SSL_ERROR_SSL; PySSL_SetError takes the same early return, raises the same stale OSError, and again does not clear. The thread stays poisoned for the life of the process.

Nothing in the ssl module recovers from this. What sometimes does is a coincidence: OpenSSL's state_machine() calls ERR_clear_error() on entry, so a post-handshake message (e.g. a TLS 1.3 NewSessionTicket) processed inside a later read cleans up by accident. That is why this is hard to see in small tests and easy to hit in a long-running client that handshakes connections on one thread and reads them on another.

Reproducer

Loopback only; needs the openssl CLI for a throwaway certificate.

"""Two TLS connections on loopback. The peer resets the first; sending on it raises ConnectionResetError,
which is correct. The second connection is healthy and idle, so recv() on it should wait for the timeout.
Instead it raises ConnectionResetError at once, and keeps doing so on every call.

Session tickets are switched off on purpose: TLS 1.3 sends them after the handshake and the client processes
them inside its next read, which runs OpenSSL's state machine, whose entry point calls ERR_clear_error() and
would hide the bug by coincidence."""
import os, socket, ssl, struct, subprocess, sys, tempfile, threading, time

print(sys.version.split()[0], ssl.OPENSSL_VERSION, sys.platform)
d = tempfile.mkdtemp()
crt, key = os.path.join(d, "cert.pem"), os.path.join(d, "key.pem")
subprocess.run(["openssl", "req", "-x509", "-newkey", "ec", "-pkeyopt", "ec_paramgen_curve:prime256v1", "-nodes",
                "-keyout", key, "-out", crt, "-days", "1", "-subj", "/CN=localhost"], check=True, capture_output=True)
server_ctx = ssl.SSLContext(ssl.PROTOCOL_TLS_SERVER)
server_ctx.load_cert_chain(crt, key)
server_ctx.num_tickets = 0
client_ctx = ssl.create_default_context()
client_ctx.check_hostname, client_ctx.verify_mode = False, ssl.CERT_NONE

listener = socket.socket()
listener.bind(("127.0.0.1", 0))
listener.listen(5)
held = []

def serve():
    while True:
        conn, _ = listener.accept()
        tls = server_ctx.wrap_socket(conn, server_side=True)
        tls.settimeout(10)
        if tls.recv(16) == b"reset":
            raw = socket.socket(fileno=tls.detach())
            raw.setsockopt(socket.SOL_SOCKET, socket.SO_LINGER, struct.pack("ii", 1, 0))
            raw.close()                      # RST, no close_notify
        else:
            held.append(tls)                 # stays open and silent

threading.Thread(target=serve, daemon=True).start()

def connect(first_message):
    c = client_ctx.wrap_socket(socket.create_connection(listener.getsockname()), server_hostname="localhost")
    c.settimeout(1)
    c.send(first_message)
    return c

def show(label, op):
    t = time.time()
    try:
        r = op()
        print(f"{label}: returned {r!r} after {time.time() - t:.2f}s")
    except Exception as e:
        print(f"{label}: {type(e).__name__}({getattr(e, 'errno', None)}) after {time.time() - t:.2f}s")

healthy = connect(b"idle")
show("healthy.recv() before      ", lambda: healthy.recv(100))     # TimeoutError after 1s: correct
broken = connect(b"reset")
time.sleep(0.2)
show("broken.send() after peer RST", lambda: broken.send(b"x"))      # ConnectionResetError: correct
for i in range(3):
    show(f"healthy.recv() after, try {i + 1}", lambda: healthy.recv(100))  # expected: TimeoutError after 1s

Python 3.12.13, OpenSSL 3.5.5, Linux aarch64 (the python:3.12-slim image):

healthy.recv() before      : TimeoutError(None) after 1.00s
broken.send() after peer RST: ConnectionResetError(104) after 0.00s
healthy.recv() after, try 1: ConnectionResetError(104) after 0.00s
healthy.recv() after, try 2: ConnectionResetError(104) after 0.00s
healthy.recv() after, try 3: ConnectionResetError(104) after 0.00s

The ConnectionResetError from broken.send() is correct. The three from healthy.recv() are not: that connection is open and idle, and recv() should have waited 1 s and raised TimeoutError, as it did before.

Same script on Python 3.12.8, OpenSSL 3.5.7, macOS (predates the backport):

healthy.recv() before      : TimeoutError(None) after 1.00s
broken.send() after peer RST: SSLError(5) after 0.00s
healthy.recv() after, try 1: TimeoutError(None) after 1.00s
healthy.recv() after, try 2: TimeoutError(None) after 1.00s
healthy.recv() after, try 3: TimeoutError(None) after 1.00s

(The SSLError for the reset connection is the older, less precise report that GH-127361 improved on; the healthy connection is unaffected because that code path ends in fill_and_set_sslerror() + ERR_clear_error().)

How it was found

slack_sdk's builtin Socket Mode client reads every WebSocket connection it will ever have on one long-lived thread and handshakes new ones on another. One TLS connection reset by a network gateway put the receive thread's queue in this state. From then on every new connection was established fine, delivered one record, and failed with BrokenPipeError on the first would-block read — reconnect, repeat every 10 seconds, for three days, with the process otherwise healthy. strace showed no failing syscall at all in those cycles (the read returned EAGAIN), and reading the thread's queue through ctypes showed two entries: SSL_R_UNEXPECTED_EOF_WHILE_READING from the reset and ERR_LIB_SYS/EPIPE from OpenSSL's attempt to send the alert (tls_retry_write_records), so SSL_get_error() answered SSL_ERROR_SSL and the early return in that case was the one taken. Calling ERR_clear_error() from the receive thread recovered the client immediately.

Suggested fix

Call ERR_clear_error() before the two return PyErr_SetFromErrno(PyExc_OSError); in the ERR_LIB_SYS branches, as the e == 0 branch of SSL_ERROR_SYSCALL already does. (The broader request to clear the queue before every I/O call is gh-81891 / bpo-37710; this report is about the specific regression from GH-127361.)

CPython versions tested on

3.12.13 (affected); 3.12.8 (not affected). The same two early returns are on main today.

Operating systems tested on

Linux (aarch64, Debian-based python:3.12-slim); macOS for the unaffected comparison.

Linked PRs

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    topic-SSLtype-bugAn unexpected behavior, bug, or error

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions