Skip to content

Freeze/thaw poisons pooled localhost connection on non-SnapStart functions → SendRequest/ConnectionReset with no retry; no config to disable pooling #830

Description

@pbuntine

Summary

On a regular (non-SnapStart) on-demand function, the execution environment freeze/thaw between invocations can leave the adapter's pooled keep-alive connection to the app server in a reset state. The next invocation reuses that dead connection, and because fetch_response forwards the request exactly once with no retry, the caller gets a 5xx. There is currently no configuration to disable the adapter↔app connection pool outside of SnapStart.

Setup

  • Adapter: public.ecr.aws/awsguru/aws-lambda-adapter:1.0.1 (x86_64), as a Lambda extension in a container image
  • App: Next.js 16 standalone server on :3000
  • Base image: node:24-bookworm-slim
  • Front: API Gateway v2, payload format 2.0, $default route
  • Region: eu-central-1
  • Traffic: low, function scales to zero frequently
  • AWS_LWA_READINESS_CHECK_PATH=/api/health, AWS_LWA_READINESS_CHECK_PROTOCOL=tcp

Symptom

Intermittent 500 {"message":"Internal Server Error"} to the client. Adapter log on the failing invocation:

ERROR Lambda runtime invoke{requestId="7ab2a21b-..."}: 
  hyper_util::client::legacy::Error(SendRequest,
  hyper::Error(Io, Os { code: 104, kind: ConnectionReset, message: "Connection reset by peer" }))

The next request to the same execution environment succeeds. The app server does not crash — no error on its side, and it keeps serving.

Timeline (one execution environment, from CloudWatch)

15:49:43.137  new exec env starts: app "Ready", "lambda-adapter State: Ready"
15:49:43.141  req A  durationMs 760   initDurationMs 562   status success   <- cold-start request, OK
15:49:52.105  req B  durationMs 85    (warm, no init)      status success   <- reuses pooled conn, OK
15:49:55.823  req C  durationMs 2486  (warm, no init)      ERROR ConnectionReset  <- 500 to caller

Req C is ~3.6s after req B, on a warm environment. The environment was frozen in between.

Analysis

The adapter builds its client with a keep-alive pool:

https://github.com/aws/aws-lambda-web-adapter/blob/v1.0.1/src/lib.rs#L588-L600

if env::var("AWS_LAMBDA_INITIALIZATION_TYPE").as_deref() == Ok("snap-start") {
    builder.pool_max_idle_per_host(0);
} else {
    builder.pool_idle_timeout(Duration::from_secs(4));
}

pool_idle_timeout is measured with Instant/CLOCK_MONOTONIC, which does not advance while the execution environment is frozen. So a connection that has been idle far longer than 4s in wall-clock time still looks fresh to the pool and is reused — but the underlying socket was reset across the freeze/thaw. fetch_response then does:

https://github.com/aws/aws-lambda-web-adapter/blob/v1.0.1/src/lib.rs#L1012

let mut app_response = self.client.request(request).await?;

— a single attempt. The connection error propagates straight out as a runtime error → 5xx.

This is the same failure mode the SnapStart branch above guards against with pool_max_idle_per_host(0), but on-demand environments freeze/thaw as well and hit it too. Related prior reports of the SendRequest / IncompleteMessage family: #415, #295.

Request

Either (ideally both):

  1. Expose the pool control as configuration — e.g. AWS_LWA_DISABLE_CONNECTION_POOL=true (or AWS_LWA_POOL_MAX_IDLE_PER_HOST=0) so non-SnapStart functions can opt into pool_max_idle_per_host(0). For localhost the cost of a fresh connection per request is negligible.
  2. Retry the forwarded request once on connection-level errors (ConnectionReset, BrokenPipe, IncompleteMessage when zero bytes were written). These are safe to retry — no request bytes reached the app — and every mainstream HTTP client does this for pooled connections.

Workarounds tried

  • AWS_LWA_READINESS_CHECK_PROTOCOL=tcp — removes the connection the readiness probe would otherwise seed into the pool, but not request-to-request reuse (this report is with tcp already set).
  • Client-side retry at the caller — works for service-to-service callers, not for browsers hitting the function directly.

Activity

  1. bnusunny commented on Sep 1, 2026

    @bnusunny
    Contributor

    @pbuntine Thanks for the detailed report — the code references and CloudWatch timeline made
    this easy to chase.

    I tried to reproduce the freeze/thaw part on a real function and couldn't, so I'd
    like to pin down what's actually closing your socket before we pick a fix.

    Test setup: LWA 1.0.1 (LambdaAdapterLayerX86:28), zip + nodejs22.x, plain Node
    http server, READINESS_CHECK_PROTOCOL=tcp, reserved concurrency 1 so every
    invoke hits the same environment.

    • CLOCK_MONOTONIC doesn't stall across on-demand freeze/thaw. Measured
      against CLOCK_REALTIME/CLOCK_BOOTTIME over freezes of 3.5s, 34s, 92s and
      181s: monotonic lost 0–1ms every time. So pool_idle_timeout was measuring real
      elapsed time, and hyper-util: client connection pool won't clean expired connections after system sleep hyperium/hyper#3810 doesn't seem to apply here.
    • The pooled connection survived it. ~49 invocations, zero failures — one
      socket served 40 requests across 39 freeze/thaw cycles over 91.8s, gaps up to
      3.78s.
    • Node's effective keep-alive is 6s, not 5s. Node adds keepAliveTimeoutBuffer
      (default 1000ms) to the socket timeout, and Next.js standalone doesn't override
      keepAliveTimeout unless KEEP_ALIVE_TIMEOUT is set. Your req C was ~3.6s idle
      (B ended ~15:49:52.19, C started 15:49:55.82), inside both windows.

    I could only reproduce it by shortening the app's keep-alive below the adapter's 4s
    reuse window — then ~3 of 11 invokes failed, with the app closing the socket in the
    same millisecond the adapter reused it. That needs
    app_keepalive < idle_gap < 4s, which is an empty range at Node's 6s.

    Three things would help:

    1. Set KEEP_ALIVE_TIMEOUT=60000 and see whether the errors stop. That's one
      variable and it either confirms or rules out the whole app-side keep-alive class.
    2. RUST_LOG=debug on a failing invocation — specifically whether
      reuse idle connection for ("http", 127.0.0.1:3000) appears right before the error.
    3. durationMs 2486 on req C doesn't fit a stale-connection reuse, which fails in
      microseconds. That looks more like the app already had the request and the socket
      died mid-flight. Function memory size, and any sign of the Node process
      restarting, would help.

    On request 1 — I'm open to exposing the pool control, since some app servers do
    default below 4s. I'd just rather land it with a root cause we both understand.
    On 2, note hyper already retries when it's provably safe (retry_canceled_requests);
    the SendRequest error is the branch where the request was already handed off, so a
    blanket retry would make POSTs at-least-once.

  2. bnusunny commented on Sep 13, 2026

    @bnusunny
    Contributor

    the connection pool control is added in #837

  3. bnusunny commented on Sep 21, 2026

    @bnusunny
    Contributor

    The connection pool timeout control (envvar AWS_LWA_POOL_IDLE_TIMEOUT_SECONDS) is released in v1.1.0.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions