Skip to content

test: add local Docker/RIE integration test harness (Level 2 tests) - #804

Merged
joeyzhao2018 merged 18 commits into
mainfrom
joey/make-integration-tests-local
Aug 27, 2026
Merged

joeyzhao2018 merged 18 commits into
mainfrom
joey/make-integration-tests-local

Conversation

@joeyzhao2018

@joeyzhao2018 joeyzhao2018 commented Jul 29, 2026 •

Copy link
Copy Markdown
Collaborator

What this is

This is the first PR in the effort to move datadog-lambda-js's core logic into dd-trace-js.
It doesn't move any code yet — it builds the safety net that every later PR will be checked against.

Per Step 2 of the migration strategy, we need three levels of test coverage. Level 2 — Docker-based integration tests — didn't exist yet: our integration tests run against real AWS Lambdas, which is too slow and too credential-bound for the dd-trace pipeline's few-minute budget. This PR adds that missing level.

What it does

Adds integration_tests_local/, which runs the container handler variants against the AWS Runtime Interface Emulator — no AWS account, no deployed lambda functions, no credentials.

For each of Node 18/20/22/24 × {cjs, esm}, it builds the image, starts it under RIE, fires 9 input events, and diffs both the return values and the normalized logs against committed snapshots. That's 80 golden snapshots (72 return values + 8 log files) capturing exactly what the library does today. We actually enhance this part to remove replications in the snapshots. so now it's much cleaner.

Those goldens are the point. Once we start moving code into dd-trace-js, they're the oracle that tells us whether migrated code still behaves identically — so they're captured from the current implementation before any migration happens, and they're never regenerated wholesale from migrated code.

Also adds a GitHub Actions workflow running all five runtime legs on every PR.

Notes for review

Most of the diff is the 80 snapshot files. The parts worth actual review are run.sh and normalize.sh.

Normalization is the load-bearing piece. It decides which volatility gets erased before diffing (timestamps, request IDs, durations, trace IDs). Erase too little and the tests flake; erase too much and the gate silently stops catching regressions. Two specific decisions:

  • Span meta/metrics keys are preserved — their stable keys and values are part of the compatibility contract.
  • Trailing-whitespace stripping is scoped to RIE's REPORT lines only, so application output stays an exact-whitespace comparison.

Node 18 and 20 legs are deliberate. dd-trace's v5.x is the only line serving those runtimes, and we ship layers for them, so they stay in the matrix for the whole migration.
(This is the same position taken when closing #791.)

Node 26 is a special case. public.ecr.aws/lambda/nodejs:26 doesn't exist yet — AWS only publishes 26-preview.* — so the base image is pinned to a specific preview tag, threaded through a runtime→image-tag indirection applied only at the docker build --build-arg (the bare major still drives image tags, container names and snapshot paths). Node 26 has no committed snapshots yet, so it's excluded from the default local run and runs in update mode
in CI, which asserts it produces exactly the expected 20 artifacts without committing them. The next PR freezes them. When AWS publishes the GA image, the Node 26 goldens need re-capturing from the pre-migration code — that's recorded in the README.

Hardening details: RIE binary pinned to v1.36 with a per-platform SHA-256 that's re-verified on every run; missing snapshots fail instead of passing silently; HTTP status is checked alongside the response body (curl -s exits 0 on a 5xx, so update mode could otherwise record error pages as goldens); diff exit codes 0/1/>1 are handled distinctly; fixed sleeps replaced with polling on both REPORT and INVOKE RTDONE, since RIE can emit them out of order.

Trying it locally

./integration_tests_local/run.sh                          # all runtimes with committed snapshots
RUNTIME_PARAM=18 VARIANT_PARAM=esm ./integration_tests_local/run.sh
UPDATE_SNAPSHOTS=true ./integration_tests_local/run.sh    # regenerate

Requires Docker. First run pulls the base images (~1 GB each) and the RIE binary.

What comes next

This is Step 2 of the strategy. The PRs that follow, roughly in order:

  1. Freeze the baseline — capture the Node 26 snapshots and tag the full 100-snapshot baseline, making all five runtime legs strict comparisons.
  2. dd-trace-js scaffolding — register an aws-lambda plugin and define a small, stable facade (wrap, sendDistributionMetric, sendDistributionMetricWithDate, getTraceHeaders) for the shim to call.
  3. Mechanical ports (Step 3) — utilities and event detection, context extractors, span inferrer and X-Ray, span pointers. Code moves as-is; specs move with it.
  4. The seams — the places copy-paste can't solve: config wiring, handler lifecycle, metrics, and cold start / DSM / AppSec.
  5. Shim conversion (Step 4) — datadog-lambda-js becomes thin and delegates to dd-trace, with the layer built from migrated logic.
  6. Release gates — goldens, ported unit tests and the real-AWS e2e suite all run against the candidate layer before anything is deleted.
  7. Cleanup and release (Step 5) — delete the old business logic, publish.

Every one of those is checked against the goldens this PR adds, and progress is tracked row-by-row in migration_parity.md.

Remove the cold_start value normalization from the local harness
pipeline. The cold->warm transition (invoke #1 cold_start:true,
invokes #2..N false) is deliberate coverage and is deterministic
locally, since proactive initialization cannot happen unless
SIMULATE_PROACTIVE_INIT=true. Snapshots regenerated with real
cold_start values.

proactive_initialization markers and the '(init: N ms)' END suffix
remain stripped: they reflect platform scheduling, not code behavior.
The real-AWS run_integration_tests.sh is intentionally untouched;
proactive-init cold_start flakes there are handled by rerun.
@joeyzhao2018
joeyzhao2018 marked this pull request as ready for review July 29, 2026 16:48
@joeyzhao2018
joeyzhao2018 requested review from a team as code owners July 29, 2026 16:48

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0df937ad33

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread integration_tests_local/run.sh Outdated
@datadog-prod-us1-6

datadog-prod-us1-6 Bot commented Jul 29, 2026 •

Copy link
Copy Markdown

Pipelines  Tests

✨ Unblock PR with BitsAI

⚠️ Warnings

❌ Your PR has failed checks. Please review the issues below and take necessary action before merging.

🚦 1 Pipeline job failed

DataDog/datadog-lambda-js | integration test (node18) — 🔧 Needs a code fix, caused by this PR

View more details · View in GitLab

ℹ️ Info

🔄 Datadog auto-retried 1 job - 1 passed on retry View in Datadog

Useful? React with 👍 / 👎

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: 1252d0b | Docs | View more details | Give us feedback!

joeyzhao2018 and others added 12 commits July 29, 2026 13:06
Adds a workflow that runs the RIE-based harness from
integration_tests_local/ on every PR and main push: one job per Node
major version (18/20/22/24), each building and testing both container
variants against the real AWS Lambda base images. No AWS credentials
needed. Uses linux/amd64 images on the x86_64 runners.
Snapshot files are now named <variant>_node<runtime>[...] (e.g.
esm_node18.log) instead of container-<variant>_node<runtime>[...].
The docker image tags and function names keep the container- prefix;
only the snapshot paths changed.
* Add Node26 runtime

* update snapshots

* fix integration test path

* node version compatibilty

* filtering snapshots

* cleanup
Prepare the docker/RIE golden harness to serve as the compatibility oracle for
the dd-trace-js migration, before any migrated code runs against it.

Oracle correctness:
- normalize.sh: `run_id_filter` defaulted to 's/$/^/', which is not a no-op --
  it appended a literal '^' to every output line, and all eight committed log
  goldens carried one. Setting RUN_ID would have flipped the filter and broken
  them all at once. Default is now a genuine pass-through, and the goldens are
  updated by that byte-provable transform only (strip the trailing '^', strip
  RIE's trailing REPORT tab); no golden was regenerated by running code.
- normalize.sh: stop wiping span `meta`/`metrics` keys -- their stable keys and
  values are part of the compatibility contract. (The filter was already dead:
  the goldens are pretty-printed and the pattern could not cross newlines.)
- Trailing-blank stripping is scoped to REPORT lines, so application output
  remains an exact-whitespace oracle.
- run.sh: drop `diff -w` from the sorted log comparison. It hid whitespace-only
  drift, e.g. a migrated emitter serializing '{"a":1}' where today it emits
  '{"a": 1}'. The tab it existed for is now removed at the source.

Reliability:
- Pin the RIE binary to v1.36 with a per-platform SHA-256, re-hashed on every
  run so a stale cached download cannot survive.
- Fail on missing snapshots unless UPDATE_SNAPSHOTS=true; validate
  RUNTIME_PARAM/VARIANT_PARAM/UPDATE_SNAPSHOTS; reject argv, which was silently
  ignored.
- Capture the HTTP status alongside the body: `curl -s` exits 0 on a 5xx, so
  update mode could have recorded error pages as goldens.
- Handle `diff` exit codes 0/1/>1 distinctly; a diff *error* previously read as
  success.
- Replace fixed sleeps with polling: a TCP readiness probe (never a POST, which
  would consume the cold-start invocation) and a wait for both REPORT and
  INVOKE RTDONE per event, since RIE can emit RTDONE after REPORT.
- Guard cleanup() against the blank array slots left by "${array[@]/$cid}".

Node 26:
- Add the leg to the harness and CI matrix. Node 26 ships to customers but has
  zero L2 coverage.
- `public.ecr.aws/lambda/nodejs:26` does not exist yet (preview only), so pin
  the multi-arch tag 26-preview.2026.08.21.22 through a runtime->image-tag
  indirection applied only at the --build-arg; $node_version also feeds image
  tags, container names and snapshot paths, which must keep the bare major.
- Node 26 is excluded from the default runtime list and runs in update mode in
  CI until PR 2 commits its snapshots. CI asserts it produces exactly 20
  artifacts and that no committed snapshot is ever mutated.
- Record the GA re-capture trigger: when AWS publishes the bare image, the Node
  26 goldens must be re-captured from the pinned pre-migration ref.

Node 18 and 20 legs are retained deliberately (C18); v5.x is the only dd-trace
line serving those runtimes.

Verified: all eight strict legs (18/20/22/24 x cjs/esm) pass byte-exact; the
node26 update-mode leg produces exactly 20 artifacts; negative tests for
missing snapshots and whitespace-only drift both fail as intended.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@joeyzhao2018 joeyzhao2018 changed the title chore: make integration tests local test: add local Docker/RIE integration test harness (Level 2 tests) Aug 27, 2026
The local harness committed one golden per (variant, runtime, event): 72
return values and 8 logs. All 72 return values were byte-identical, and the
8 logs held only one genuinely runtime-specific line each. 80 files carried
3 distinct expectations, and 80 files had to be re-read whenever one of them
legitimately changed.

Collapse them to 3 shared goldens covering all five runtimes and both
variants:

  snapshots/return_values/default.json
  snapshots/logs/cjs.log
  snapshots/logs/esm.log

Two changes make sharing safe rather than lossy:

- Assert `runtime:nodejsNN.x` on every invocation with the major actually
  under test, then collapse it to `nodejsXX.x`. Without the assertion,
  sharing the golden would silently stop checking that the library reports
  the runtime it is running on.
- Drop the runtime major from AWS_LAMBDA_FUNCTION_NAME. It propagates into
  service, resource, resource_names, functionname, function_arn,
  _dd.base_service and _dd.tags.process, so embedding the runtime there made
  every runtime's golden differ in ~100 lines of pure fixture naming, which
  buried the one line that was actually runtime-specific.

Divergence is expressed by adding an override file, never by loosening a
comparison:

  snapshots/logs/${variant}_node${major}.log
  snapshots/return_values/${variant}_node${major}_${event}.json

An override wins for its leg alone, so a real behavioral difference appears
as a new file in review, where widening a normalization filter to absorb it
would appear as nothing. In update mode a leg that disagrees with an
existing shared golden now fails instead of overwriting it; otherwise the
last runtime to run would define the expectation for all of them. That
agreement check uses the same sort-insensitive comparison as strict mode,
since RIE's platform logging races application output.

Node 26 consequently needs no snapshots of its own: its normalized output
matches the shared goldens. That removes the update-mode CI override and the
node26 baseline assertion, making all five workflow legs strict.

The consolidation is auditable rather than a blind re-capture: reverse-
substituting the runtime major and function name into the 3 shared files
reproduces all 72 original return-value goldens byte-for-byte and all 8
original per-runtime logs with an identical set of lines. The only residual
differences are the ordering of RIE's own platform lines (INVOKE RTDONE,
START/END/REPORT), which is already compared order-insensitively.

Also fixes a latent normalization hole this exposed: function_arn is
normalized by a pattern whose character class excludes `_`, so names ending
in `_node18` halted the match and the goldens recorded a partially
normalized `"function_arn":"XXXX_node18"`. With the suffix gone the ARN
normalizes completely.

Verified with a full strict run: 5 runtimes x 2 variants, 110 assertions,
no failures.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
run_id_filter="s/${RUN_ID}/XXXX/g"
fi

node "$repo_dir/integration_tests/parse-json.js" |

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've found in a couple of the layers (i think node and ruby) that when I run integration tests locally sometimes I get cold start logs and sometimes I don't, so olivier and I have removed them (example) in these repos. You might want to consider doing the same here, or just leave it unless it becomes a problem.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

although now that I read further not sure if we can do that given the proactive init simulation

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

When you previously run integration tests "locally", it actually still uses real lambdas. For real lambda, the platform decides when to run init, so sometimes it proactively initializes the sandbox and cold_start flips to false.
This PR is real 'local' using docker and no need for aws credentials whatsoever. And under RIE there's no platform scheduler — we start the container and drive all 9 invocations ourselves, so the cold→warm transition is deterministic.

As for the SIMULATE_PROACTIVE_INIT mode. It is a manual diagnostic, not part of the test gate. In other words, it's not used (for now). To make it more clear, i updated the README for its section.

joeyzhao2018 and others added 2 commits August 27, 2026 09:38
The readiness poll waits for both a REPORT and an INVOKE RTDONE per
invocation. RIE's managed-instances path — the one SIMULATE_PROACTIVE_INIT
switches into to get eager init — emits REPORT but never RTDONE, so the poll
could never be satisfied and the mode timed out at 9/0 instead of producing
a usable diff.

Verified directly against both RIE paths with a single invocation:

  default RIE          REPORT=1  INVOKE_RTDONE=1
  managed-instances    REPORT=1  INVOKE_RTDONE=0

Require RTDONE only on the default path. The mode now runs to completion and
surfaces exactly what it is meant to: cold_start flips to false on the first
invocation, plus managed-instances platform noise (structured bootstrap logs
and a different emulated region in ARNs).

Also correct the proactive_initialization row in migration_parity.md, which
claimed RIE cannot produce a proactively-initialized sandbox. It can: raw
un-normalized logs from a managed-instances run with a 15s init->invoke gap
contain "proactive_initialization":1, proactive_initialization:true and
cold_start:false. The feature is unprotected because normalization drops
those markers and the goldens come from immediate-invoke runs, not because
the sandbox state is unreachable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The section documented the mechanism in detail but never said what its
status is, so a reader could reasonably assume it is part of the suite. It
is not set in CI, not set in the default run, and nothing gates on it; a run
with it is expected to differ from the goldens rather than match them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

@rithikanarayan rithikanarayan left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@joeyzhao2018
joeyzhao2018 merged commit dabb8e7 into main Aug 27, 2026
53 of 55 checks passed
@joeyzhao2018
joeyzhao2018 deleted the joey/make-integration-tests-local branch August 27, 2026 17:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants