Repository navigation
test: add local Docker/RIE integration test harness (Level 2 tests) - #804
Conversation
Remove the cold_start value normalization from the local harness pipeline. The cold->warm transition (invoke #1 cold_start:true, invokes #2..N false) is deliberate coverage and is deterministic locally, since proactive initialization cannot happen unless SIMULATE_PROACTIVE_INIT=true. Snapshots regenerated with real cold_start values. proactive_initialization markers and the '(init: N ms)' END suffix remain stripped: they reflect platform scheduling, not code behavior. The real-AWS run_integration_tests.sh is intentionally untouched; proactive-init cold_start flakes there are handled by rerun.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 0df937ad33
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Adds a workflow that runs the RIE-based harness from integration_tests_local/ on every PR and main push: one job per Node major version (18/20/22/24), each building and testing both container variants against the real AWS Lambda base images. No AWS credentials needed. Uses linux/amd64 images on the x86_64 runners.
Snapshot files are now named <variant>_node<runtime>[...] (e.g. esm_node18.log) instead of container-<variant>_node<runtime>[...]. The docker image tags and function names keep the container- prefix; only the snapshot paths changed.
* Add Node26 runtime * update snapshots * fix integration test path * node version compatibilty * filtering snapshots * cleanup
Prepare the docker/RIE golden harness to serve as the compatibility oracle for
the dd-trace-js migration, before any migrated code runs against it.
Oracle correctness:
- normalize.sh: `run_id_filter` defaulted to 's/$/^/', which is not a no-op --
it appended a literal '^' to every output line, and all eight committed log
goldens carried one. Setting RUN_ID would have flipped the filter and broken
them all at once. Default is now a genuine pass-through, and the goldens are
updated by that byte-provable transform only (strip the trailing '^', strip
RIE's trailing REPORT tab); no golden was regenerated by running code.
- normalize.sh: stop wiping span `meta`/`metrics` keys -- their stable keys and
values are part of the compatibility contract. (The filter was already dead:
the goldens are pretty-printed and the pattern could not cross newlines.)
- Trailing-blank stripping is scoped to REPORT lines, so application output
remains an exact-whitespace oracle.
- run.sh: drop `diff -w` from the sorted log comparison. It hid whitespace-only
drift, e.g. a migrated emitter serializing '{"a":1}' where today it emits
'{"a": 1}'. The tab it existed for is now removed at the source.
Reliability:
- Pin the RIE binary to v1.36 with a per-platform SHA-256, re-hashed on every
run so a stale cached download cannot survive.
- Fail on missing snapshots unless UPDATE_SNAPSHOTS=true; validate
RUNTIME_PARAM/VARIANT_PARAM/UPDATE_SNAPSHOTS; reject argv, which was silently
ignored.
- Capture the HTTP status alongside the body: `curl -s` exits 0 on a 5xx, so
update mode could have recorded error pages as goldens.
- Handle `diff` exit codes 0/1/>1 distinctly; a diff *error* previously read as
success.
- Replace fixed sleeps with polling: a TCP readiness probe (never a POST, which
would consume the cold-start invocation) and a wait for both REPORT and
INVOKE RTDONE per event, since RIE can emit RTDONE after REPORT.
- Guard cleanup() against the blank array slots left by "${array[@]/$cid}".
Node 26:
- Add the leg to the harness and CI matrix. Node 26 ships to customers but has
zero L2 coverage.
- `public.ecr.aws/lambda/nodejs:26` does not exist yet (preview only), so pin
the multi-arch tag 26-preview.2026.08.21.22 through a runtime->image-tag
indirection applied only at the --build-arg; $node_version also feeds image
tags, container names and snapshot paths, which must keep the bare major.
- Node 26 is excluded from the default runtime list and runs in update mode in
CI until PR 2 commits its snapshots. CI asserts it produces exactly 20
artifacts and that no committed snapshot is ever mutated.
- Record the GA re-capture trigger: when AWS publishes the bare image, the Node
26 goldens must be re-captured from the pinned pre-migration ref.
Node 18 and 20 legs are retained deliberately (C18); v5.x is the only dd-trace
line serving those runtimes.
Verified: all eight strict legs (18/20/22/24 x cjs/esm) pass byte-exact; the
node26 update-mode leg produces exactly 20 artifacts; negative tests for
missing snapshots and whitespace-only drift both fail as intended.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The local harness committed one golden per (variant, runtime, event): 72
return values and 8 logs. All 72 return values were byte-identical, and the
8 logs held only one genuinely runtime-specific line each. 80 files carried
3 distinct expectations, and 80 files had to be re-read whenever one of them
legitimately changed.
Collapse them to 3 shared goldens covering all five runtimes and both
variants:
snapshots/return_values/default.json
snapshots/logs/cjs.log
snapshots/logs/esm.log
Two changes make sharing safe rather than lossy:
- Assert `runtime:nodejsNN.x` on every invocation with the major actually
under test, then collapse it to `nodejsXX.x`. Without the assertion,
sharing the golden would silently stop checking that the library reports
the runtime it is running on.
- Drop the runtime major from AWS_LAMBDA_FUNCTION_NAME. It propagates into
service, resource, resource_names, functionname, function_arn,
_dd.base_service and _dd.tags.process, so embedding the runtime there made
every runtime's golden differ in ~100 lines of pure fixture naming, which
buried the one line that was actually runtime-specific.
Divergence is expressed by adding an override file, never by loosening a
comparison:
snapshots/logs/${variant}_node${major}.log
snapshots/return_values/${variant}_node${major}_${event}.json
An override wins for its leg alone, so a real behavioral difference appears
as a new file in review, where widening a normalization filter to absorb it
would appear as nothing. In update mode a leg that disagrees with an
existing shared golden now fails instead of overwriting it; otherwise the
last runtime to run would define the expectation for all of them. That
agreement check uses the same sort-insensitive comparison as strict mode,
since RIE's platform logging races application output.
Node 26 consequently needs no snapshots of its own: its normalized output
matches the shared goldens. That removes the update-mode CI override and the
node26 baseline assertion, making all five workflow legs strict.
The consolidation is auditable rather than a blind re-capture: reverse-
substituting the runtime major and function name into the 3 shared files
reproduces all 72 original return-value goldens byte-for-byte and all 8
original per-runtime logs with an identical set of lines. The only residual
differences are the ordering of RIE's own platform lines (INVOKE RTDONE,
START/END/REPORT), which is already compared order-insensitively.
Also fixes a latent normalization hole this exposed: function_arn is
normalized by a pattern whose character class excludes `_`, so names ending
in `_node18` halted the match and the goldens recorded a partially
normalized `"function_arn":"XXXX_node18"`. With the suffix gone the ARN
normalizes completely.
Verified with a full strict run: 5 runtimes x 2 variants, 110 assertions,
no failures.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| run_id_filter="s/${RUN_ID}/XXXX/g" | ||
| fi | ||
|
|
||
| node "$repo_dir/integration_tests/parse-json.js" | |
There was a problem hiding this comment.
I've found in a couple of the layers (i think node and ruby) that when I run integration tests locally sometimes I get cold start logs and sometimes I don't, so olivier and I have removed them (example) in these repos. You might want to consider doing the same here, or just leave it unless it becomes a problem.
There was a problem hiding this comment.
although now that I read further not sure if we can do that given the proactive init simulation
There was a problem hiding this comment.
When you previously run integration tests "locally", it actually still uses real lambdas. For real lambda, the platform decides when to run init, so sometimes it proactively initializes the sandbox and cold_start flips to false.
This PR is real 'local' using docker and no need for aws credentials whatsoever. And under RIE there's no platform scheduler — we start the container and drive all 9 invocations ourselves, so the cold→warm transition is deterministic.
As for the SIMULATE_PROACTIVE_INIT mode. It is a manual diagnostic, not part of the test gate. In other words, it's not used (for now). To make it more clear, i updated the README for its section.
The readiness poll waits for both a REPORT and an INVOKE RTDONE per invocation. RIE's managed-instances path — the one SIMULATE_PROACTIVE_INIT switches into to get eager init — emits REPORT but never RTDONE, so the poll could never be satisfied and the mode timed out at 9/0 instead of producing a usable diff. Verified directly against both RIE paths with a single invocation: default RIE REPORT=1 INVOKE_RTDONE=1 managed-instances REPORT=1 INVOKE_RTDONE=0 Require RTDONE only on the default path. The mode now runs to completion and surfaces exactly what it is meant to: cold_start flips to false on the first invocation, plus managed-instances platform noise (structured bootstrap logs and a different emulated region in ARNs). Also correct the proactive_initialization row in migration_parity.md, which claimed RIE cannot produce a proactively-initialized sandbox. It can: raw un-normalized logs from a managed-instances run with a 15s init->invoke gap contain "proactive_initialization":1, proactive_initialization:true and cold_start:false. The feature is unprotected because normalization drops those markers and the goldens come from immediate-invoke runs, not because the sandbox state is unreachable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The section documented the mechanism in detail but never said what its status is, so a reader could reasonably assume it is part of the suite. It is not set in CI, not set in the default run, and nothing gates on it; a run with it is expected to differ from the goldens rather than match them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
What this is
This is the first PR in the effort to move
datadog-lambda-js's core logic intodd-trace-js.It doesn't move any code yet — it builds the safety net that every later PR will be checked against.
Per Step 2 of the migration strategy, we need three levels of test coverage. Level 2 — Docker-based integration tests — didn't exist yet: our integration tests run against real AWS Lambdas, which is too slow and too credential-bound for the dd-trace pipeline's few-minute budget. This PR adds that missing level.
What it does
Adds
integration_tests_local/, which runs the container handler variants against the AWS Runtime Interface Emulator — no AWS account, no deployed lambda functions, no credentials.For each of Node 18/20/22/24 × {cjs, esm}, it builds the image, starts it under RIE, fires 9 input events, and diffs both the return values and the normalized logs against committed snapshots.
That's 80 golden snapshots (72 return values + 8 log files) capturing exactly what the library does today.We actually enhance this part to remove replications in the snapshots. so now it's much cleaner.Those goldens are the point. Once we start moving code into
dd-trace-js, they're the oracle that tells us whether migrated code still behaves identically — so they're captured from the current implementation before any migration happens, and they're never regenerated wholesale from migrated code.Also adds a GitHub Actions workflow running all five runtime legs on every PR.
Notes for review
Most of the diff is the 80 snapshot files.The parts worth actual review arerun.shandnormalize.sh.Normalization is the load-bearing piece. It decides which volatility gets erased before diffing (timestamps, request IDs, durations, trace IDs). Erase too little and the tests flake; erase too much and the gate silently stops catching regressions. Two specific decisions:
meta/metricskeys are preserved — their stable keys and values are part of the compatibility contract.REPORTlines only, so application output stays an exact-whitespace comparison.Node 18 and 20 legs are deliberate. dd-trace's
v5.xis the only line serving those runtimes, and we ship layers for them, so they stay in the matrix for the whole migration.(This is the same position taken when closing #791.)
Node 26 is a special case.
public.ecr.aws/lambda/nodejs:26doesn't exist yet — AWS only publishes26-preview.*— so the base image is pinned to a specific preview tag, threaded through a runtime→image-tag indirection applied only at thedocker build --build-arg(the bare major still drives image tags, container names and snapshot paths). Node 26 has no committed snapshots yet, so it's excluded from the default local run and runs in update modein CI, which asserts it produces exactly the expected 20 artifacts without committing them. The next PR freezes them. When AWS publishes the GA image, the Node 26 goldens need re-capturing from the pre-migration code — that's recorded in the README.
Hardening details: RIE binary pinned to v1.36 with a per-platform SHA-256 that's re-verified on every run; missing snapshots fail instead of passing silently; HTTP status is checked alongside the response body (
curl -sexits 0 on a 5xx, so update mode could otherwise record error pages as goldens);diffexit codes 0/1/>1 are handled distinctly; fixed sleeps replaced with polling on bothREPORTandINVOKE RTDONE, since RIE can emit them out of order.Trying it locally
Requires Docker. First run pulls the base images (~1 GB each) and the RIE binary.
What comes next
This is Step 2 of the strategy. The PRs that follow, roughly in order:
aws-lambdaplugin and define a small, stable facade (wrap,sendDistributionMetric,sendDistributionMetricWithDate,getTraceHeaders) for the shim to call.datadog-lambda-jsbecomes thin and delegates to dd-trace, with the layer built from migrated logic.Every one of those is checked against the goldens this PR adds, and progress is tracked row-by-row in
migration_parity.md.