Skip to content

Latest commit

Β 

History

History
714 lines (627 loc) Β· 42.6 KB

File metadata and controls

714 lines (627 loc) Β· 42.6 KB

Changelog

Unreleased

1.10.0

  • One subpackage per provider. Every sandbox moved under code_sandboxes.sandboxes.<provider> β€” the implementation in <provider>/<provider>.py, re-exported by the package's __init__, and the modules only it needs beside it (Kaggle's kernel client, executors and live session; Marimo's reactive graph and cells driver; Google Colab's kernel client). The public names are unchanged: import them from code_sandboxes as before. Breaking: the old flat module paths (code_sandboxes.datalayer_sandbox, code_sandboxes.marimo_sandbox, …) are gone, with no aliases; import from code_sandboxes or the new paths.
  • The Marimo kernel helper is real code. sandboxes/marimo/reactive_kernel.py is a typed, importable, self-contained module; what a sandbox sends to the kernel is that file's own text (reactive.KERNEL_HELPER_SOURCE reads it from the installed package). The type-checker and a new test suite (tests/test_marimo_reactive_kernel.py) now read exactly what the kernel runs: replacement, conflicts, cycles, plans, removal, re-execution keeping the graph, and the stdout wire encoding.
  • Docs: every provider page imports from its subpackage; the Marimo page describes the real-code helper and points at the hosted marimo toolset.

1.9.39

  • The GitHub workflow files are all .yaml now (build, py-tests, py-code-style, py-typing, reusable-python, environments-live), the spelling release.yaml already had; the reusable workflow's callers and the contributing page follow.
  • Marimo reactivity through the Jupyter-shaped API (#37). run_code names its cell through the execution context (Context(id=...); the same id replaces the cell) and the result says what happened: cell_id, and reactions β€” every cell re-run because of it, each a Reaction with its code and its own ExecutionResult. CodeSandboxClient.execute, execute_code and execute_code_streaming take cell_id; a reply carries the reactions under marimo with their own Jupyter-shaped outputs; execute_interactive emits them after the cell's own, each message tagged metadata.marimo = {cell_id, reaction}; the stream tags every event with marimo_cell_id. The client passes run_cell, register_cell, remove_cell, plan, graph and cells through to a reactive sandbox and says so with reactive. The marimo_reactions extra attribute is replaced by the typed reactions field. Nothing changes on the wire.

1.9.38

  • The environments line and main are one branch again: everything released from feat/env-custom as 1.9.13 through 1.9.37 (custom environments: resolve, build, attest, sign, the egress proxy, conda artifacts, the contract's working directory) is on main, rebased onto the marimo sandbox. The 1.9.13 entry below is what PyPI's 1.9.13 shipped; main's own "1.9.13: the marimo sandbox" never reached PyPI (the version was taken) and is this release instead:
  • marimo sandbox (#34, #35): a Jupyter server sandbox whose kernel also holds Marimo's dataflow graph. run_cell(cell_id, code) runs a cell, then the cells that depend on what it defined, in dependency order; a failing cell stops the reaction. register_cell, remove_cell, plan, graph and cells expose the graph; run_code keeps working, each call a cell of its own, with the re-run cell ids on the result as marimo_reactions. The graph is Marimo's own (marimo._runtime.dataflow), installed into the kernel by one execute request and asked over the Jupyter protocol; marimo is pip-installed into a kernel that lacks it unless install_marimo=False. get_manager("marimo") is the jupyter-server manager under the marimo variant. The same helper source drives jupyter-react's variant="marimo".
  • Releases are cut by a v* tag now (.github/workflows/release.yaml, trusted publishing); see RELEASE.md.

1.9.18

  • A lock no longer carries a wall clock, so the build cache can hit (environments/resolve.py, resolve_conda.py; PLAN_ENVS.md E1-26, D-12). The lock's header carried a # resolved-at: line, and its digest is over the whole text β€” so two resolves of the same spec, in the same base, pinning the same 320 packages, produced two different digests. Found by resolving one environment twice on r1 on 2026-09-16: the texts differed in exactly that one line out of 5,388. Section 5's cache key is over the lock digest, so D-12's build cache could never hit, and it never had: environments.cache.lookups read hit=false twelve times out of twelve. When a lock was resolved is on the lock document Runtimes stores, in its created_at, which is where it belongs. resolved_at is gone from lock_document, conda_lock_document, resolve_environment and resolve_conda_environment; the two tests that asserted determinism by freezing the clock now assert it without one.

1.9.17

  • An artifact's size is read from the registry (environments/attest.py; PLAN_ENVS.md E1-25). attest_artifact took size_bytes from its caller and nobody ever passed one β€” the builder answers a reference, not a weight β€” so every artefact was recorded with sizeBytes: null and environments.artifact.bytes, the series section 14 tracks the artifact size in, had no point in it although artifacts had been recorded (seen on r1, 2026-09-16, through the OTEL query API). Attestor.size_of() asks the registry, with the client the scan is already read from, and a size that cannot be read is logged rather than raised: a missing number on a dashboard is not a reason to refuse an artifact that is otherwise signed. 3 new tests.

1.9.15

  • A restart restarts the kernel, not just this client's socket (jupyter_server_sandbox, client; PLAN_ENV.md E0-09, Appendix B check 7). CodeSandboxClient.restart() was stop() then start(), which is right for a sandbox this process owns β€” it is destroyed and recreated, and nothing survives β€” and wrong for one attached to a Jupyter server somebody else runs, which is every Datalayer runtime pod: stopping drops the websocket while the kernel process keeps running, so the reconnect lands in the same interpreter with every global still set. Check 7 is "nothing is assumed to persist across restarts", and it read state survived the restart ('True') for exactly this reason β€” found live on r1, 2026-09-16, the first drill whose smoke test reached the check. JupyterServerSandbox.restart_kernel() now asks the server's own POST /api/kernels/{id}/restart (the way _do_interrupt already uses the API rather than the client's lifecycle) and reconnects onto the new kernel; restart() prefers it and falls back to the lifecycle for every variant that draws no such distinction. 7 new tests.
  • datalayer/python-cpu:2026.09 repinned to sha256:122d3e31f5e2507251457cbf47871c39ac1753adb1d83777ab0743fa11cd6148: the contract layer now sets MappingKernelManager.root_dir, so kernels start in /home/datalayer/content. The image already declared WORKDIR there and sandbox-contract/v1's User row already required it, but a kernel's cwd is the Jupyter server's to choose and jupyter-python's config roots it at $HOME β€” so every environment's kernel ran in /home/datalayer and Appendix B check 2 read cwd is '/home/datalayer', not '/home/datalayer/content'. The file browser stays rooted at $HOME, where a person expects to see everything they have; only the kernel moves.
  • Two assertions that had rotted through three base releases are pinned in one place again: the channel's digest and its apt snapshot were duplicated across test_environment_bases.py and test_environment_resolve.py, and 2026-09-15's and 2026-09-16's releases left both red rather than catching anything.

1.9.14

  • datalayer/python-cpu:2026.09 base channel repinned to the rebuilt jupyter-python:0.2.2 (now carrying jupyter-kernels==1.2.23) plus the contract layer, digest sha256:aa5413000bb5b6ecd0a0cf03959b107f0d572f65bf230c08bbdf9a4569775545, released 2026-09-16 to environments/base/python-cpu. Every variant pins the same digest.

1.9.13

  • jupyter-kernels==1.2.23 forced into sandbox-contract/v1: it carries the pooled kernel manager the runtime's Jupyter config selects (kernel_manager_class = jupyter_kernels.pool.mapping.PooledMappingKernelManager), replacing the deprecated private datalayer-kernels. PyPI serves it, so a resolve satisfies it from the index and the wheelhouse carries no wheel for it; the pin keeps uv pip sync --require-hashes from stripping it out of a user environment's image.

1.9.12

  • owner_repository, owner_cache_repository and ECR_ENVIRONMENT_PREFIX moved from environments.adapters.datalayer to environments.builders (PLAN_ENV E1-14): Runtimes now validates a smoke-test launch's artifact against this same repository shape, and builders is the neutral module a service may import β€” adapters.datalayer is not, and importing it from a service binds that service to the Datalayer provider the way check_provider_boundary.py exists to prevent. Re-exported from the adapter unchanged, so every existing from .datalayer import owner_repository still works. No behavior change; 78 tests still pass.

1.9.11

  • 1.9.10's own fix did not work: --registry-referrers-mode only ever governs reading referrers, never where sign writes one β€” sign --help says so outright ("mode for fetching references"), and a second live sign on r1, 2026-09-14, confirmed it: the signature still landed as an OCI 1.1 referrer, no legacy tag. Attestor no longer tries to force cosign's own storage choice. Its replay check and its own signature_ref now ask cosign directly, the same way the Operator's own verify will: can_verify(reference) runs cosign verify --key <key> --insecure-ignore-tlog=true <reference> and answers its exit code, and signature_ref is the digest reference itself β€” what cosign verify takes, not a tag it may or may not have written. Signing the same digest twice no longer errors the way a second push under the old immutable tag once would have, but can_verify still avoids it, since two valid signatures claiming to be Datalayer's own word on one artifact is not the design either.

1.9.10

  • cosign signs the legacy tag again, not just an OCI referrer (environments.attest; PLAN_ENV.md E1-09). cosign 3.1.3 defaults to the OCI 1.1 referrers API for a signature's own storage, not the classic sha256-<hex>.sig sidecar tag. Found live on r1, 2026-09-14: cosign sign reported success β€” the first artifact this pipeline ever actually signed β€” but pushed no such tag; _signature_exists's own replay check and the signature_ref this workflow records both assume one exists. --registry-referrers-mode=legacy restores it. cosign verify needs no matching flag: it already looks for the tag by default.

1.9.9

  • cosign gets the same registry credential the build pushed with (environments.attest; PLAN_ENV.md D-17, E1-09). cosign has no AWS credential chain of its own for ECR, unlike the boto3 client the scan is read with. Found live on r1, 2026-09-14: cosign sign reached the repository anonymously and was refused, plain 401 Unauthorized, on every real artifact this pipeline ever tried to sign. Attestor takes a registry_auth β€” the same {"DOCKER_CONFIG": <dir>} shape the resolver and the builder already read off a BuildCredential β€” and sets it on cosign's own subprocess, added to this process's own environment rather than replacing it.

1.9.8

  • Every page of a scan's findings is read before deciding (environments.attest; PLAN_ENV.md E1-08, D-11). describe_image_scan_ findings paginates, and the attestor called it exactly once. Found live on r1, 2026-09-14: the first real image with more findings than one page had 1,547 enhanced findings, 31 of them critical, and the single-page read saw a small enough slice that the decision passed β€” an image with real, unreviewed critical vulnerabilities would have been signed and started. A hard cap of 50 pages keeps a pathological registry from paginating forever; hitting it logs that more findings went unread rather than pretending the scan was complete.

1.9.7

  • cosign signs again under --use-signing-config=false (environments .attest; PLAN_ENV.md E1-09). cosign 3.1.3 defaults --use-signing-config to true: a TUF-provided signing config now names the service URLs, including a transparency log, and --tlog-upload=false alone no longer overrides that β€” cosign refused the combination outright (found live on r1, 2026-09-14, the first real signature this pipeline ever attempted). Turning the signing config off restores the plain, flag-driven behavior --tlog-upload=false already asks for.

1.9.6

  • datalayer/python-cpu's 2026.09 channel points at the patched base (environments.bases; PLAN_ENV.md E1-05, E1-08). jupyter-python 0.2.2 upgrades the Ubuntu security packages and conda's own OpenSSL, and drops JupyterLab's staging yarn.lock β€” the 31 fixable-critical findings that blocked every build from the prior digest under the default scan policy.

1.9.5

  • Attest waits for a just-pushed image's scan to exist (environments.attest; PLAN_ENV.md E1-08). Enhanced scanning starts after the push, so the first answers for a new image are ScanNotFoundException. Attest read that as an image with no scan and failed the build at once (found live on r1, 2026-09-14). It is now waited for like a running scan, within the same bound.
  • apt is pinned at, and installed from, the base channel's Ubuntu snapshot (environments.bases, environments.resolve, environments.adapters.datalayer; PLAN_ENV.md E1-04). A channel records the snapshot.ubuntu.com id its image was upgraded at. The solve runs apt-get with --snapshot at that id, the lock records it (# datalayer-apt-snapshot:), and the Datalayer builder installs its pins from the same snapshot. This replaces a deb line that named only main (gdal-bin is in universe) and left the live mirror enabled beside it. DATALAYER_APT_SNAPSHOT must now be a snapshot id.
  • A private registry's credential is resolved and used (environments .image_import, environments.resolve; PLAN_ENV.md E3-04's second half). image.credentialSecretId used to only clear the allowlist check; nothing fetched or used the credential it named. It is now fetched through the same IAM route a build secret is (E3-05), and sent as HTTP Basic on the registry's own token exchange β€” the way docker login authenticates a private pull β€” never to the registry named in the reference, and never logged. resolve_environment and resolve_image_base take owner_uid and resolve_secret to reach it; unused by a public registry or any other source.
  • The Datalayer base channel's own five ESM-locked advisories are allowed (environments.policy; PLAN_ENV.md E1-08). Amazon Inspector marks all five fixable, but the fix is an Ubuntu ESM package version a plain apt-get upgrade cannot reach. Reviewed and added to DEFAULT_POLICY .allowed, 2026-09-14, so they are recorded as allowed rather than missed, and any other critical still blocks.

1.9.4

  • A pushed Environment image is decided by its linux/amd64 image's scan (environments.attest; PLAN_ENV.md E1-08). The Datalayer builder pushes with SBOM and provenance attestations, so the digest it records is an OCI image index. Amazon Inspector scans the image inside it and answers UNSUPPORTED_IMAGE for the index itself, so every real build failed at attest with DL_ENV_PROVIDER_ERROR (found live on r1, 2026-09-14, the first build to get past the scan-read permission). The attestor now reads the manifest and, for an index, reads the scan of its linux/amd64 image, skipping the attestation manifest. The signature stays on the index, which is what a pod pulls.
  • Modal attaches a build secret to the postInstall steps that name it (environments.adapters.modal; PLAN_ENV.md E3-05). Modal used to refuse every buildSecrets spec. Each declared secret is now resolved from IAM before Modal is touched, made into a Modal Secret of the build's own, passed as secrets= to exactly the run_commands steps whose command names it, and deleted after the build like the base-reader secret. Its value is redacted from the log and from a failed build's error. A mountAs: file secret is refused, because Modal passes secrets as environment variables and a file would land in a layer.

1.9.3

  • A build secret is mounted only on the postInstall commands that name it (environments.spec.command_names_secret, environments.adapters.datalayer; PLAN_ENV.md E3-05). Every declared secret used to be mounted on every postInstall RUN. A command now gets a secret's mount only when the secret's name is in it as a whole word ($NAME, ${NAME}, --key-env NAME, /run/secrets/NAME). A declared secret that no command names is a DL_ENV_SPEC_INVALID finding on spec.buildSecrets[i], rather than a build that runs with it empty.

1.9.2

  • A real ECR repository name is lowercase; a real owner uid is not (environments.adapters.datalayer.owner_repository, new owner_cache_repository; PLAN_ENV.md E1-07). Found live, 2026-09-14, the first real Datalayer build ever run for a real account's own uid rather than a lowercase test fixture: DescribeRepositories refused outright, "Invalid parameter at 'repositoryName'" β€” this project's own uids are ULIDs, conventionally uppercase, and ECR repository names match [a-z0-9]+((\.|_|__|-+)[a-z0-9]+)* per path segment. Both the owner's own repository and the owner's cache repository are lowered at the one place each is built, so every caller (create, inspect, exists, delete, the cache import/export) stays consistent with itself.

1.9.1

  • mTLS to the build pool's buildkitd (environments.adapters.datalayer.Builder, environments.resolve.BuildkitResolveRunner; PLAN_ENV.md E1-06/E1-07). Both drivers of buildctl took only --addr, which is what every prior live drill needed against a plain-socket, ephemeral buildkitd β€” never a real one. The build pool's real daemon on r1 takes mTLS connections only, so reaching it needed three more flags neither driver had: --tlscert, --tlskey and --tlscacert, each read from DATALAYER_BUILDKIT_TLSCERT, DATALAYER_BUILDKIT_TLSKEY and DATALAYER_BUILDKIT_TLSCACERT when not passed explicitly (Builder, matching how it already reads DATALAYER_BUILDKIT_ADDR) or passed in by the caller (BuildkitResolveRunner, constructed by durable's own activities_environments.py, which now reads and forwards the same three). All three or none β€” a partial set is treated as none, since a buildkitd requiring mTLS refuses a client carrying only some of them at the daemon, with a less useful error than refusing here.

1.9.0

  • The Daytona builder (environments.adapters.daytona, PLAN_ENV.md E2-04). A snapshot built from the approved Datalayer base β€” uv pip sync --require-hashes against the resolved lock, your files, postInstall commands, and an explicit tini -- sleep infinity entrypoint β€” with the owner's own Daytona organization, never this package's ambient credentials. Confirmed live against a real, hash-verified lock: a real snapshot reaches Active and a sandbox launched from it passes the full sandbox contract, doctor included β€” the one managed variant, of the three landing here, whose build-time state and launched state fully agree.
  • The E2B builder (environments.adapters.e2b, PLAN_ENV.md E2-03). A template built from code-interpreter-v1 β€” E2B's own proprietary code-interpreter server ships baked into it only, so the build starts there rather than from the Datalayer base and reconciles whatever it shipped to exactly what the lock pins. Two live-found identity bugs are fixed: code-interpreter-v1 already holds an account at uid 1000 and a group at gid 100 of its own, so the old useradd was silently landing elsewhere β€” the build now renames the existing account instead β€” and the base sets no locale at all, now set explicitly. A real build passes the contract's doctor check in full. Still open: a launched sandbox currently runs as root, not the contract's 1000:100 β€” the two systemd services that actually execute a launched sandbox's code have no User= of their own, and E2B's private server hardcodes /home/user as the working directory, independent of anything the build sets.
  • The Modal builder (environments.adapters.modal, PLAN_ENV.md E2-05). An image built from_aws_ecr off the approved base, with a per-build Secret made and torn down around it. The contract's own identity is put back at launch: modal_sandbox.py's session driver now drops itself to 1000:100 for a sandbox launched from a built Environments artifact specifically (gated on image_id, so general Modal sandbox usage elsewhere is unaffected) β€” Modal ignores the image's own USER, so this is the only place it can be restored. A separate, previously silently-fatal bug is also fixed: the real launcher creates the sandbox with no command arguments at all, so an entrypoint relying on exec "$@" alone was a no-op and every launched environment exited within seconds; the entrypoint now falls back to sleep infinity when given none. Still open: two of the core tier's checks (imports, filesystem) still fail on a launched sandbox, for a reason not yet root-caused β€” confirmed specific to a contract-built image under repeated exec, not the identity fix.
  • A real, shared bug in all three managed launchers, found in review: start() creates the remote sandbox well before it marks itself started, while stop() was guarded on that flag rather than on the resource itself β€” a failure in between left a real, running sandbox that stop() skipped entirely, orphaned for good. Fixed identically in daytona_sandbox.py, modal_sandbox.py and e2b_sandbox.py.
  • A nightly live matrix for the three managed builders (.github/workflows/environments-live.yaml, PLAN_ENV.md E2-11). Builds a real artifact on each real provider from a real, hash-verified lock, launches a sandbox from it, and runs the formal core tier β€” opening (or commenting on) an issue naming the provider and the check on a genuine failure. E2B and Modal run xfail(strict=False) for their own already-documented gaps above; Daytona carries no such marker. Needs provider secrets added to the repository before it does anything for real.

1.8.2

  • A sandbox can be asked whether it is still alive, and the answer no longer comes from a flag we set ourselves. CodeSandboxClient.is_alive() returned is_started, which records that start() ran in this process and nothing else, so it kept answering True after the backend was gone. Sandbox now carries an is_alive() that variants override; the Jupyter Server variant looks its kernel id up on the server the way _find_existing_kernel lists them, and the client asks the sandbox instead of reading its own flag. Variants that have no way to ask their provider inherit the base answer, is_started, which is all they can honestly say.
  • Build secrets (environments.build_secrets, PLAN_ENV.md E3-05). spec.buildSecrets now resolves for real on the Datalayer variant: each declared secret's value is fetched from IAM's own internal route at the moment the postInstall step runs, mounted with BuildKit's --mount=type=secret (an environment variable or a file under /run/secrets/, per mountAs) in its own directory outside the build context, and passed to buildctl by file path β€” never in argv, never in a step result, never an ARG/ENV that would bake it into the image's history. E2B and Daytona refuse a spec naming one outright: E0-04's spike found only a registry login for the private base on either, never a per-step arbitrary secret. A version with any build secret can never be published to the Library (D-12). New codes DL_ENV_BUILD_SECRET_UNAVAILABLE and DL_ENV_PUBLICATION_BLOCKED.

1.8.0

  • The Datalayer builder (environments.adapters.datalayer, PLAN_ENV.md E1-07). The Dockerfile is generated from the lock β€” uv pip sync --require-hashes, apt at the versions the lock recorded, env before anything installs, postInstall as uid 1000 with no network β€” and the push is by digest with an SBOM and provenance attestation, under the operability tag v<n>-<build_uid> so a retried build cannot collide with the attempt before it. inspect, resolve, exists and delete go through the ECR API. Needs the environments-builder extra.
  • The scan and the signature (environments.policy, environments.attest, E1-08, E1-09). A critical finding with a fixed version blocks and one nothing fixes is recorded; the decision record keeps the threshold it was decided under, so what stopped a build reads a month later. cosign signs only once the scan passed, and a replay finds the signature rather than pushing a second.
  • A sandbox launches from an artifact (E2-02): Sandbox.create(artifact=…) hands each variant its own argument β€” an E2B template build, a Daytona snapshot, a Modal image id through Image.from_id.
  • Fixed, and a live defect: Daytona's adapter took the image branch whenever resources were requested, so a snapshot asked for with cpu= came up from a plain Debian image running none of the snapshot's content, with nothing saying so. The combination is refused (correction 13).

1.7.0

  • A version is resolved into one lock (code_sandboxes.environments.resolve, PLAN_ENV.md E1-04, D-9). Datalayer's protected constraints are merged over the user's requirements β€” a requirement that agrees with a pin is dropped for it, one that contradicts it is DL_ENV_PROTECTED_PACKAGE with the supported range β€” the base is resolved to a digest per requested variant, and uv's refusals are read into the error taxonomy: a conflict with the pair that cannot hold, a package no index has, a protected pin, and DL_ENV_PROVIDER_ERROR, which is retryable, for a failure that is not about the version.

    The lock is uv's hashed output with the apt pins and the protected pins above it as comments: one document that says everything a build installs, and still a requirements file. BuildkitResolveRunner is D-9's solve, FROM the resolved base digest; LocalResolveRunner runs uv where it is called, for plane local, and refuses to pin apt rather than lock another distribution's versions.

  • Every variant has a real interrupt now, and can be reached to give it. Sandbox.interrupt is two gates β€” _executing_event must be set, then _do_interrupt must answer β€” and seven variants failed one of them, each in a way that looked like success:

    • docker, google_colab, kaggle and monty had no _do_interrupt, so the base default ran: set a flag, return True, stop nothing. True means "the interrupt was delivered", and nothing had been. google_colab and kaggle did read the flag, but only after the run, to label a finished result as interrupted β€” labelling a run is not stopping one.
    • cloudflare, coreweave, daytona, e2b and modal answered honestly (False: this provider takes no interrupt) but never marked their execution window, so interrupt() returned at the first gate and their answer was never reached β€” and is_executing was always False, which every status above them believed.

    Docker, Colab and Kaggle now interrupt for real, delegating to the JupyterKernelClient that already knows how to authenticate to their server rather than rebuilding the request. Monty answers False with its reason: feed_run executes a snippet in one blocking call with no limits and no callable hook, so there is no point at which a flag could be read.

    The execution window is @marks_execution on run_code rather than a line in each body, because it has to close on every path out β€” including the early return ExecutionResult(...) each adapter uses for an infrastructure failure β€” and a finally in the decorator cannot be forgotten in one branch of one variant. Both invariants are asserted across the package with no pinned exceptions left.

  • A Datalayer sandbox says when it is running code, so an interrupt can reach it. Sandbox.interrupt refuses before it delegates β€” if not self._executing_event.is_set(): return False β€” and run_code never set that event, so is_executing was always False, every interrupt was refused at the door, and the refusal was reported as "no code was running" to a caller watching a cell run.

    This is the other half of the interrupt fix released in 1.4.4, and the measurement said so: with _do_interrupt implemented and deployed, cancelling a two-minute cell ten seconds in still left the kernel busy for 60.5 seconds. Implementing the interrupt changed nothing while nothing could call it. Both are needed, and the package's tests now hold them together.

    Five variants β€” cloudflare, coreweave, daytona, e2b, modal β€” implement _do_interrupt and never mark execution either, so their interrupts are unreachable in the same way. Pinned as known rather than fixed blind, since none can be measured from here.

  • DatalayerSandbox can be interrupted. It neither implemented _do_interrupt nor read the flag the base class sets, so the default ran instead: it sets a flag, returns True, and stops nothing. Everything above believed it β€” tasks/cancel in the MCP gateway marks a task cancelled and calls the interrupt execute_cell registered, which is this. Measured on prod1 on 2026-09-07: a two-minute cell cancelled ten seconds in answered cancelled from both tasks/cancel and tasks/get, and the next execute_code on that session took 60.8 seconds and came back empty, where a free kernel answers in about a second. The cell ran to completion on a runtime that went on being billed. It now delegates to the runtime's own sandbox_client β€” the jupyter-server client run_code already executes through, which interrupts the kernel over the REST API β€” and answers whether the interrupt was delivered, never raising: a cancel that cannot reach the kernel is a cancel that failed, and the caller decides what that means.

    docker and monty have the same hole and are pinned as known in tests/test_a_datalayer_sandbox_can_be_interrupted.py, so the set cannot grow quietly. (google_colab and kaggle are entitled to the default: they poll _interrupt_requested and stop cooperatively.)

  • provider_catalog takes a names argument, so a caller can describe the providers it serves and pay for only those. Added under 1.3.1 without a version bump, which is what broke the operator: its image installs this package unpinned, PyPI's newest was 1.3.0, and the two-argument call landed on a one-argument function β€” TypeError: provider_catalog() takes from 0 to 1 positional arguments but 2 were given, and datalayer envs ls answered 500. A new public parameter is a feature; released as 1.4.0 so a dependant can ask for it.

  • CodeExecutionOutcome now carries outputs: the rich results as Jupyter outputs, with their mime bundles intact. It called itself a faithful superset of the raw ExecutionResult and was not β€” every representation but text/plain was dropped on the way through, so a matplotlib figure reached its callers as the string <Figure size 640x480 with 1 Axes> and anything wanting to draw it had nothing to draw. results is unchanged, for callers that only print.

  • Exported the lifecycle vocabulary from the package root, so a consumer writes from code_sandboxes import SandboxLifecycle rather than reaching into code_sandboxes.lifecycle β€” the import path is the part that cannot be changed afterwards, which is why it moved before anything depended on it.

    The vocabulary gained update (the Runtimes API's PUT) and split in two. SandboxLifecycle is one sandbox β€” start, stop, pause, resume, snapshot, run_code; SandboxManagerLifecycle is whoever hands them out β€” create, list, get, update. They were one protocol that quietly disagreed with LIFECYCLE_OPERATIONS, because Sandbox.create is a classmethod and a client's create is not; INSTANCE_OPERATIONS and MANAGER_OPERATIONS now say which verb belongs to which shape, and a test holds them to covering every verb exactly once.

    It also gained the URL builders β€” runtimes_url, runtime_url, runtime_pause_url, runtime_resume_url, sandbox_snapshots_url, sandbox_snapshot_url, runtime_checkpoints_url β€” so every Python caller of the Runtimes API builds a path from one place instead of its own f-string. snapshot is recorded against POST /sandbox-snapshots, which is the route that exists; it had been documented as a sub-path of the runtime, which was not.

  • Numbered the prompt's examples, and made one runnable by its number: :examples lists them 1., 2., … and :examples:2 prints the second and then executes it, for a reader who wants the answer rather than the paste. They are now declared where the sandbox is made β€” Sandbox.create(..., examples=[...]), carried on SandboxConfig β€” so run_repl(sandbox) finds them without being told twice; passing them to run_repl still overrides for one prompt. Snippets are printed with Rich's markup off, since [...] was being read as a style tag and a snippet holding list[str] printed as list = [], wrong exactly where someone was about to copy it.

  • Added :examples to the sandbox prompt. run_repl(sandbox, examples=[...]) takes title-and-code pairs and prints them on request, for a reader to copy into the prompt; every REPL example under examples/repl ships its own, and the ones that can take a GPU offer device discovery and a timed matmul instead of their general set when --gpu was asked for. The snippets avoid blocks on purpose: the prompt reads one line at a time, so a pasted for or def would arrive without its body.

  • Fixed a daytona GPU sandbox failing to be created at all unless it was also asking for preemptible capacity. Daytona requires every GPU sandbox to be ephemeral β€” "GPU sandboxes must be ephemeral; set autoDeleteInterval to 0" β€” and auto_delete_interval=0 was being set only on the spot=True path, so a plain gpu="H100" was refused by the API. It now follows the GPU itself, which is what Daytona ties it to.

  • Added three cloud variants: e2b, coreweave and cloudflare.

    e2b runs in a Firecracker microVM through E2B's code interpreter SDK, so it holds a Jupyter kernel per context β€” x = 1 in one call is still there in the next β€” and answers with rich display data: a figure comes back as an image, an HTML repr as HTML. It needs E2B_API_KEY and pip install code-sandboxes[e2b]. set_timeout() extends the life of a running sandbox and get_host(port) gives the public host of a port inside.

    coreweave runs a container on CoreWeave's GPU cloud. What the SDK offers is exec β€” a process at a time β€” so a namespace is held here instead: one python -u -c session is started with the sandbox and fed JSON lines on stdin, the same arrangement the modal variant uses, and snippets share a namespace as they do everywhere else. A session that cannot start, or that goes away, drops back to a process per snippet rather than failing. It needs CWSANDBOX_API_KEY and pip install code-sandboxes[coreweave].

    cloudflare runs a container on Cloudflare's edge. Cloudflare's own SDK is a Workers binding written in TypeScript, which a Python process cannot hold, so this variant drives the SANDBOX BRIDGE β€” the Worker Cloudflare publishes to expose the SDK over HTTP. Deploy it once with npm create cloudflare -- sandbox-bridge --template=cloudflare/sandbox-sdk/bridge/worker, then set CLOUDFLARE_SANDBOX_API_URL and CLOUDFLARE_SANDBOX_API_KEY. The bridge gives a started process nothing to write to, so each snippet runs in one of its own and state does not carry between calls β€” put what shares state in one snippet, or keep it in a file, which does persist. Its manager creates, gets and deletes; it cannot list, because the bridge has no endpoint that enumerates sandboxes, and says so rather than answering with an empty list.

  • Hardened the three new variants against silently doing something other than what was asked. cloudflare now carries SandboxConfig.env_vars into every snippet β€” the bridge takes no environment when it creates a sandbox, so they had been accepted and dropped β€” refuses a network_policy it cannot apply rather than leaving a sandbox believed to be cut off connected, refuses get_variable with the reason instead of answering the misleading "no such variable", and serves files.read/files.write through the bridge's own file endpoints so they need no session at all. coreweave refuses the variable APIs when there is no session process β€” under stateful=False, or after one was lost β€” rather than reporting a successful set that vanishes with the process, and a snippet that runs past its timeout now has its session STOPPED rather than left running and changing the namespace behind a call that already returned.

  • A GPU asked of a variant that has none is now REFUSED rather than dropped. --gpu reaches coreweave, datalayer, daytona, kaggle and modal, and code-sandboxes exec -v e2b --gpu H100 says which variants can give one instead of running on a CPU as though nothing had been asked β€” a sandbox that looks as though it asked for an H100 and did not is one whose timings mean nothing. --gpu was previously accepted and silently ignored for every other variant.

  • Corrected the module docstring of the modal variant, which still described the process-per-snippet behaviour that the session process replaced: modal keeps a namespace between snippets, and falls back to a process per snippet only when the session cannot be held.

  • Added GPU support to the daytona variant. gpu= takes Daytona's own flavors, gpu_count= how many, and several names comma-separated are an ordered list of preferences Daytona falls back along β€” gpu="H100,H200" takes an H200 when no H100 is free. spot=True runs on preemptible capacity, which is far cheaper and outside the GPU quota; it is GPU-only and built from an image with auto_delete_interval=0, both checked before the request rather than left to come back as an API error. DaytonaSandbox.preempted_at() answers when a spot sandbox was reclaimed, and run_code asks on your behalf so that an eviction is not reported as a dropped connection. code-sandboxes exec/repl --spot reaches it from the CLI.

  • Renamed the google_colab variant to google-colab, so every canonical variant name is spelled the one way (jupyter-server already was). Any spelling is still accepted everywhere a variant is named β€” normalize_variant now folds to the canonical dashed form rather than to underscores, which is what a dispatcher compares against, so the two can no longer drift apart.

  • Added code-sandboxes exec, which runs one snippet in a fresh sandbox of any variant and exits with the status the code earned β€” 0 when it ran cleanly, 1 when it raised β€” so it composes in a shell. The code comes from an argument, from --file, or from standard input; --quiet prints only what the code produced. exec and repl take the same options.

  • Moved the machinery for showing a run β€” show_code, show_result, show_and_run, run_repl, repl_prompt β€” into code_sandboxes.console, exported from the package. It existed three times over: in the CLI, in the REPL examples and in the exec examples, disagreeing about whether the value of a trailing expression is shown, whether stderr is told apart from stdout, and which words end a session. The examples now import it like any other consumer, and examples/*/[exec|repl]_common.py are gone.

  • Added the daytona sandbox variant (DaytonaSandbox), running code in a Daytona cloud sandbox. It drives the sandbox's code interpreter rather than process.code_run, so state persists between calls and create_context() gives a namespace Daytona keeps apart. The value of a trailing expression is captured and returned as ExecutionResult.text, which the interpreter itself does not report. GPUs, cpu/memory and the network policy map onto Daytona's own settings; binary files go through its filesystem API. Authenticate with DAYTONA_API_KEY (or DAYTONA_JWT_TOKEN with DAYTONA_ORGANIZATION_ID) and install with pip install code-sandboxes[daytona]. get_manager("daytona") answers the CRUD verbs over an organization's sandboxes.

  • Added the kaggle sandbox variant (KaggleSandbox) to connect to a Kaggle interactive notebook runtime via jupyter-kernel-client's KaggleKernelClient. Authenticate with a Kaggle API token (token argument or the KAGGLE_API_TOKEN environment variable) β€” omitting kernel_id then creates a new kernel. Alternatively, connect to an existing session with a server_url/kernel_id or a notebook session channels_url (the signed JWT in the proxied URL provides the authentication). Install with pip install code-sandboxes[kaggle].

  • Enhanced KaggleSandbox with a transparent batch primitive: when no runtime connection details are provided, it automatically executes code through KaggleKernelExecutor (submit/poll/download) so integrations like jupyter-mcp-server can run on Kaggle without requiring interactive runtime wiring.

  • Added Kaggle accelerator forwarding in batch mode: Sandbox.create(variant="kaggle", gpu=...) now passes the value to KaggleKernelExecutor.execute(accelerator=...), supporting both Kaggle API values (NvidiaTeslaT4, ...) and friendly aliases (T4, P100, ...).

  • Updated ColabSandbox to be reuse-only for existing Colab runtimes and added channels_url parsing support for extracting server_url / kernel_id / proxy_token directly from the Colab WebSocket channels URL.

  • Breaking change: sandbox variant names are eval, docker, jupyter, and datalayer.

  • Removed support for the older local-* variant names from the public API and documentation.

  • Clarified in the documentation that Sandbox.create() defaults to datalayer.