diff --git a/README.rst b/README.rst index 405cadd..5e07861 100644 --- a/README.rst +++ b/README.rst @@ -48,6 +48,19 @@ of ``path``, the cache is ignored. be a sequence of names of arguments of the decorated function that will be ignored for caching purposes. +``memoize_path()`` can also take a ``custom_fingerprint`` callable to use +instead of ``stat()``; it returns ``None`` to fall back to ``stat()``. For +example, ``fscacher.annex_key_fingerprint`` fingerprints locked git-annex'ed +files by their keys, so results survive their content being dropped: + +.. code:: python + + from fscacher import annex_key_fingerprint + + @cache.memoize_path(custom_fingerprint=annex_key_fingerprint) + def foo(path, ...): + ... + Caches are stored on-disk and thus persist between Python runs. To clear a given ``PersistentCache`` and erase its data store, call the ``clear()`` method. diff --git a/docs/design/git-annex-content.md b/docs/design/git-annex-content.md new file mode 100644 index 0000000..dbb2252 --- /dev/null +++ b/docs/design/git-annex-content.md @@ -0,0 +1,133 @@ +# Custom fingerprints and git-annex content + +Design notes for `memoize_path(custom_fingerprint=...)` and +`fscacher.annex_key_fingerprint`, introduced in +[#113](https://github.com/con/fscacher/pull/113). The docstrings and README +only say what these are; this records why they work the way they do. + +## Motivation + +`memoize_path` keys its cache on a `stat()` of the path argument (mtime, +ctime, size, inode). Some resources can vouch for their content better: + +- A *locked* git-annex'ed file is a symlink whose target names a key; for + content-hash backends, the key pins the content. It also survives the file + being moved, re-cloned, or its content dropped. +- Non-path objects (e.g. dandi-cli's `Readable`s, which may stream remote + content) have no `stat()` at all, but may know a fingerprint of their + content. + +## `custom_fingerprint`: an alternative to `stat()`, not a second code path + +The callable replaces only the `stat()`-based fingerprint; everything else +(falling back to a direct call when there is no fingerprint, injecting the +fingerprint into the cache key, tokens) is shared. `_get_fingerprint()` +returns either a `CustomFingerprint` or a `PathFingerprint`, or `None`, and the +rest of `memoize_path` does not care which. `PathFingerprint` keys exactly as +before, so existing caches stay valid. + +The contract for a callable: + +- **It returns `None` to fall back to `stat()`**, for anything it does not + recognize, including plain paths unless it fingerprints them. It must not + raise for such values: an exception propagates, so the decorated call + fails. This is deliberate, as an exception there most likely is a bug in + the callable, which should not silently disable caching. (A way to say "do + not cache", e.g. a named `fscacher.NO_CACHE` constant, can be added if ever + needed.) +- **It runs on every call** of the decorated function, so it must be cheap. + `annex_key_fingerprint` only reads the symlink; it never runs git-annex. +- **Equal fingerprints share results, wherever they are.** The path argument + itself is not part of the cache key: in `stat()` mode the (dereferenced) path + is part of the fingerprint, but a custom fingerprint is used as is. So a + callable must include the path in its fingerprint unless the result does not + depend on it, as `annex_key_fingerprint` does by default (see + `pair_with_path` below). For the same reason the value need not be a path + at all. +- **Fingerprints must be picklable with a stable `repr()`**, as joblib hashes + them into the key (e.g. a string or a tuple of strings). +- **There is no "modified just now" window.** With `stat()`, a file modified + within `_min_dtime` is not cached, since a quick later modification might + not change the fingerprint. A custom fingerprint is trusted to change + whenever the result may change; the callable is responsible for that. This + is only justified for fingerprints that pin the content, like hash-based + annex keys. + +## Which git-annex files are fingerprinted by key + +Only **locked** files: a symlink whose target looks like +`.../annex/objects/*/*/KEY/KEY`. The link is read with `readlink()` only. + +Only **content-hash backends**: `SHA1`, `SHA224`/`256`/`384`/`512`, +`SHA3_224`/`256`/`384`/`512`, `SKEIN256`/`512`, `BLAKE2B*`/`BP*`/`S*`/`SP*`, +and `MD5`, each with or without the `E` suffix (which adds the extension). +Everything else falls back to `stat()`: + +- `WORM` keys are made of size, mtime and file name, so they do not pin the + content. +- `URL` and `VURL` keys name a URL whose content may change. +- External `X*` backends are unknown. + +**Unlocked files** also fall back to `stat()`: an unlocked file is a regular +file whose key is not updated until it is re-added, so an edit would not +change the key. This includes every file on an *adjusted unlocked branch*, +which git-annex uses on filesystems without usable symlinks or permissions, +notably native Windows. There, `annex_key_fingerprint` is a no-op. + +WSL behaves like Linux on its own filesystem (e.g. under `~`). On a Windows +drive under `/mnt/c`, git-annex is likely to treat the filesystem as crippled +and use an adjusted branch, so the same `stat()` fallback applies. Neither is +tested in CI. + +## `pair_with_path` + +By default the fingerprint is `(absolute path, key)`: the path as given, not +dereferenced (dereferencing would give the object path, identical for all +files with the key). Results are then only shared by files at the same path, +which is needed whenever the result depends on the path, e.g. on the extension +(the key's extension may differ from the file's) or on neighboring files. +The path is made absolute with `os.path.abspath`, which does not resolve +symlinks (neither the file's nor its parent directories'), so the same file +reached through different paths gets separate entries; that errs on the safe +side. + +With `pair_with_path=False`, the fingerprint is the key alone, and results are +shared by all files with the same content: twins in a dataset, and the same +file across clones. Only use it for results that depend on the content only. + +## Dropped content + +With `stat()`, a locked file whose content was dropped is a broken symlink: it +cannot be fingerprinted, so the function is called directly every time. + +With the key, it is still fingerprinted, so a result cached while the content +was present is returned after `git annex drop`. This is useful (e.g. metadata +of files no longer present locally) and deliberate. A file whose content was +never present has no cached result, so the function is called and fails as +before. + +## Directories + +`memoize_path` fingerprints a directory argument by walking it and combining +the `stat()` of every entry (this predates custom fingerprints). With `stat()` +alone, a single dropped annexed file (a broken symlink) makes the whole +directory unfingerprintable, and the mtimes of annexed files feed the +"modified just now" window. + +**The callable only sees the top-level argument.** For a directory, it may +return a fingerprint of the whole tree itself; if it returns `None` (as +`annex_key_fingerprint` does), the directory is walked and `stat()`-ed exactly +as before, dropped files included. This keeps a single call site and a single +kind of argument, so the callable stays a plain alternative to `stat()`. + +An earlier iteration also consulted the callable for each entry met during the +walk (passing the `os.DirEntry`, so that `annex_key_fingerprint` could skip +regular files without a system call), so that trees with dropped annexed files +could be cached. It was dropped: it made the callable part of the `stat()` +code path through a second, implicit contract, and it is better designed +together with the question below. + +**Open question.** The walk is O(number of entries) on every call. For very +large trees (Zarrs with millions of chunks, possibly with dropped chunks) a +cheaper fingerprint of the whole tree would be needed, e.g. the git tree hash +of a committed directory, or per-entry fingerprints within the walk as above. diff --git a/src/fscacher/__init__.py b/src/fscacher/__init__.py index 340e8cd..f024125 100644 --- a/src/fscacher/__init__.py +++ b/src/fscacher/__init__.py @@ -5,6 +5,7 @@ """ from ._version import get_versions +from .annex import annex_key_fingerprint from .cache import PersistentCache __version__ = get_versions()["version"] @@ -13,4 +14,4 @@ __license__ = "MIT" __url__ = "https://github.com/con/fscacher" -__all__ = ["PersistentCache"] +__all__ = ["PersistentCache", "annex_key_fingerprint"] diff --git a/src/fscacher/annex.py b/src/fscacher/annex.py new file mode 100644 index 0000000..f7591c8 --- /dev/null +++ b/src/fscacher/annex.py @@ -0,0 +1,52 @@ +"""Fingerprinting of git-annex'ed files by their keys""" + +import os +import os.path as op +import re + +#: Backends whose keys are hashes of the content, and thus pin it (with or +#: without the ``E`` suffix, which adds the file extension to the key). Others +#: do not: ``WORM`` keys are made of size, mtime and file name, ``URL`` and +#: ``VURL`` keys of a URL whose content may change, and external ``X*`` +#: backends are unknown. +CONTENT_HASH_BACKEND_RE = re.compile( + r"(SHA(1|224|256|384|512)|SHA3_(224|256|384|512)|SKEIN(256|512)" + r"|BLAKE2(B|BP|S|SP)\d+|MD5)E?" +) + + +def annex_key_fingerprint(path, *, pair_with_path=True): + """ + Fingerprint a locked git-annex'ed file by its key, for use as the + ``custom_fingerprint`` of `PersistentCache.memoize_path` + + Returns the key if ``path`` is a symbolic link into a git-annex object + store (``.../annex/objects/.../KEY/KEY``) and the key's backend hashes the + content (see `CONTENT_HASH_BACKEND_RE`); otherwise returns `None`, so that + the file is fingerprinted by ``stat()`` as usual. That is the case of an + unlocked file (including files on an adjusted branch, as on Windows), + whose key is not updated when the file is modified. Only the link is read, + so this is cheap, and a file whose content is not present (e.g., dropped) + is still fingerprinted: cached results are then returned for it. + + Parameters + ---------- + pair_with_path: bool, optional + If true (the default), the fingerprint is the pair of the absolute path + (as given, not dereferenced) and the key, so results are only shared by + files at the same path, as needed when they depend on the path (e.g., on + the extension or on neighboring files). If false, results are shared by + all files with the same key, e.g., across clones of a dataset. + """ + try: + path = os.fsdecode(os.fspath(path)) + target = os.fsdecode(os.readlink(path)) + except (TypeError, ValueError, OSError): + return None + parts = target.replace("\\", "/").split("/") + if len(parts) < 6 or parts[-6:-4] != ["annex", "objects"] or parts[-1] != parts[-2]: + return None + key = parts[-1] + if not CONTENT_HASH_BACKEND_RE.fullmatch(key.split("-", 1)[0]): + return None + return (op.abspath(path), key) if pair_with_path else key diff --git a/src/fscacher/cache.py b/src/fscacher/cache.py index b0a1758..3739565 100644 --- a/src/fscacher/cache.py +++ b/src/fscacher/cache.py @@ -84,9 +84,30 @@ def memoize(self, f=None, *, exclude_kwargs=None): return f return self._memory.cache(f, ignore=exclude_kwargs) - def memoize_path(self, f=None, *, exclude_kwargs=None): + def memoize_path(self, f=None, *, exclude_kwargs=None, custom_fingerprint=None): + """ + Memoize a function whose first argument is a path, keyed on a + fingerprint of the file or directory at that path + + Parameters + ---------- + exclude_kwargs: list of str, optional + Names of arguments of the decorated function to ignore for caching + purposes + custom_fingerprint: callable, optional + An alternative to the built-in ``stat()``-based fingerprint, e.g. + `fscacher.annex.annex_key_fingerprint`. It is called with the value + of the first argument only, and returns either a fingerprint or + `None` to fall back to ``stat()``. For a directory, it may return a + fingerprint of the whole tree, or `None` to fingerprint it by + ``stat()``-ing each file as usual. + """ if f is None: - return partial(self.memoize_path, exclude_kwargs=exclude_kwargs) + return partial( + self.memoize_path, + exclude_kwargs=exclude_kwargs, + custom_fingerprint=custom_fingerprint, + ) if self._ignore_cache: return f @@ -128,38 +149,26 @@ def fingerprinter(*args, **kwargs): bound = sig.bind(*args, **kwargs) bound.apply_defaults() path_orig = bound.arguments[path_arg] - try: - path = op.realpath(path_orig) - except TypeError: + fprint = self._get_fingerprint(path_orig, custom_fingerprint) + if fprint is None: lgr.debug( - "Calling %s directly since argument is not a path-like object", f + "Calling %s directly since no fingerprint for %r", f, path_orig ) - return f(*args, **kwargs) - if path != path_orig: - lgr.log(5, "Dereferenced %r into %r", path_orig, path) - if op.isdir(path): - fprint = self._get_dir_fingerprint(path) - else: - fprint = self._get_file_fingerprint(path) - if fprint is None: - lgr.debug("Calling %s directly since no fingerprint for %r", f, path) - # just call the function -- we have no fingerprint, - # probably does not exist or permissions are wrong + # just call the function -- we have no fingerprint, probably + # not a path, does not exist, or permissions are wrong ret = f(*args, **kwargs) # We should still pass through if file was modified just now, # since that could mask out quick modifications. # Target use cases will not be like that. elif fprint.modified_in_window(self._min_dtime): - lgr.debug("Calling %s directly since too short for %r", f, path) + lgr.debug("Calling %s directly since too short for %r", f, path_orig) ret = f(*args, **kwargs) else: - lgr.debug("Calling memoized version of %s for %s", f, path) + lgr.debug("Calling memoized version of %s for %r", f, path_orig) # If there is a fingerprint -- inject it into the signature kwargs_ = kwargs.copy() - kwargs_[fingerprint_kwarg] = ( - (path,) - + fprint.to_tuple() - + (tuple(self._tokens) if self._tokens else ()) + kwargs_[fingerprint_kwarg] = fprint.to_tuple() + ( + tuple(self._tokens) if self._tokens else () ) ret = fingerprinted(*args, **kwargs_) lgr.log(1, "Returning value %r", ret) @@ -168,6 +177,29 @@ def fingerprinter(*args, **kwargs): # and we memoize actually that function return fingerprinter + def _get_fingerprint(self, value, custom_fingerprint=None): + """ + Fingerprint of ``value``: the custom one if any, else one of the file + or directory at its dereferenced path, or `None` + """ + if custom_fingerprint is not None: + custom = custom_fingerprint(value) + if custom is not None: + lgr.log(5, "Custom fingerprint for %r: %r", value, custom) + return CustomFingerprint(custom) + try: + path = op.realpath(value) + except TypeError: + lgr.debug("Cannot fingerprint %r: not a path-like object", value) + return None + if path != value: + lgr.log(5, "Dereferenced %r into %r", value, path) + if op.isdir(path): + fprint = self._get_dir_fingerprint(path) + else: + fprint = self._get_file_fingerprint(path) + return None if fprint is None else PathFingerprint(path, fprint) + @staticmethod def _get_file_fingerprint(path): """Simplistic generic file fingerprinting based on ctime, mtime, and size""" @@ -214,6 +246,35 @@ def to_tuple(self): return tuple(self) +class PathFingerprint: + """A fingerprint of a file or directory, together with its path""" + + def __init__(self, path, fprint): + self.path = path + self.fprint = fprint + + def modified_in_window(self, min_dtime): + return self.fprint.modified_in_window(min_dtime) + + def to_tuple(self): + return (self.path,) + self.fprint.to_tuple() + + +class CustomFingerprint: + """A fingerprint returned by a ``custom_fingerprint`` callable""" + + def __init__(self, value): + self.value = value + + def modified_in_window(self, min_dtime): # noqa: U100 + # it is up to the callable to change the fingerprint on modification + return False + + def to_tuple(self): + # a path key starts with an absolute path, never with this marker + return ("custom", self.value) + + class DirFingerprint: def __init__(self): self.last_modified = None diff --git a/src/fscacher/tests/test_annex.py b/src/fscacher/tests/test_annex.py new file mode 100644 index 0000000..1dd2622 --- /dev/null +++ b/src/fscacher/tests/test_annex.py @@ -0,0 +1,274 @@ +from functools import partial +import os +import os.path as op +from pathlib import Path +import shutil +import subprocess +import time +import pytest +from .. import PersistentCache, annex_key_fingerprint + +KEY = "SHA256E-s7--0123456789abcdef.dat" + + +@pytest.fixture(scope="function") +def cache(tmp_path_factory): + return PersistentCache(path=tmp_path_factory.mktemp("cache")) + + +def annex_link(repo: Path, relpath: str, key: str, content="content") -> Path: + """ + Create a locked annexed file in a fake annex layout; with ``content=None``, + its content is missing (as if dropped) + """ + obj = repo / ".git" / "annex" / "objects" / "Xx" / "Yy" / key / key + obj.parent.mkdir(parents=True, exist_ok=True) + if content is not None: + obj.write_text(content) + link = repo / relpath + link.parent.mkdir(parents=True, exist_ok=True) + try: + os.symlink(op.relpath(obj, link.parent), link) + except OSError: + pytest.skip("symlinks are not supported here") + return link + + +def drop(link: Path) -> None: + os.unlink(op.realpath(link)) + + +@pytest.mark.parametrize( + "key", + [ + KEY, + "SHA256-s7--0123456789abcdef", + "SHA1-s7--0123", + "SHA512E-s7--0123.nwb", + "SHA3_256E-s7--0123.dat", + "SKEIN512-s7--0123", + "BLAKE2B256E-s7--0123.dat", + "BLAKE2SP224-s7--0123", + "MD5E-s7--0123.dat", + ], +) +def test_annex_key_fingerprint_content_hash_backends(tmp_path, monkeypatch, key): + link = annex_link(tmp_path, "sub/file.dat", key) + assert annex_key_fingerprint(link) == (str(link), key) + assert annex_key_fingerprint(str(link)) == (str(link), key) + assert annex_key_fingerprint(os.fsencode(link)) == (str(link), key) + assert annex_key_fingerprint(link, pair_with_path=False) == key + # Relative paths are paired as absolute ones + monkeypatch.chdir(tmp_path) + assert annex_key_fingerprint(op.join("sub", "file.dat")) == (str(link), key) + + +@pytest.mark.parametrize( + "key", + [ + "WORM-s7-m1700000000--file.dat", + "URL--https&c%%example.com%file.dat", + "VURL--https&c%%example.com%file.dat", + "XFOO-s7--0123", + "SHA256Z-s7--0123", + ], +) +def test_annex_key_fingerprint_other_backends(tmp_path, key): + assert annex_key_fingerprint(annex_link(tmp_path, "file.dat", key)) is None + + +def test_annex_key_fingerprint_not_annexed(tmp_path): + regular = tmp_path / "regular.dat" + regular.write_text("content") + other = tmp_path / "other.dat" + try: + # a link to a file named like a key, but not in an annex object store + (tmp_path / KEY).mkdir() + (tmp_path / KEY / KEY).write_text("content") + os.symlink(op.join(KEY, KEY), other) + except OSError: + pytest.skip("symlinks are not supported here") + for value in [regular, other, tmp_path, tmp_path / "missing", 42, None]: + assert annex_key_fingerprint(value) is None + + +def test_annex_key_fingerprint_dropped(tmp_path): + link = annex_link(tmp_path, "file.dat", KEY, content=None) + assert not op.exists(link) + assert annex_key_fingerprint(link) == (str(link), KEY) + + +def make_reader(cache, calls, **kwargs): + @cache.memoize_path(**kwargs) + def read(path): + calls.append(str(path)) + with open(path) as f: + return f"{op.basename(path)}:{f.read()}" + + return read + + +@pytest.mark.parametrize("pair_with_path", [True, False]) +def test_memoize_path_annex_twins(cache, tmp_path, pair_with_path): + calls = [] + read = make_reader( + cache, + calls, + custom_fingerprint=partial( + annex_key_fingerprint, pair_with_path=pair_with_path + ), + ) + # Twins: same key at different paths, with different extensions + a = annex_link(tmp_path, "a.dat", KEY) + b = annex_link(tmp_path, "sub/b.nwb", KEY) + assert read(a) == "a.dat:content" + assert read(a) == "a.dat:content" + assert calls == [str(a)] + if pair_with_path: + # each path has its own entry, as the result may depend on it + assert read(b) == "b.nwb:content" + assert read(b) == "b.nwb:content" + assert calls == [str(a), str(b)] + else: + # the result is shared, even though here it depends on the path + assert read(b) == "a.dat:content" + assert calls == [str(a)] + + +def test_memoize_path_annex_dropped(cache, tmp_path): + calls = [] + read = make_reader(cache, calls, custom_fingerprint=annex_key_fingerprint) + link = annex_link(tmp_path, "file.dat", KEY) + assert read(link) == "file.dat:content" + drop(link) + # The result cached while the content was present is still returned + assert read(link) == "file.dat:content" + assert calls == [str(link)] + # A file whose content was never present is read (and fails) every time + missing = annex_link(tmp_path, "missing.dat", "SHA256E-s3--fedcba.dat", None) + for _ in range(2): + with pytest.raises(FileNotFoundError): + read(missing) + assert calls == [str(link), str(missing), str(missing)] + + +def test_memoize_path_annex_fallback_to_stat(cache, tmp_path): + calls = [] + read = make_reader(cache, calls, custom_fingerprint=annex_key_fingerprint) + # A WORM key does not vouch for the content: stat() is used, which cannot + # fingerprint a dropped file, so the function is called every time + worm = annex_link(tmp_path, "worm.dat", "WORM-s7-m1--worm.dat", None) + for _ in range(2): + with pytest.raises(FileNotFoundError): + read(worm) + assert calls == [str(worm)] * 2 + # An unlocked file (a regular file) is fingerprinted by stat(), so its + # modifications are noticed + calls.clear() + unlocked = tmp_path / "unlocked.dat" + unlocked.write_text("content") + time.sleep(cache._min_dtime * 1.1) + assert read(unlocked) == "unlocked.dat:content" + assert read(unlocked) == "unlocked.dat:content" + time.sleep(cache._min_dtime * 1.1) + unlocked.write_text("edited") + time.sleep(cache._min_dtime * 1.1) + assert read(unlocked) == "unlocked.dat:edited" + assert calls == [str(unlocked)] * 2 + + +def test_memoize_path_annex_directory(cache, tmp_path): + seen = [] + + def fingerprint(path): + seen.append(path) + return annex_key_fingerprint(path) + + calls = [] + + @cache.memoize_path(custom_fingerprint=fingerprint) + def listdir(path): + calls.append(str(path)) + return sorted( + (p.relative_to(path).as_posix(), op.exists(p)) + for p in Path(path).rglob("*") + if not p.is_dir() or p.is_symlink() + ) + + ds = tmp_path / "ds" + # the object store lives outside of the directory, as for a .zarr in a + # dataset + for name in ["a.dat", "sub/b.dat"]: + annex_link(tmp_path, f"ds/{name}", f"SHA256E-s7--{name[-5]}.dat") + (ds / "regular.txt").write_text("regular") + time.sleep(cache._min_dtime * 1.1) + # The callable only sees the top-level argument, for which it returns None, + # so the directory is fingerprinted by stat()-ing each file, as without it + expected = [("a.dat", True), ("regular.txt", True), ("sub/b.dat", True)] + assert listdir(ds) == listdir(ds) == expected + assert calls == [str(ds)] + assert seen == [ds, ds] + # ... and thus a dropped file prevents fingerprinting the directory + drop(ds / "sub" / "b.dat") + expected[-1] = ("sub/b.dat", False) + assert listdir(ds) == listdir(ds) == expected + assert calls == [str(ds)] * 3 + + +@pytest.mark.skipif(shutil.which("git-annex") is None, reason="git annex required") +def test_memoize_path_git_annex(cache, tmp_path, monkeypatch): + for var in ["GIT_AUTHOR_NAME", "GIT_COMMITTER_NAME"]: + monkeypatch.setenv(var, "Test") + for var in ["GIT_AUTHOR_EMAIL", "GIT_COMMITTER_EMAIL"]: + monkeypatch.setenv(var, "test@example.com") + + def git(*args): + subprocess.run(["git", *args], cwd=tmp_path, check=True) + + git("init", "-q") + git("annex", "init", "-q") + git("config", "annex.backend", "SHA256E") + (tmp_path / "a.dat").write_text("content") + (tmp_path / "sub").mkdir() + (tmp_path / "sub" / "b.dat").write_text("content") + (tmp_path / "c.dat").write_text("other") + (tmp_path / "w.dat").write_text("worm") + git("annex", "add", "-q", "a.dat", "sub", "c.dat") + git("-c", "annex.backend=WORM", "annex", "add", "-q", "w.dat") + git("commit", "-q", "-m", "Add files") + a, b, c, w = (tmp_path / p for p in ["a.dat", "sub/b.dat", "c.dat", "w.dat"]) + if not op.islink(a): + pytest.skip("git-annex does not lock files here (adjusted branch?)") + + key = annex_key_fingerprint(a, pair_with_path=False) + assert key.startswith("SHA256E-s7--") + assert annex_key_fingerprint(b, pair_with_path=False) == key + assert annex_key_fingerprint(w) is None + + calls = [] + read = make_reader( + cache, + calls, + custom_fingerprint=partial(annex_key_fingerprint, pair_with_path=False), + ) + assert read(a) == "a.dat:content" + assert read(b) == "a.dat:content" + assert calls == [str(a)] + + # Unlocked, a file is fingerprinted by stat(), and edits are noticed + assert read(c) == "c.dat:other" + git("annex", "unlock", "-q", "c.dat") + assert annex_key_fingerprint(c) is None + time.sleep(cache._min_dtime * 1.1) + assert read(c) == "c.dat:other" + time.sleep(cache._min_dtime * 1.1) + c.write_text("edited") + time.sleep(cache._min_dtime * 1.1) + assert read(c) == "c.dat:edited" + assert calls == [str(a), str(c), str(c), str(c)] + + # Dropped, a locked file's cached result is still returned + git("annex", "drop", "-q", "--force", "a.dat") + assert not op.exists(a) + assert read(a) == "a.dat:content" + assert len(calls) == 4 diff --git a/src/fscacher/tests/test_cache.py b/src/fscacher/tests/test_cache.py index 13e585e..ac7d3b6 100644 --- a/src/fscacher/tests/test_cache.py +++ b/src/fscacher/tests/test_cache.py @@ -8,6 +8,7 @@ import subprocess import sys import time +from typing import Optional import pytest from .. import PersistentCache from ..cache import DirFingerprint, FileFingerprint @@ -558,3 +559,93 @@ def memoread_extra(path, arg, kwarg=None, extra=None): (path, 1, "quux", "bar"), (path, 1, None, "foo"), ] + + +@dataclass +class Blob: + """A non-path resource that may know a fingerprint of its content""" + + content: str + fingerprint: Optional[str] = None + + +def blob_fingerprint(x): + return x.fingerprint if isinstance(x, Blob) else None + + +def test_memoize_path_custom_fingerprint(cache, tmp_path): + calls = [] + + @cache.memoize_path(custom_fingerprint=blob_fingerprint) + def memoread(src, arg=0, kwarg=None): + calls.append((src, arg, kwarg)) + if isinstance(src, Blob): + return f"{src.content}:{arg}:{kwarg}" + with open(src) as f: + return f"{f.read()}:{arg}:{kwarg}" + + # A value without a fingerprint (nor a path) is not cached + assert memoread(Blob("content")) == "content:0:None" + assert memoread(Blob("content")) == "content:0:None" + assert len(calls) == 2 + + # A value with one is cached by it: a twin is served from the cache, + # however its other arguments are passed + assert memoread(Blob("content", "A"), 1) == "content:1:None" + assert len(calls) == 3 + assert memoread(Blob("never read", "A"), 1) == "content:1:None" + assert memoread(Blob("never read", "A"), arg=1) == "content:1:None" + assert memoread(src=Blob("never read", "A"), arg=1) == "content:1:None" + assert len(calls) == 3 + + # The fingerprint and the other arguments are part of the key + assert memoread(Blob("other", "B"), 1) == "other:1:None" + assert memoread(Blob("content", "A"), 1, kwarg="q") == "content:1:q" + assert len(calls) == 5 + + # Paths, for which the callable returns None, are fingerprinted by stat() + path = tmp_path / "file.dat" + path.write_text("content") + time.sleep(cache._min_dtime * 1.1) + assert memoread(path, 1) == "content:1:None" + assert memoread(path, 1) == "content:1:None" + assert len(calls) == 6 + time.sleep(cache._min_dtime * 1.1) + path.write_text("changed") + time.sleep(cache._min_dtime * 1.1) + assert memoread(path, 1) == "changed:1:None" + assert len(calls) == 7 + + +def test_memoize_path_custom_fingerprint_tokens(tmp_path_factory): + calls = [] + + def memoread(src): + calls.append(src) + return src.content + + path = tmp_path_factory.mktemp("cache") + c1 = PersistentCache(path=path, tokens=["1"]) + c2 = PersistentCache(path=path, tokens=["2"]) + m1 = c1.memoize_path(memoread, custom_fingerprint=blob_fingerprint) + m2 = c2.memoize_path(memoread, custom_fingerprint=blob_fingerprint) + assert m1(Blob("content", "A")) == "content" + assert m1(Blob("never read", "A")) == "content" + assert len(calls) == 1 + assert m2(Blob("content", "A")) == "content" + assert len(calls) == 2 + + +def test_memoize_path_custom_fingerprint_ignored(monkeypatch, tmp_path): + monkeypatch.setenv("FSCACHER_CACHE", "ignore") + cache = PersistentCache(path=tmp_path) + calls = [] + + @cache.memoize_path(custom_fingerprint=blob_fingerprint) + def memoread(src): + calls.append(src) + return src.content + + assert memoread(Blob("content", "A")) == "content" + assert memoread(Blob("content", "A")) == "content" + assert len(calls) == 2