Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 13 additions & 0 deletions README.rst
Original file line number Diff line number Diff line change
Expand Up @@ -48,6 +48,19 @@ of ``path``, the cache is ignored.
be a sequence of names of arguments of the decorated function that will be
ignored for caching purposes.

``memoize_path()`` can also take a ``custom_fingerprint`` callable to use
instead of ``stat()``; it returns ``None`` to fall back to ``stat()``. For
example, ``fscacher.annex_key_fingerprint`` fingerprints locked git-annex'ed
files by their keys, so results survive their content being dropped:

.. code:: python

from fscacher import annex_key_fingerprint

@cache.memoize_path(custom_fingerprint=annex_key_fingerprint)
def foo(path, ...):
...

Caches are stored on-disk and thus persist between Python runs. To clear a
given ``PersistentCache`` and erase its data store, call the ``clear()``
method.
Expand Down
133 changes: 133 additions & 0 deletions docs/design/git-annex-content.md

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This file is for more of those lower-level design concepts for the AI to refer to (a human can perhaps glance and clarify major details)

It also contains thoughts/discussions of edge cases like WSL for reference

(hiding away in commit messages makes it a tad harder to introspect / search IMO)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

in general -- I do not mind those. But that is where I feel adhering to some framework like spec-kit is benefitical since keeping such files in sync with the code is part of the process. Otherwise they can get "out of sync"

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what part of spec-kit keeps already-specified docs up-to-date?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the /speckit-analyze

Sure but that doesn't work automatically 'on its own' and I bet it would work on design docs like these just fine; all one would need is a 'routine' (weekly? monthly? when releases are cut?) to regularly scan the repo with that command

Original file line number Diff line number Diff line change
@@ -0,0 +1,133 @@
# Custom fingerprints and git-annex content

Design notes for `memoize_path(custom_fingerprint=...)` and
`fscacher.annex_key_fingerprint`, introduced in
[#113](https://github.com/con/fscacher/pull/113). The docstrings and README
only say what these are; this records why they work the way they do.

## Motivation

`memoize_path` keys its cache on a `stat()` of the path argument (mtime,
ctime, size, inode). Some resources can vouch for their content better:

- A *locked* git-annex'ed file is a symlink whose target names a key; for
content-hash backends, the key pins the content. It also survives the file
being moved, re-cloned, or its content dropped.
- Non-path objects (e.g. dandi-cli's `Readable`s, which may stream remote
content) have no `stat()` at all, but may know a fingerprint of their
content.

## `custom_fingerprint`: an alternative to `stat()`, not a second code path

The callable replaces only the `stat()`-based fingerprint; everything else
(falling back to a direct call when there is no fingerprint, injecting the
fingerprint into the cache key, tokens) is shared. `_get_fingerprint()`
returns either a `CustomFingerprint` or a `PathFingerprint`, or `None`, and the
rest of `memoize_path` does not care which. `PathFingerprint` keys exactly as
before, so existing caches stay valid.

The contract for a callable:

- **It returns `None` to fall back to `stat()`**, for anything it does not
recognize, including plain paths unless it fingerprints them. It must not
raise for such values: an exception propagates, so the decorated call
fails. This is deliberate, as an exception there most likely is a bug in
the callable, which should not silently disable caching. (A way to say "do
not cache", e.g. a named `fscacher.NO_CACHE` constant, can be added if ever
needed.)
- **It runs on every call** of the decorated function, so it must be cheap.
`annex_key_fingerprint` only reads the symlink; it never runs git-annex.
- **Equal fingerprints share results, wherever they are.** The path argument
itself is not part of the cache key: in `stat()` mode the (dereferenced) path
is part of the fingerprint, but a custom fingerprint is used as is. So a
callable must include the path in its fingerprint unless the result does not
depend on it, as `annex_key_fingerprint` does by default (see
`pair_with_path` below). For the same reason the value need not be a path
at all.
- **Fingerprints must be picklable with a stable `repr()`**, as joblib hashes
them into the key (e.g. a string or a tuple of strings).
- **There is no "modified just now" window.** With `stat()`, a file modified
within `_min_dtime` is not cached, since a quick later modification might
not change the fingerprint. A custom fingerprint is trusted to change
whenever the result may change; the callable is responsible for that. This
is only justified for fingerprints that pin the content, like hash-based
annex keys.

## Which git-annex files are fingerprinted by key

Only **locked** files: a symlink whose target looks like
`.../annex/objects/*/*/KEY/KEY`. The link is read with `readlink()` only.

Only **content-hash backends**: `SHA1`, `SHA224`/`256`/`384`/`512`,
`SHA3_224`/`256`/`384`/`512`, `SKEIN256`/`512`, `BLAKE2B*`/`BP*`/`S*`/`SP*`,
and `MD5`, each with or without the `E` suffix (which adds the extension).
Everything else falls back to `stat()`:

- `WORM` keys are made of size, mtime and file name, so they do not pin the
content.
- `URL` and `VURL` keys name a URL whose content may change.
- External `X*` backends are unknown.

**Unlocked files** also fall back to `stat()`: an unlocked file is a regular
file whose key is not updated until it is re-added, so an edit would not
change the key. This includes every file on an *adjusted unlocked branch*,
which git-annex uses on filesystems without usable symlinks or permissions,
notably native Windows. There, `annex_key_fingerprint` is a no-op.

WSL behaves like Linux on its own filesystem (e.g. under `~`). On a Windows
drive under `/mnt/c`, git-annex is likely to treat the filesystem as crippled
and use an adjusted branch, so the same `stat()` fallback applies. Neither is
tested in CI.

## `pair_with_path`

By default the fingerprint is `(absolute path, key)`: the path as given, not
dereferenced (dereferencing would give the object path, identical for all
files with the key). Results are then only shared by files at the same path,
which is needed whenever the result depends on the path, e.g. on the extension
(the key's extension may differ from the file's) or on neighboring files.
The path is made absolute with `os.path.abspath`, which does not resolve
symlinks (neither the file's nor its parent directories'), so the same file
reached through different paths gets separate entries; that errs on the safe
side.

With `pair_with_path=False`, the fingerprint is the key alone, and results are
shared by all files with the same content: twins in a dataset, and the same
file across clones. Only use it for results that depend on the content only.

## Dropped content

With `stat()`, a locked file whose content was dropped is a broken symlink: it
cannot be fingerprinted, so the function is called directly every time.

With the key, it is still fingerprinted, so a result cached while the content
was present is returned after `git annex drop`. This is useful (e.g. metadata
of files no longer present locally) and deliberate. A file whose content was
never present has no cached result, so the function is called and fails as
before.

## Directories

`memoize_path` fingerprints a directory argument by walking it and combining
the `stat()` of every entry (this predates custom fingerprints). With `stat()`
alone, a single dropped annexed file (a broken symlink) makes the whole
directory unfingerprintable, and the mtimes of annexed files feed the
"modified just now" window.

**The callable only sees the top-level argument.** For a directory, it may
return a fingerprint of the whole tree itself; if it returns `None` (as
`annex_key_fingerprint` does), the directory is walked and `stat()`-ed exactly
as before, dropped files included. This keeps a single call site and a single
kind of argument, so the callable stays a plain alternative to `stat()`.

An earlier iteration also consulted the callable for each entry met during the
walk (passing the `os.DirEntry`, so that `annex_key_fingerprint` could skip
regular files without a system call), so that trees with dropped annexed files
could be cached. It was dropped: it made the callable part of the `stat()`
code path through a second, implicit contract, and it is better designed
together with the question below.

**Open question.** The walk is O(number of entries) on every call. For very
large trees (Zarrs with millions of chunks, possibly with dropped chunks) a
cheaper fingerprint of the whole tree would be needed, e.g. the git tree hash
of a committed directory, or per-entry fingerprints within the walk as above.
3 changes: 2 additions & 1 deletion src/fscacher/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@
"""

from ._version import get_versions
from .annex import annex_key_fingerprint
from .cache import PersistentCache

__version__ = get_versions()["version"]
Expand All @@ -13,4 +14,4 @@
__license__ = "MIT"
__url__ = "https://github.com/con/fscacher"

__all__ = ["PersistentCache"]
__all__ = ["PersistentCache", "annex_key_fingerprint"]
52 changes: 52 additions & 0 deletions src/fscacher/annex.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,52 @@
"""Fingerprinting of git-annex'ed files by their keys"""

import os
import os.path as op
import re

#: Backends whose keys are hashes of the content, and thus pin it (with or
#: without the ``E`` suffix, which adds the file extension to the key). Others
#: do not: ``WORM`` keys are made of size, mtime and file name, ``URL`` and
#: ``VURL`` keys of a URL whose content may change, and external ``X*``
#: backends are unknown.
CONTENT_HASH_BACKEND_RE = re.compile(
r"(SHA(1|224|256|384|512)|SHA3_(224|256|384|512)|SKEIN(256|512)"
r"|BLAKE2(B|BP|S|SP)\d+|MD5)E?"
)


def annex_key_fingerprint(path, *, pair_with_path=True):
"""
Fingerprint a locked git-annex'ed file by its key, for use as the
``custom_fingerprint`` of `PersistentCache.memoize_path`

Returns the key if ``path`` is a symbolic link into a git-annex object
store (``.../annex/objects/.../KEY/KEY``) and the key's backend hashes the
content (see `CONTENT_HASH_BACKEND_RE`); otherwise returns `None`, so that
the file is fingerprinted by ``stat()`` as usual. That is the case of an
unlocked file (including files on an adjusted branch, as on Windows),
whose key is not updated when the file is modified. Only the link is read,
so this is cheap, and a file whose content is not present (e.g., dropped)
is still fingerprinted: cached results are then returned for it.

Parameters
----------
pair_with_path: bool, optional
If true (the default), the fingerprint is the pair of the absolute path
(as given, not dereferenced) and the key, so results are only shared by
files at the same path, as needed when they depend on the path (e.g., on
the extension or on neighboring files). If false, results are shared by
all files with the same key, e.g., across clones of a dataset.
"""
try:
path = os.fsdecode(os.fspath(path))

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

need to remember about this fsdecode... not sure I ever used it ...

target = os.fsdecode(os.readlink(path))
except (TypeError, ValueError, OSError):
return None
parts = target.replace("\\", "/").split("/")
if len(parts) < 6 or parts[-6:-4] != ["annex", "objects"] or parts[-1] != parts[-2]:
return None
key = parts[-1]
if not CONTENT_HASH_BACKEND_RE.fullmatch(key.split("-", 1)[0]):
return None
return (op.abspath(path), key) if pair_with_path else key

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

note to paranoid me (smth negging in the back of my mind) -- abspath does not resolve links!

❯ python -c 'import os.path; print(os.path.abspath("sub-qa_ses-20200102_acq-MPRAGE_T1w.nii.gz"))'
/home/yoh/datalad/dbic/QA/sub-qa/ses-20200102/anat/sub-qa_ses-20200102_acq-MPRAGE_T1w.nii.gz
❯ dg sub-qa_ses-20200102_acq-MPRAGE_T1w.nii.gz
get(ok): sub-qa/ses-20200102/anat/sub-qa_ses-20200102_acq-MPRAGE_T1w.nii.gz (file) [from origin...]                                                            
❯ python -c 'import os.path; print(os.path.abspath("sub-qa_ses-20200102_acq-MPRAGE_T1w.nii.gz"))'
/home/yoh/datalad/dbic/QA/sub-qa/ses-20200102/anat/sub-qa_ses-20200102_acq-MPRAGE_T1w.nii.gz
❯ python -c 'import os.path; print(os.path.abspath("./sub-qa_ses-20200102_acq-MPRAGE_T1w.nii.gz"))'
/home/yoh/datalad/dbic/QA/sub-qa/ses-20200102/anat/sub-qa_ses-20200102_acq-MPRAGE_T1w.nii.gz

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeh, many reasons why I personally prefer pathlib, which then has .resolve() as an action to be explicit

107 changes: 84 additions & 23 deletions src/fscacher/cache.py
Original file line number Diff line number Diff line change
Expand Up @@ -84,9 +84,30 @@ def memoize(self, f=None, *, exclude_kwargs=None):
return f
return self._memory.cache(f, ignore=exclude_kwargs)

def memoize_path(self, f=None, *, exclude_kwargs=None):
def memoize_path(self, f=None, *, exclude_kwargs=None, custom_fingerprint=None):
"""
Memoize a function whose first argument is a path, keyed on a
fingerprint of the file or directory at that path

Parameters
----------
exclude_kwargs: list of str, optional
Names of arguments of the decorated function to ignore for caching
purposes
custom_fingerprint: callable, optional
An alternative to the built-in ``stat()``-based fingerprint, e.g.
`fscacher.annex.annex_key_fingerprint`. It is called with the value
of the first argument only, and returns either a fingerprint or
`None` to fall back to ``stat()``. For a directory, it may return a
fingerprint of the whole tree, or `None` to fingerprint it by
``stat()``-ing each file as usual.
"""
if f is None:
return partial(self.memoize_path, exclude_kwargs=exclude_kwargs)
return partial(
self.memoize_path,
exclude_kwargs=exclude_kwargs,
custom_fingerprint=custom_fingerprint,
)
if self._ignore_cache:
return f

Expand Down Expand Up @@ -128,38 +149,26 @@ def fingerprinter(*args, **kwargs):
bound = sig.bind(*args, **kwargs)
bound.apply_defaults()
path_orig = bound.arguments[path_arg]
try:
path = op.realpath(path_orig)
except TypeError:
fprint = self._get_fingerprint(path_orig, custom_fingerprint)
if fprint is None:
lgr.debug(
"Calling %s directly since argument is not a path-like object", f
"Calling %s directly since no fingerprint for %r", f, path_orig
)
return f(*args, **kwargs)
if path != path_orig:
lgr.log(5, "Dereferenced %r into %r", path_orig, path)
if op.isdir(path):
fprint = self._get_dir_fingerprint(path)
else:
fprint = self._get_file_fingerprint(path)
if fprint is None:
lgr.debug("Calling %s directly since no fingerprint for %r", f, path)
# just call the function -- we have no fingerprint,
# probably does not exist or permissions are wrong
# just call the function -- we have no fingerprint, probably
# not a path, does not exist, or permissions are wrong
ret = f(*args, **kwargs)
# We should still pass through if file was modified just now,
# since that could mask out quick modifications.
# Target use cases will not be like that.
elif fprint.modified_in_window(self._min_dtime):
lgr.debug("Calling %s directly since too short for %r", f, path)
lgr.debug("Calling %s directly since too short for %r", f, path_orig)
ret = f(*args, **kwargs)
else:
lgr.debug("Calling memoized version of %s for %s", f, path)
lgr.debug("Calling memoized version of %s for %r", f, path_orig)
# If there is a fingerprint -- inject it into the signature
kwargs_ = kwargs.copy()
kwargs_[fingerprint_kwarg] = (
(path,)
+ fprint.to_tuple()
+ (tuple(self._tokens) if self._tokens else ())
kwargs_[fingerprint_kwarg] = fprint.to_tuple() + (
tuple(self._tokens) if self._tokens else ()
)
ret = fingerprinted(*args, **kwargs_)
lgr.log(1, "Returning value %r", ret)
Expand All @@ -168,6 +177,29 @@ def fingerprinter(*args, **kwargs):
# and we memoize actually that function
return fingerprinter

def _get_fingerprint(self, value, custom_fingerprint=None):
"""
Fingerprint of ``value``: the custom one if any, else one of the file
or directory at its dereferenced path, or `None`
"""
if custom_fingerprint is not None:
custom = custom_fingerprint(value)
if custom is not None:
lgr.log(5, "Custom fingerprint for %r: %r", value, custom)
return CustomFingerprint(custom)
try:
path = op.realpath(value)
except TypeError:
lgr.debug("Cannot fingerprint %r: not a path-like object", value)
return None
if path != value:
lgr.log(5, "Dereferenced %r into %r", value, path)
if op.isdir(path):
fprint = self._get_dir_fingerprint(path)
else:
fprint = self._get_file_fingerprint(path)
return None if fprint is None else PathFingerprint(path, fprint)

@staticmethod
def _get_file_fingerprint(path):
"""Simplistic generic file fingerprinting based on ctime, mtime, and size"""
Expand Down Expand Up @@ -214,6 +246,35 @@ def to_tuple(self):
return tuple(self)


class PathFingerprint:
"""A fingerprint of a file or directory, together with its path"""

def __init__(self, path, fprint):
self.path = path
self.fprint = fprint

def modified_in_window(self, min_dtime):
return self.fprint.modified_in_window(min_dtime)

def to_tuple(self):
return (self.path,) + self.fprint.to_tuple()


class CustomFingerprint:
"""A fingerprint returned by a ``custom_fingerprint`` callable"""

def __init__(self, value):
self.value = value

def modified_in_window(self, min_dtime): # noqa: U100
# it is up to the callable to change the fingerprint on modification
return False

def to_tuple(self):
# a path key starts with an absolute path, never with this marker
return ("custom", self.value)


class DirFingerprint:
def __init__(self):
self.last_modified = None
Expand Down
Loading
Loading