Repository navigation
Add custom_fingerprint to memoize_path, and a git-annex key fingerprint #113
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
b006d05
a43b1a0
a8ed903
c43232e
7dd8e7d
302f55d
a5a181a
e303605
ad4f67b
37ce6fd
e0eca3e
6468a25
0ae8e87
4f89cd7
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,133 @@ | ||
| # Custom fingerprints and git-annex content | ||
|
|
||
| Design notes for `memoize_path(custom_fingerprint=...)` and | ||
| `fscacher.annex_key_fingerprint`, introduced in | ||
| [#113](https://github.com/con/fscacher/pull/113). The docstrings and README | ||
| only say what these are; this records why they work the way they do. | ||
|
|
||
| ## Motivation | ||
|
|
||
| `memoize_path` keys its cache on a `stat()` of the path argument (mtime, | ||
| ctime, size, inode). Some resources can vouch for their content better: | ||
|
|
||
| - A *locked* git-annex'ed file is a symlink whose target names a key; for | ||
| content-hash backends, the key pins the content. It also survives the file | ||
| being moved, re-cloned, or its content dropped. | ||
| - Non-path objects (e.g. dandi-cli's `Readable`s, which may stream remote | ||
| content) have no `stat()` at all, but may know a fingerprint of their | ||
| content. | ||
|
|
||
| ## `custom_fingerprint`: an alternative to `stat()`, not a second code path | ||
|
|
||
| The callable replaces only the `stat()`-based fingerprint; everything else | ||
| (falling back to a direct call when there is no fingerprint, injecting the | ||
| fingerprint into the cache key, tokens) is shared. `_get_fingerprint()` | ||
| returns either a `CustomFingerprint` or a `PathFingerprint`, or `None`, and the | ||
| rest of `memoize_path` does not care which. `PathFingerprint` keys exactly as | ||
| before, so existing caches stay valid. | ||
|
|
||
| The contract for a callable: | ||
|
|
||
| - **It returns `None` to fall back to `stat()`**, for anything it does not | ||
| recognize, including plain paths unless it fingerprints them. It must not | ||
| raise for such values: an exception propagates, so the decorated call | ||
| fails. This is deliberate, as an exception there most likely is a bug in | ||
| the callable, which should not silently disable caching. (A way to say "do | ||
| not cache", e.g. a named `fscacher.NO_CACHE` constant, can be added if ever | ||
| needed.) | ||
| - **It runs on every call** of the decorated function, so it must be cheap. | ||
| `annex_key_fingerprint` only reads the symlink; it never runs git-annex. | ||
| - **Equal fingerprints share results, wherever they are.** The path argument | ||
| itself is not part of the cache key: in `stat()` mode the (dereferenced) path | ||
| is part of the fingerprint, but a custom fingerprint is used as is. So a | ||
| callable must include the path in its fingerprint unless the result does not | ||
| depend on it, as `annex_key_fingerprint` does by default (see | ||
| `pair_with_path` below). For the same reason the value need not be a path | ||
| at all. | ||
| - **Fingerprints must be picklable with a stable `repr()`**, as joblib hashes | ||
| them into the key (e.g. a string or a tuple of strings). | ||
| - **There is no "modified just now" window.** With `stat()`, a file modified | ||
| within `_min_dtime` is not cached, since a quick later modification might | ||
| not change the fingerprint. A custom fingerprint is trusted to change | ||
| whenever the result may change; the callable is responsible for that. This | ||
| is only justified for fingerprints that pin the content, like hash-based | ||
| annex keys. | ||
|
|
||
| ## Which git-annex files are fingerprinted by key | ||
|
|
||
| Only **locked** files: a symlink whose target looks like | ||
| `.../annex/objects/*/*/KEY/KEY`. The link is read with `readlink()` only. | ||
|
|
||
| Only **content-hash backends**: `SHA1`, `SHA224`/`256`/`384`/`512`, | ||
| `SHA3_224`/`256`/`384`/`512`, `SKEIN256`/`512`, `BLAKE2B*`/`BP*`/`S*`/`SP*`, | ||
| and `MD5`, each with or without the `E` suffix (which adds the extension). | ||
| Everything else falls back to `stat()`: | ||
|
|
||
| - `WORM` keys are made of size, mtime and file name, so they do not pin the | ||
| content. | ||
| - `URL` and `VURL` keys name a URL whose content may change. | ||
| - External `X*` backends are unknown. | ||
|
|
||
| **Unlocked files** also fall back to `stat()`: an unlocked file is a regular | ||
| file whose key is not updated until it is re-added, so an edit would not | ||
| change the key. This includes every file on an *adjusted unlocked branch*, | ||
| which git-annex uses on filesystems without usable symlinks or permissions, | ||
| notably native Windows. There, `annex_key_fingerprint` is a no-op. | ||
|
|
||
| WSL behaves like Linux on its own filesystem (e.g. under `~`). On a Windows | ||
| drive under `/mnt/c`, git-annex is likely to treat the filesystem as crippled | ||
| and use an adjusted branch, so the same `stat()` fallback applies. Neither is | ||
| tested in CI. | ||
|
|
||
| ## `pair_with_path` | ||
|
|
||
| By default the fingerprint is `(absolute path, key)`: the path as given, not | ||
| dereferenced (dereferencing would give the object path, identical for all | ||
| files with the key). Results are then only shared by files at the same path, | ||
| which is needed whenever the result depends on the path, e.g. on the extension | ||
| (the key's extension may differ from the file's) or on neighboring files. | ||
| The path is made absolute with `os.path.abspath`, which does not resolve | ||
| symlinks (neither the file's nor its parent directories'), so the same file | ||
| reached through different paths gets separate entries; that errs on the safe | ||
| side. | ||
|
|
||
| With `pair_with_path=False`, the fingerprint is the key alone, and results are | ||
| shared by all files with the same content: twins in a dataset, and the same | ||
| file across clones. Only use it for results that depend on the content only. | ||
|
|
||
| ## Dropped content | ||
|
|
||
| With `stat()`, a locked file whose content was dropped is a broken symlink: it | ||
| cannot be fingerprinted, so the function is called directly every time. | ||
|
|
||
| With the key, it is still fingerprinted, so a result cached while the content | ||
| was present is returned after `git annex drop`. This is useful (e.g. metadata | ||
| of files no longer present locally) and deliberate. A file whose content was | ||
| never present has no cached result, so the function is called and fails as | ||
| before. | ||
|
|
||
| ## Directories | ||
|
|
||
| `memoize_path` fingerprints a directory argument by walking it and combining | ||
| the `stat()` of every entry (this predates custom fingerprints). With `stat()` | ||
| alone, a single dropped annexed file (a broken symlink) makes the whole | ||
| directory unfingerprintable, and the mtimes of annexed files feed the | ||
| "modified just now" window. | ||
|
|
||
| **The callable only sees the top-level argument.** For a directory, it may | ||
| return a fingerprint of the whole tree itself; if it returns `None` (as | ||
| `annex_key_fingerprint` does), the directory is walked and `stat()`-ed exactly | ||
| as before, dropped files included. This keeps a single call site and a single | ||
| kind of argument, so the callable stays a plain alternative to `stat()`. | ||
|
|
||
| An earlier iteration also consulted the callable for each entry met during the | ||
| walk (passing the `os.DirEntry`, so that `annex_key_fingerprint` could skip | ||
| regular files without a system call), so that trees with dropped annexed files | ||
| could be cached. It was dropped: it made the callable part of the `stat()` | ||
| code path through a second, implicit contract, and it is better designed | ||
| together with the question below. | ||
|
|
||
| **Open question.** The walk is O(number of entries) on every call. For very | ||
| large trees (Zarrs with millions of chunks, possibly with dropped chunks) a | ||
| cheaper fingerprint of the whole tree would be needed, e.g. the git tree hash | ||
| of a committed directory, or per-entry fingerprints within the walk as above. |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,52 @@ | ||
| """Fingerprinting of git-annex'ed files by their keys""" | ||
|
|
||
| import os | ||
| import os.path as op | ||
| import re | ||
|
|
||
| #: Backends whose keys are hashes of the content, and thus pin it (with or | ||
| #: without the ``E`` suffix, which adds the file extension to the key). Others | ||
| #: do not: ``WORM`` keys are made of size, mtime and file name, ``URL`` and | ||
| #: ``VURL`` keys of a URL whose content may change, and external ``X*`` | ||
| #: backends are unknown. | ||
| CONTENT_HASH_BACKEND_RE = re.compile( | ||
| r"(SHA(1|224|256|384|512)|SHA3_(224|256|384|512)|SKEIN(256|512)" | ||
| r"|BLAKE2(B|BP|S|SP)\d+|MD5)E?" | ||
| ) | ||
|
|
||
|
|
||
| def annex_key_fingerprint(path, *, pair_with_path=True): | ||
| """ | ||
| Fingerprint a locked git-annex'ed file by its key, for use as the | ||
| ``custom_fingerprint`` of `PersistentCache.memoize_path` | ||
|
|
||
| Returns the key if ``path`` is a symbolic link into a git-annex object | ||
| store (``.../annex/objects/.../KEY/KEY``) and the key's backend hashes the | ||
| content (see `CONTENT_HASH_BACKEND_RE`); otherwise returns `None`, so that | ||
| the file is fingerprinted by ``stat()`` as usual. That is the case of an | ||
| unlocked file (including files on an adjusted branch, as on Windows), | ||
| whose key is not updated when the file is modified. Only the link is read, | ||
| so this is cheap, and a file whose content is not present (e.g., dropped) | ||
| is still fingerprinted: cached results are then returned for it. | ||
|
|
||
| Parameters | ||
| ---------- | ||
| pair_with_path: bool, optional | ||
| If true (the default), the fingerprint is the pair of the absolute path | ||
| (as given, not dereferenced) and the key, so results are only shared by | ||
| files at the same path, as needed when they depend on the path (e.g., on | ||
| the extension or on neighboring files). If false, results are shared by | ||
| all files with the same key, e.g., across clones of a dataset. | ||
| """ | ||
| try: | ||
| path = os.fsdecode(os.fspath(path)) | ||
|
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. need to remember about this |
||
| target = os.fsdecode(os.readlink(path)) | ||
| except (TypeError, ValueError, OSError): | ||
| return None | ||
| parts = target.replace("\\", "/").split("/") | ||
| if len(parts) < 6 or parts[-6:-4] != ["annex", "objects"] or parts[-1] != parts[-2]: | ||
| return None | ||
| key = parts[-1] | ||
| if not CONTENT_HASH_BACKEND_RE.fullmatch(key.split("-", 1)[0]): | ||
| return None | ||
| return (op.abspath(path), key) if pair_with_path else key | ||
|
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. note to paranoid me (smth negging in the back of my mind) -- abspath does not resolve links!
Contributor
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Yeh, many reasons why I personally prefer |
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
This file is for more of those lower-level design concepts for the AI to refer to (a human can perhaps glance and clarify major details)
It also contains thoughts/discussions of edge cases like WSL for reference
(hiding away in commit messages makes it a tad harder to introspect / search IMO)
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
in general -- I do not mind those. But that is where I feel adhering to some framework like spec-kit is benefitical since keeping such files in sync with the code is part of the process. Otherwise they can get "out of sync"
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
what part of spec-kit keeps already-specified docs up-to-date?
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
the
/speckit-analyze:https://github.github.com/spec-kit/quickstart.html#step-7-speckit-analyze--check-consistency
https://github.com/github/spec-kit/blob/main/templates/commands/analyze.md?plain=1
there is also apparently now a new command -
/speckit-convergeas wellhttps://github.github.com/spec-kit/quickstart.html#step-9-speckit-converge--verify-completeness
which relates
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Sure but that doesn't work automatically 'on its own' and I bet it would work on design docs like these just fine; all one would need is a 'routine' (weekly? monthly? when releases are cut?) to regularly scan the repo with that command