Tracking issue for the nightly test-annex (nfs-home, ubuntu-24.04) failure, root-caused with the instrumented harness in #293. Nothing here needs a patch in this repository; it is a git-annex upstream matter plus a workaround for NFS users.
Symptom
test-annex (nfs-home) fails on roughly two nightlies out of three, in a different test each night, and is usually the only failing job in the run. The messages alternate between:
<file> failed to link to annex — git annex add of an unlocked file
git-annex: unlock failed — git annex unlock
Exception: content changed while it was being sent — testremote transfers
Root cause
All three are the same assumption. git-annex records a file's InodeCache — (inode, size, high-resolution mtime) — copies the file, stats it again, and compares the two with compareStrong, i.e. exact equality including the nanosecond field. If they differ it concludes the file changed underneath, deletes the destination and fails.
On NFS a client can return two different mtimes for a file nothing wrote to. Instrumented runs caught it four times, always with inode and size unchanged and only the mtime moving:
| file |
mtime before → after |
delta |
worktree file (import) |
1790084329.078602968 → .090671216 |
+12.07 ms |
worktree file (add subdirs) |
1790178434.030902619 → .309979362 |
+279 ms |
annex object (edit) |
1790084752.023778495 → 1790084806.769318488 |
+54.75 s |
annex object (conversion annexed to git) |
1790178735.562779131 → 1790178776.778504714 |
+41.22 s |
The sub-second ones are the client's value being replaced by the server's within the same second; the tens-of-seconds ones are a cached attribute going stale (default acregmax is 60 s). In the object cases nothing in the test writes to the file at all.
Two call sites are affected, so this is not confined to one command:
Annex/Content.hs:646 — checksrcunchanged in linkAnnex → failed to link to annex / unlock failed
Annex/Content.hs:730 — sameInodeCache in prepSendAnnex → content changed while it was being sent
compareInodeCaches only relaxes to compareWeak (2-second tolerance) when the inode sentinal says inodes changed, which never happens here, so the comparison stays exact.
Confirmed by experiment
Run 35882579073 ran the nfs-home flavor three ways, three times each, same build and same runner batch:
| arm |
mount |
result |
default |
as the nightlies mount |
3/3 failed |
actimeo0 |
-o actimeo=0 |
3/3 passed |
noac |
-o noac |
3/3 passed |
actimeo=0 disables attribute caching and changes nothing else, so this is the attribute cache, not the write path. Cost is roughly 4×: a passing default job did the same 26 groups in 8m20s against 32m01s with actimeo=0.
Workaround for NFS users: mount with actimeo=0. noac also works but additionally forces synchronous writes and is not needed.
Is it a regression?
Not in the git-annex code that fails — in the CI environment or in timing.
test-annex (nfs-home) passed on 2026-03-13 (run 23036066273, where only test-datalad failed) and was failing by 2026-08-08 (run 31238024732). Most June–July nightlies failed at build-package, so the suite did not run and the transition is masked in the GitHub-side logs; the tinuous archive has those runs if anyone wants to bisect the date exactly.
- On the git-annex side nothing moved: since 2026-01-01 exactly one commit touched
Annex/Content*.hs, Utility/InodeCache.hs, Annex/InodeSentinal.hs or Utility/CopyFile.hs — 95520e391 "generalize type", which only changes downloadUrl. No commit has touched the compareStrong usage since at least 2025-01-01, and both the strict comparison and the high-resolution mtime predate 2022-08 (as far back as this repo's upstream/master mirror can be deepened).
So the code has assumed this for years; what changed is how often the NFS client actually violates it on the runners.
Reproducing
Two scripts, proposed for eval-under in con/eval-under#11:
mtime-stability — write, stat, copy as git-annex copies, stat again, compare; reports a rate, no git-annex involved. This is the minimal reproducer: the failure is visible with four lines of Python.
git-annex-linkannex — loops unlock and unlocked add, reporting a failure rate in minutes rather than a pass/fail of the whole suite.
That PR also adds eval-under nfs --mount-opts, without which actimeo=0 could not be tested locally at all.
Suggested next step
Report upstream at git-annex.branchable.com: exact equality on a high-resolution mtime is not a safe assumption on NFS, at both checksrcunchanged and sameInodeCache, with the numbers above and the actimeo=0 control. A fix confined to linkAnnex would leave transfers failing.
Tracking issue for the nightly
test-annex (nfs-home, ubuntu-24.04)failure, root-caused with the instrumented harness in #293. Nothing here needs a patch in this repository; it is a git-annex upstream matter plus a workaround for NFS users.Symptom
test-annex (nfs-home)fails on roughly two nightlies out of three, in a different test each night, and is usually the only failing job in the run. The messages alternate between:<file> failed to link to annex—git annex addof an unlocked filegit-annex: unlock failed—git annex unlockException: content changed while it was being sent—testremotetransfersRoot cause
All three are the same assumption. git-annex records a file's
InodeCache—(inode, size, high-resolution mtime)— copies the file, stats it again, and compares the two withcompareStrong, i.e. exact equality including the nanosecond field. If they differ it concludes the file changed underneath, deletes the destination and fails.On NFS a client can return two different mtimes for a file nothing wrote to. Instrumented runs caught it four times, always with inode and size unchanged and only the mtime moving:
import)1790084329.078602968→.090671216add subdirs)1790178434.030902619→.309979362edit)1790084752.023778495→1790084806.769318488conversion annexed to git)1790178735.562779131→1790178776.778504714The sub-second ones are the client's value being replaced by the server's within the same second; the tens-of-seconds ones are a cached attribute going stale (default
acregmaxis 60 s). In the object cases nothing in the test writes to the file at all.Two call sites are affected, so this is not confined to one command:
Annex/Content.hs:646—checksrcunchangedinlinkAnnex→failed to link to annex/unlock failedAnnex/Content.hs:730—sameInodeCacheinprepSendAnnex→content changed while it was being sentcompareInodeCachesonly relaxes tocompareWeak(2-second tolerance) when the inode sentinal says inodes changed, which never happens here, so the comparison stays exact.Confirmed by experiment
Run 35882579073 ran the
nfs-homeflavor three ways, three times each, same build and same runner batch:defaultactimeo0-o actimeo=0noac-o noacactimeo=0disables attribute caching and changes nothing else, so this is the attribute cache, not the write path. Cost is roughly 4×: a passingdefaultjob did the same 26 groups in 8m20s against 32m01s withactimeo=0.Workaround for NFS users: mount with
actimeo=0.noacalso works but additionally forces synchronous writes and is not needed.Is it a regression?
Not in the git-annex code that fails — in the CI environment or in timing.
test-annex (nfs-home)passed on 2026-03-13 (run 23036066273, where onlytest-dataladfailed) and was failing by 2026-08-08 (run 31238024732). Most June–July nightlies failed atbuild-package, so the suite did not run and the transition is masked in the GitHub-side logs; the tinuous archive has those runs if anyone wants to bisect the date exactly.Annex/Content*.hs,Utility/InodeCache.hs,Annex/InodeSentinal.hsorUtility/CopyFile.hs—95520e391"generalize type", which only changesdownloadUrl. No commit has touched thecompareStrongusage since at least 2025-01-01, and both the strict comparison and the high-resolution mtime predate 2022-08 (as far back as this repo'supstream/mastermirror can be deepened).So the code has assumed this for years; what changed is how often the NFS client actually violates it on the runners.
Reproducing
Two scripts, proposed for eval-under in con/eval-under#11:
mtime-stability— write, stat, copy as git-annex copies, stat again, compare; reports a rate, no git-annex involved. This is the minimal reproducer: the failure is visible with four lines of Python.git-annex-linkannex— loopsunlockand unlockedadd, reporting a failure rate in minutes rather than a pass/fail of the whole suite.That PR also adds
eval-under nfs --mount-opts, without whichactimeo=0could not be tested locally at all.Suggested next step
Report upstream at git-annex.branchable.com: exact equality on a high-resolution mtime is not a safe assumption on NFS, at both
checksrcunchangedandsameInodeCache, with the numbers above and theactimeo=0control. A fix confined tolinkAnnexwould leave transfers failing.