Objects treated as missing despite being present, due to race with geometric repacking - #2207
Open
newren wants to merge 2 commits into
Open
Objects treated as missing despite being present, due to race with geometric repacking#2207newren wants to merge 2 commits into
newren wants to merge 2 commits into
Conversation
When objects involved in the merge cannot be read, the merge machinery will return early with result.clean = -1, and result.tree left as NULL. pick_regular_commit() tested only "if (!result->clean)", ignoring the case where "clean < 0". That causes the code to try to use result->tree, resulting in a SIGSEGV. Handle clean < 0 explicitly; the merge machinery will already have printed messages such as "Could not read <object>" and "collecting merge info failed for trees...", so we don't need to add much detail beyond the fact that the merge failed. Signed-off-by: Elijah Newren <newren@gmail.com>
When a geometric repack runs concurrently with other git processes, it can write a new pack and multi-pack-index and then delete older packs that the new one subsumes. One or more of those older packs may have been indexed by the previous multi-pack-index. A process that already had the previous multi-pack-index open keeps using it, and that stale index still records the removed pack(s) as owning some objects. Because a multi-pack-index attributes each object to exactly one pack, an object that exists in multiple covered packs is served only through its recorded owner. If that owner is the pack a concurrent repack just removed, find_pack_entry() cannot serve the object: fill_midx_entry() routes the lookup to the missing pack (prepare_midx_pack() fails), and the regular pack fallback deliberately skips every multi-pack-index covered pack. The object is reported missing even though a perfectly good copy survives in another covered pack -- for example a large "base" pack that geometric repacking intentionally kept. The false negative is not limited to one caller. Any reader (cat-file, rev-list, pack-objects, ...) can spuriously fail with "unable to read object", and callers that only ask whether an object exists get a wrong answer too, since the OBJECT_INFO_QUICK path never retries. Writers that merge in-core, such as "git replay", are hit hardest: merge-ort treats the unreadable tree as a premature abort, sets result.clean < 0, and returns without a result tree. Teach find_pack_entry() to recover. After the normal multi-pack-index lookup and the regular pack fallback both miss, check whether the object is nonetheless present in a covered multi-pack-index (bsearch_midx()). If it is, its recorded owner must have become unavailable, so scan that index's packs directly for a surviving copy. The bsearch gate keeps genuine misses (i.e. objects absent from the index) on the fast path, and because the recovery lives in find_pack_entry() itself it also fixes the OBJECT_INFO_QUICK callers that never reprepare. This recovers the object without touching the multi-pack-index itself. Reloading the stale index would be a more complete fix but would be much more involved: other code (pack bitmaps, object name disambiguation) borrows and caches the "struct multi_pack_index *" across object reads, so freeing it underneath them would be a use-after-free. Refreshing the index with proper invalidation of those borrowers is left for future work. Signed-off-by: Elijah Newren <newren@gmail.com>
Author
|
/submit |
|
Submitted as pull.2207.git.1787092446.gitgitgadget@gmail.com To fetch this version into To fetch this version to local tag |
| @@ -327,6 +327,13 @@ static struct commit *pick_regular_commit(struct repository *repo, | |||
| merge_opt->ancestor = NULL; | |||
There was a problem hiding this comment.
Junio C Hamano wrote on the Git mailing list (how to reply to this email):
"Elijah Newren via GitGitGadget" <gitgitgadget@gmail.com> writes:
> From: Elijah Newren <newren@gmail.com>
>
> When objects involved in the merge cannot be read, the merge machinery
> will return early with result.clean = -1, and result.tree left as NULL.
> pick_regular_commit() tested only "if (!result->clean)", ignoring the
> case where "clean < 0". That causes the code to try to use
> result->tree, resulting in a SIGSEGV.
>
> Handle clean < 0 explicitly; the merge machinery will already have printed
> messages such as "Could not read <object>" and "collecting merge info
> failed for trees...", so we don't need to add much detail beyond the
> fact that the merge failed.
>
> Signed-off-by: Elijah Newren <newren@gmail.com>
> ---
> replay.c | 7 +++++++
> t/t3650-replay-basics.sh | 35 +++++++++++++++++++++++++++++++++++
> 2 files changed, 42 insertions(+)
>
> diff --git a/replay.c b/replay.c
> index 463c900d6c..33e21b2032 100644
> --- a/replay.c
> +++ b/replay.c
> @@ -327,6 +327,13 @@ static struct commit *pick_regular_commit(struct repository *repo,
> merge_opt->ancestor = NULL;
> merge_opt->branch2 = NULL;
>
> + if (result->clean < 0) {
> + error(_("merge of %s onto %s failed"),
> + oid_to_hex(&pickme->object.oid),
> + oid_to_hex(&replayed_base->object.oid));
> + return NULL;
> + }
> +
> if (!result->clean)
> return NULL;
Hmph, so anything but "0 < result->clean" is a failure, but we by
mistake took any non-zero value as OK? That is an obvious mistake.
Well spotted and fixed.
> + # Ensure replay gracefully handles the missing object
> + test_must_fail git replay --onto onto base..side 2>err &&
> + test_grep ! "[Ss]egmentation" err &&
> + test_grep "Could not read\|collecting merge info failed" err
"test_must_fail" means "the tested command must fail voluntarily and
in a controlled way", so a segfaulting git-replay invocation would
not pass test_must_fail. Hence, there is no need to separately
test "test_grep ! '[sS]egmentation'".
Besides, the spelling used by strsignal() is implementation-defined,
so you cannot reliably grep for it anyway.
> + )
> +'
> +
> test_done
Thanks.| @@ -31,6 +31,35 @@ static int find_pack_entry(struct odb_source_packed *store, | |||
| } | |||
There was a problem hiding this comment.
Junio C Hamano wrote on the Git mailing list (how to reply to this email):
"Elijah Newren via GitGitGadget" <gitgitgadget@gmail.com> writes:
> @@ -31,6 +31,35 @@ static int find_pack_entry(struct odb_source_packed *store,
> }
> }
>
> + /*
> + * Recovery for a concurrent-repack race: a MIDX can name an owning
> + * pack for an object that a simultaneous repack has since deleted,
> + * even though the object still exists in another pack the same MIDX
> + * covers (e.g. a kept base pack that geometric repack did not rewrite).
> + * If the object is present in a MIDX yet none of the paths above could
> + * serve it, its recorded owning pack has become unavailable. The
> + * regular fallback above deliberately skips MIDX-covered packs, so
> + * scan this MIDX's packs directly to find the surviving copy. The
> + * bsearch gate keeps genuine misses (objects absent from the MIDX) on
> + * the fast path.
> + */
> + if (store->midx) {
> + struct multi_pack_index *m = store->midx;
> + uint32_t midx_pos, i;
> +
> + if (bsearch_midx(oid, m, &midx_pos)) {
> + for (i = 0; i < m->num_packs + m->num_packs_in_base; i++) {
> + struct packed_git *p;
> +
> + if (prepare_midx_pack(m, i))
> + continue;
> + p = nth_midxed_pack(m, i);
> + if (p && packfile_fill_entry(p, oid, e))
> + return 1;
> + }
> + }
> + }
> +
> return 0;
> }
I'll prepare an evil-merge to rewrite this line to
if (p && packfile_fill_entry(p, oid, e, bad_pack))
to adjust to the API change another topic in-flight brings in when
merging these patches to 'seen'.
This is strictly FYI. You do not need to rebase on top of the other
topic, until I and/or the author of the other topic ask you.
Thanks.
| @@ -31,6 +31,35 @@ static int find_pack_entry(struct odb_source_packed *store, | |||
| } | |||
There was a problem hiding this comment.
Patrick Steinhardt wrote on the Git mailing list (how to reply to this email):
On Tue, Aug 18, 2026 at 10:34:06PM +0000, Elijah Newren via GitGitGadget wrote:
> From: Elijah Newren <newren@gmail.com>
>
> When a geometric repack runs concurrently with other git processes, it
> can write a new pack and multi-pack-index and then delete older packs
> that the new one subsumes. One or more of those older packs may have
> been indexed by the previous multi-pack-index. A process that already
> had the previous multi-pack-index open keeps using it, and that stale
> index still records the removed pack(s) as owning some objects.
>
> Because a multi-pack-index attributes each object to exactly one pack,
> an object that exists in multiple covered packs is served only through
> its recorded owner. If that owner is the pack a concurrent repack just
> removed, find_pack_entry() cannot serve the object: fill_midx_entry()
> routes the lookup to the missing pack (prepare_midx_pack() fails), and
> the regular pack fallback deliberately skips every multi-pack-index
> covered pack. The object is reported missing even though a perfectly
> good copy survives in another covered pack -- for example a large "base"
> pack that geometric repacking intentionally kept.
Okay. Rephrasing in my own words: the object in question exists in two
packs covered by the MIDX. We rewrite one of those two packs, and the
MIDX used to reference the object via the pack we're about to rewrite.
Consequently, the MIDX is stale now and it cannot be used to find the
object anymore because its pack has disappeared. And as we know to skip
searching packfiles for the object that are already covered by the MIDX
we won't be able to find it via the second packfile, either.
> The false negative is not limited to one caller. Any reader
> (cat-file, rev-list, pack-objects, ...) can spuriously fail with
> "unable to read object", and callers that only ask whether an object
> exists get a wrong answer too, since the OBJECT_INFO_QUICK path never
> retries. Writers that merge in-core, such as "git replay", are hit
> hardest: merge-ort treats the unreadable tree as a premature abort, sets
> result.clean < 0, and returns without a result tree.
Hm. Isn't there a slight variant of the race though for any caller that
does not use OBJECT_INFO_QUICK?
Namely, the packfile containing our object disappears and is being
written to a new packfile, and that file is the only one containing it.
Without OBJECT_INFO_QUICK we would be fine: we notice the object could
not be found, and then we perform a second read that makes the "packed"
backend reload its packfiles. It would find the new packfile, and
because it's not covered by its MIDX it would use it to surface the
object. But without OBJECT_INFO_QUICK that's not the case, as we would
skip reloading packfiles altogether, and hence we would not be able to
find that object at all.
As far as I can see though, we don't seem to pass OBJECT_INFO_QUICK in
any of the mentioned readers. I could very well be missing something
here, but I would have thought that those readers are fine in this
scenario?
> diff --git a/odb/source-packed.c b/odb/source-packed.c
> index 0890704e76..de96215069 100644
> --- a/odb/source-packed.c
> +++ b/odb/source-packed.c
> @@ -31,6 +31,35 @@ static int find_pack_entry(struct odb_source_packed *store,
> }
> }
>
> + /*
> + * Recovery for a concurrent-repack race: a MIDX can name an owning
> + * pack for an object that a simultaneous repack has since deleted,
> + * even though the object still exists in another pack the same MIDX
> + * covers (e.g. a kept base pack that geometric repack did not rewrite).
> + * If the object is present in a MIDX yet none of the paths above could
> + * serve it, its recorded owning pack has become unavailable. The
> + * regular fallback above deliberately skips MIDX-covered packs, so
> + * scan this MIDX's packs directly to find the surviving copy. The
> + * bsearch gate keeps genuine misses (objects absent from the MIDX) on
> + * the fast path.
> + */
> + if (store->midx) {
> + struct multi_pack_index *m = store->midx;
> + uint32_t midx_pos, i;
> +
> + if (bsearch_midx(oid, m, &midx_pos)) {
Okay. I was initially worried that we now unconditionally search through
all packfiles a second time, as that could have an impact on
performance. But we really only do this in case we have a MIDX and we
know that the MIDX _should_ have contained the object, but didn't yield
it.
> + for (i = 0; i < m->num_packs + m->num_packs_in_base; i++) {
> + struct packed_git *p;
> +
> + if (prepare_midx_pack(m, i))
> + continue;
> + p = nth_midxed_pack(m, i);
> + if (p && packfile_fill_entry(p, oid, e))
> + return 1;
> + }
And here we now loop through all packs covered by the MIDX and manually
try to look up the object in those. Makes sense.
> + }
> + }
I was wondering whether a preferable fix would be to eagerly load
any packfile referenced by the MIDX when loading the MIDX itself. And if
that fails, we'd ignore the MIDX altogether. This would guarantee that
the MIDX remains valid, and we wouldn't have to worry about any
disappearing packfiles.
The downside is of course that we now eagerly open packfiles, and we
didn't have to do that before. So I think your fix is preferable, as we
can rather easily detect the case where the MIDX should've yielded the
object but didn't, and consequently the additional search only triggers
in very specific edge cases.
Overall I think this patch looks good to me. The one thing that I'm a
bit puzzled about is the above discussion around OBJECT_INFO_QUICK. I
feel like I'm missing something there.
Thanks!
Patrick|
User |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
When an object is found in multiple packs that are in a multi-pack-index, and a subsequent geometric repacking creates a new multi-pack-index and removes the pack that was considered the owner of the object in the old multi-pack-index, then an already-running process that had opened the old multi-pack-index and hadn't yet opened the removed packfile will not be able to access the object -- lookups will return it as missing. Additionally, replay has a separate bug where a missing object causes a SIGSEGV rather than an error message.
This appears to affect a very small percentage of git operations in production since it is a tiny window, but I've found evidence of it occurring in at least eight distinct server-side operations, covering seven different git commands:
There are also commands that could be changing behavior without throwing an error -- e.g. object negotiation thinking an object doesn't exist and instead negotiating based on an older common commit, or cat-file --batch reporting that some objects don't exist.
This series fixes the replay bug first, since it's simpler; investigating it, together with my other recent repacking work, is what led me to the underlying multi-pack-index issue that 2/2 addresses.
cc: Patrick Steinhardt ps@pks.im