Skip to content

Batch Git base-text reads across files when loading multibuffer diffs #153

Description

@kjanat

Problem

The initial load of a large multibuffer diff is reported as slow. One verified source of overhead in the uncommitted-changes path is serial Git base-text loading with a fresh process for each file.

Source evidence

Investigated at 5ee612d; these loading paths are unchanged by PR #152.

  • load_committed_text requests HEAD and index contents together for one file through send_job.
  • The local repository worker awaits each queued job before starting the next.
  • load_revisions starts a new git cat-file --batch process for each invocation. Batching currently combines revisions within a file, not across files.

Read-only measurement

Used the first 60 added/modified paths from git diff --name-only --diff-filter=AM bffee761017f23531a264a1ae033b07e18c9e524 5ee612d5ca701d8c41c62f83a3bfff535c34421b, requesting HEAD:path and :path at local head ffe31ca.

Compared 60 sequential git cat-file --batch invocations with one invocation receiving all 120 revision requests. Three paired runs returned byte-identical output (19,357,226 bytes, checked by SHA-256):

Mode Run 1 Run 2 Run 3
One process per file 284.5 ms 277.9 ms 277.0 ms
One batch for all files 62.6 ms 56.6 ms 51.1 ms

The median ratio is approximately 4.9x for this isolated Git-read operation. This is not an end-to-end Zed benchmark, does not isolate cold-cache performance, and does not establish that Git reads dominate the reported delay.

Proposed change

Coalesce compatible base-text requests across files, or reuse a batch reader, without weakening ordering around repository mutations. Prefer bounded batches so the first useful file can appear promptly rather than waiting for the whole repository.

Acceptance criteria

  • Equivalent HEAD/index contents and errors, including missing revisions, additions, deletions, and partially staged files.
  • Preserve correctness when repository state changes during loading.
  • Measure process count, time to first displayed file, and total load time on many-file diffs, with cold and warm cases recorded separately.
  • Demonstrate the benefit in the actual viewer; do not substitute the isolated benchmark for that validation.

Upstream research (2026-09-29)

  • PR #59357 was merged on 2026-07-07 and is already an ancestor of this fork's investigated master. It batches repository-state reloads across open buffers and combines HEAD/index requests for individual initial loads. This issue targets the remaining initial-load batching across files, not reimplementing that merged work.
  • Issue #41054 reports multibuffer diff opening taking tens of seconds over SSH despite a small diff. It was closed for inactivity, not with a verified fix. Remote latency differs from the local measurement above.
  • Open PR #61132 documents the same serial repository-queue contention in the worktree picker. Its proposed bypass applies to that picker, not directly to mutable HEAD/index reads; preserve ordering here.
  • Open PR #63456 addresses large-repository status refresh and history loading. Those can compete for resources but are separate from initial diff base-text batching.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions