Skip to content

Unbounded memory growth in libvmaf 3.2.0 when the decoder outpaces feature extraction (regression from 3.1.0) #1587

Description

@csk61-gh

Under ffmpeg's libvmaf filter, libvmaf 3.2.0 accepts pictures faster than it scores them and queues the surplus without bound, growing until the OS kills the process. libvmaf 3.1.0, with the same ffmpeg binary, command, and input, holds flat memory indefinitely.

This only manifests when the decoder can outrun feature extraction — here, two 4K HEVC 10-bit streams via hardware decode (~167 fps each) feeding libvmaf (~21 fps).

Environment:

  • macOS 26.6.2, Apple M1 Max (10 cores), 64 GB RAM
  • ffmpeg 8.1.2 (Homebrew), --enable-libvmaf --enable-videotoolbox --enable-gpl --enable-version3
  • libvmaf 3.2.0 (fails) / 3.1.0 (works) — Homebrew bottles, swapped via DYLD_LIBRARY_PATH
  • Input: two 3840×2160 yuv420p10le HEVC files, 24 fps, 2:36:22, 225,177 frames

Reproduce:

ffmpeg -hwaccel videotoolbox -i reference.mkv -hwaccel videotoolbox -i distorted.mkv \
  -lavfi "[0:v][1:v]libvmaf=n_threads=10:shortest=1" -an -sn -f null -

Monitor with footprint -p <pid> — not ps -o rss, which falls as pages swap out and hides the growth.

Result:
3.2.0 - 26GB at 20s, 131GB at ~8 min, then SIGKILL by jetsam. ~44fps vmaf speed
3.1.0 - flat 2.1GB for 3 minutes, swap unmoved, ~21fps vmaf speed

Growth is ~0.27 GB/s, consistent with (decode rate − scoring rate) × 24.9 MB per 4K 10-bit frame. Failure is deterministic: two independent full runs died at frame 9,916 and frame 10,065.

The reported fps is itself diagnostic. 3.2.0's ~44 fps counts pictures accepted into the queue, not scored. 3.1.0's ~21 fps is the true rate, and matches a full run that completed on 3.1.0 in 3h38m (0.717x).

As a control:
the same framesync path with 'ssim' is fine.

ffmpeg -hwaccel videotoolbox -i reference.mkv -hwaccel videotoolbox -i distorted.mkv \
  -lavfi "[0:v][1:v]ssim" -an -sn -f null -

Flat at 918 MB for 3 minutes at 113 fps — a 54 fps producer/consumer gap absorbed with no growth. So ffmpeg's framesync layer applies backpressure correctly; the difference is libvmaf.

Decode alone is also clean: one input to -f null holds 304 MB at 274 fps; two inputs hold 590 MB at 335 fps combined.

n_threads makes it worse, not better

n_threads=1 reached 72 GB in 20 seconds (vs 26 GB at n_threads=10), which is consistent with n_threads setting the drain rate of an unbounded queue rather than its capacity.

Scores are unchanged

Over 1,440 frames with vmaf_4k_v0.6.1: mean 98.419877 (3.1.0) vs 98.419869 (3.2.0); min 91.551662 vs 91.551591; harmonic mean 98.380873 vs 98.380864. Differences ~1e-5 — this is purely a memory/flow-control regression.

Ruled out

  • PTS mismatch / frame-count disparity. Both inputs start at PTS 0 with identical timebases. Growth is identical when timestamps are regenerated from frame index (settb=AVTB,setpts=N/(24*TB) on both), and occurs at frame ~10,000 of 225,000 — nowhere near EOF.
  • shortest=1 — removing it changes nothing (27 GB in 20 s).
  • -thread_queue_size (8 and 32), -extra_hw_frames (2 and 8) — no effect.
  • format=yuv420p before the filter — halves frame size, growth continues (21 GB in 20 s).
  • Software decode — narrows the gap so growth is slower, but it continues.
  • ffmpeg version — 8.1.1 fails identically when linked against libvmaf 3.2.0, which is how we isolated libvmaf as the variable.

Possibly related in the 3.2.0 release notes

VMAF_BATCH_THREADING / VMAF_PICTURE_POOL threading modes, and USE_DIRECT_READ ("eliminate intermediate buffer and memcpy"). I bisected the version but did not read the diff, so I have not identified the specific change.

vmaf_read_pictures() is documented as queueing pictures and taking ownership, so queueing is by design — the question is whether that call still applies pressure when the queue is saturated.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions