Skip to content

QuantizedCache(backend="hqq") crashes in _remove_answer_from_cache with TypeError #289

Description

@bonginn

Bug

KVPressTextGenerationPipeline's removeanswer_from_cache() crashes when used with QuantizedCache(backend="hqq", ...). It runs after every generated answer to trim the question+answer tokens back out of the cache, and unconditionally does:

cache.layers[layer_idx]._quantized_keys = cache.layers[layer_idx]._quantized_keys[:, :, :sequence_length]

This assumes ._quantized_keys is a single sliceable tensor, which holds for the quanto backend (WeightQBitsTensor, still shaped [batch, kv_heads, seq_len, head_dim]), but not for hqq: HQQQuantizedLayer._quantize() (transformers/cache_utils.py) returns a (qtensor, meta) tuple, and slicing a plain Python tuple with a 3-axis index raises:

TypeError: tuple indices must be integers or slices, not tuple

qtensor itself has also been reshaped/bit-packed into a flat [bytes_per_group, num_groups] layout (e.g. [32, 176] for a 22-token, 8-head, 128-dim key tensor at 4 bits) where seq_len is no longer a distinct axis, so even a tuple-aware slice couldn't correspond to "the first N tokens" - the fix needs to dequantize() back to the real shape first, slice there, then quantize() again.

To Reproduce

import torch
from transformers import QuantizedCache, pipeline
import kvpress

pipe = pipeline("kv-press-text-generation", model="meta-llama/Llama-3.2-3B-Instruct", device="cuda", dtype=torch.float16)
cache = QuantizedCache(backend="hqq", config=pipe.model.config, nbits=4)
pipe("The Eiffel Tower is located in Paris, France.", question="\nWhere is the Eiffel Tower located?", cache=cache, max_new_tokens=20)

Repository version

7331c23

Activity

  1. Momoyeyu commented on Sep 26, 2026

    @Momoyeyu

    Taking this. The fix: handle hqq's per-layer quantized storage (tuple, not a single sliceable tensor) in _remove_answer_from_cache. Will add a regression test. 🤖🤖🤖

  2. bonginn commented on Sep 26, 2026

    @bonginn
    ContributorAuthor

    Taking this. The fix: handle hqq's per-layer quantized storage (tuple, not a single sliceable tensor) in _remove_answer_from_cache. Will add a regression test. 🤖🤖🤖

    Thanks! This is already addressed in #290. It handles the hqq (qtensor, meta) tuple in _remove_answer_from_cache (dequantize → slice → re-quantize) and adds hqq to test_pipeline_with_quantized_cache as a regression test.

  3. Momoyeyu commented on Sep 26, 2026

    @Momoyeyu

    Understood — I missed #290, which was already open before my comment. Standing down; the dequantize→slice→requantize approach there matches what I had in mind anyway. 🤖🤖🤖

  4. Momoyeyu commented on Sep 26, 2026

    @Momoyeyu

    Verified locally: the hqq-parametrized test in #290 fails on main (TypeError: tuple indices) and passes on the PR branch with real hqq (0.2.8) + danube3-500m. #289 is genuinely covered. 🤖🤖🤖

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions