You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Opening an issue to track potential issues in kernels found while verifying kernels in Lean.
kernel_gdn_chunk_cumsum.cpp
no V → MTE2 edge. The loop body carries MTE2→V, V→MTE3 and MTE3→V, but nothing
orders Vec against the next iteration's TLOAD. Iteration k+1 loads into the g arena while
iteration k is still reading it with TMOV(acc_ub, g_row_0) and TADD(acc_ub, acc_ub,
g_row_i). Four unordered pairs are exhibited. The overlap is real at row granularity, not
an artifact of coarse footprints: chunk 0 reads rows 0 to 3 and chunk 1's load writes
rows 0 and 1.
TFILLPAD_INPLACE sits before the handshake it needs. It is a Vec instruction on
the g arena issued before set_flag(PIPE_MTE2, PIPE_V), so it is unordered against the
load it is meant to pad. Reachable only on a ragged final chunk or a head count that is
not a multiple of eight, neither of which the test suite exercises. This one is
granularity-dependent and I state it as such: if the real instruction writes only the
padding region it is harmless, and settling that needs the ISA's footprint. What is
certain is that it is on the wrong side of the flag.
Opening an issue to track potential issues in kernels found while verifying kernels in Lean.
kernel_gdn_chunk_cumsum.cppno V → MTE2 edge. The loop body carries MTE2→V, V→MTE3 and MTE3→V, but nothing
orders Vec against the next iteration's TLOAD. Iteration k+1 loads into the g arena while
iteration k is still reading it with TMOV(acc_ub, g_row_0) and TADD(acc_ub, acc_ub,
g_row_i). Four unordered pairs are exhibited. The overlap is real at row granularity, not
an artifact of coarse footprints: chunk 0 reads rows 0 to 3 and chunk 1's load writes
rows 0 and 1.
TFILLPAD_INPLACE sits before the handshake it needs. It is a Vec instruction on
the g arena issued before set_flag(PIPE_MTE2, PIPE_V), so it is unordered against the
load it is meant to pad. Reachable only on a ragged final chunk or a head count that is
not a multiple of eight, neither of which the test suite exercises. This one is
granularity-dependent and I state it as such: if the real instruction writes only the
padding region it is harmless, and settling that needs the ISA's footprint. What is
certain is that it is on the wrong side of the flag.