Skip to content

Perf: cache hash and prefetch chain in TensorMap lookup/insert - #5

Open
chenshengxin2026 wants to merge 3 commits into
mainfrom
perf/orch-lookup-prefetch-hash-cache
Open

chenshengxin2026 wants to merge 3 commits into
mainfrom
perf/orch-lookup-prefetch-hash-cache

Conversation

@chenshengxin2026

Copy link
Copy Markdown
Owner

Summary

  • Cache the hash(addr) result from lookup() and reuse it in the subsequent insert() call for INOUT tensors, eliminating a redundant 64-bit multiply per tensor
  • Add software prefetch of next_in_bucket during chain traversal to hide memory latency on chains longer than one entry
  • Add lookup/insert/link_entry overloads that accept precomputed hash

Benchmark Results

Ascend910 (device 11, 100 rounds, 3 runs averaged):

Example Baseline (us) Optimized (us) Delta
alternating_matmul_add 977.9 971.4 -0.7%
benchmark_bgemm 747.6 719.2 -3.8%
paged_attention_unroll Case1 1165.0 1158.2 -0.6%
paged_attention_unroll Case2 555.2 554.0 -0.2%
batch_paged_attention 3259.2 3239.0 -0.6%

The bgemm improvement is expected: it has the highest lookup+dep percentage (45.8% of orch time) and uses INOUT tensors extensively.

Testing

  • Simulation tests pass (20/20 a2a3sim)
  • Hardware tests pass

hw-native-sys-bot and others added 2 commits April 7, 2026 15:11
…-sys#456)

The pin-commit retry was inside device_worker_main (subprocess), so a
segfault in ChipWorker killed the subprocess before the retry could run.

- Split run_hw_tasks_subprocess into sim (no retry) and hw (3 retries)
- Move PTO-ISA pin retry to main() where the parent process controls it
- Subprocess runs once and exits; all retry decisions made by parent
- Remove --max-attempts CLI arg (subprocess always single-shot)

Co-authored-by: Chao Wang <26245345+ChaoWao@users.noreply.github.com>
…e-sys#462)

- Quarantine device on first failure; re-enqueue task for healthy devices (up to MAX_RETRIES across devices), matching ci.sh semantics
- Print subprocess logs on failure: sim during pin-commit retry, hw on last retry during pin-commit
- Add progress logging ([n/total] PASS/FAIL) for both sim and hw paths
- Remove unused `--parallel` flag (parallelism determined by `-d` device count)
- Remove unused `PYTHONDONTWRITEBYTECODE` env setting

Co-authored-by: Chao Wang <26245345+ChaoWao@users.noreply.github.com>
@chenshengxin2026
chenshengxin2026 force-pushed the perf/orch-lookup-prefetch-hash-cache branch from 7bbd763 to c6601f7 Compare April 7, 2026 09:23
- Cache the hash(addr) result from lookup() and reuse it in the
  subsequent insert() call for INOUT tensors, eliminating a redundant
  64-bit multiply per tensor
- Add software prefetch of next_in_bucket during chain traversal to
  hide memory latency on chains longer than one entry
- Add lookup/insert/link_entry overloads that accept precomputed hash

Benchmarked on Ascend910 (device 11, 100 rounds, 3 runs averaged):
benchmark_bgemm -3.8%, other workloads -0.2% to -0.7%.
@chenshengxin2026
chenshengxin2026 force-pushed the perf/orch-lookup-prefetch-hash-cache branch from b688858 to a0eff4b Compare April 7, 2026 11:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants