Repository navigation
Skip backtrace capture in trap handlers to eliminate lock contention - #11
Merged
Merged
Conversation
ma2bd
approved these changes
Mar 17, 2026
|
I'm confused — does that mean we have some code path that is trapping in a loop? How could we be spending 67% of CPU time in any error-handling path? (Nit: we execute Wasm in threads, not tasks — Wasm execution is currently fully synchronous, so blocks the thread) |
ma2bd
pushed a commit
to linera-io/linera-protocol
that referenced
this pull request
Mar 17, 2026
## Motivation Under high concurrency, wasmer's `Backtrace::new_unresolved()` in trap handlers takes a process-wide mutex, causing severe lock contention. In production we observed 67% of CPU time spent in `native_queued_spin_lock_slowpath`. ## Proposal Bump `linera-wasmer` and `linera-wasmer-compiler-singlepass` from `4.4.0-linera.7` to `4.4.0-linera.8`, which eliminates all `Backtrace::new_unresolved()` calls in trap handling paths. See linera-io/wasmer#11 for the upstream changes. ## Test Plan CI
Author
|
Yes, I was confused by that too, but the CPU profile seemed to point to the CPU usage coming from the WASM panics. That turned out to be wrong 😅 Even though it could, it wasn't in this particular situation. |
github-merge-queue Bot
pushed a commit
to linera-io/linera-protocol
that referenced
this pull request
Apr 1, 2026
…re (#5721) (#5844) ## Motivation Under high concurrency, wasmer's Backtrace::new_unresolved() in trap handlers takes a process-wide mutex, causing severe lock contention. In production we observed 67% of CPU time spent in native_queued_spin_lock_slowpath. ## Proposal Bump linera-wasmer and linera-wasmer-compiler-singlepass from 4.4.0-linera.7 to 4.4.0-linera.8, which eliminates all Backtrace::new_unresolved() calls in trap handling paths. See linera-io/wasmer#11 for the upstream changes. Frontport of #5721. ## Test Plan CI
github-merge-queue Bot
pushed a commit
to linera-io/linera-protocol
that referenced
this pull request
Apr 1, 2026
…re (#5721) (#5844) ## Motivation Under high concurrency, wasmer's Backtrace::new_unresolved() in trap handlers takes a process-wide mutex, causing severe lock contention. In production we observed 67% of CPU time spent in native_queued_spin_lock_slowpath. ## Proposal Bump linera-wasmer and linera-wasmer-compiler-singlepass from 4.4.0-linera.7 to 4.4.0-linera.8, which eliminates all Backtrace::new_unresolved() calls in trap handling paths. See linera-io/wasmer#11 for the upstream changes. Frontport of #5721. ## Test Plan CI
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Eliminate all
Backtrace::new_unresolved()calls in trap handling paths.Backtrace::new_unresolved()calls_Unwind_Backtrace, which invokes_Unwind_Find_FDEfor each stack frame. This takes a process-wide mutex (eitherobject_mutexin libgcc ordl_iterate_phdr's internal lock in glibc). Under highconcurrency — hundreds of async tasks executing WASM contracts simultaneously — this
causes severe lock contention. In production we observed 67% of CPU time spent in
native_queued_spin_lock_slowpath(the kernel futex spinlock beneath the contendedmutex), pinning workers at 100% CPU utilization.
Changes
lib/vm/src/trap/traphandlers.rs: Always useBacktrace::from(vec![])in thesignal-based trap handler instead of conditionally calling
Backtrace::new_unresolved()(previously only skipped for stack overflow traps).
lib/vm/src/trap/trap.rs:Trap::lib()andTrap::oom()constructors use emptybacktraces.
lib/compiler/src/engine/trap/stack.rs:Trap::Uservariant returns an emptyWASM trace instead of capturing a new backtrace during
Trap→RuntimeErrorconversion.
Trade-off
The WASM-level stack trace (
FrameInfoframes shown inRuntimeErrorDisplay) is lost— error messages will show
RuntimeError: unreachablewithout theat function_name (module[N]:0xOFFSET)lines. The trap code and the original panic/error message arefully preserved.