You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
In real-time robotics and navigation loops running at 500–1000 Hz, state and covariance matrix multiplications (typically 4x4 to 16x16) suffer from NumPy/BLAS fixed dispatch latency (3–5 µs per @ call).
We created a minimalist C microkernel called NanoGEMM designed specifically for sub-microsecond execution on small matrices via direct AVX2+FMA and ARM NEON register tiling.
Micro-benchmark on 8x8 & 16x16 covariance update:
NumPy @: ~3.2 µs per iteration
NanoGEMM: ~1.1 µs per iteration (2.9x faster)
It can be integrated as an optional drop-in acceleration with a zero-friction fallback:
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Hi @rlabbe and contributors,
Thank you for maintaining
filterpy!In real-time robotics and navigation loops running at 500–1000 Hz, state and covariance matrix multiplications (typically 4x4 to 16x16) suffer from NumPy/BLAS fixed dispatch latency (3–5 µs per
@call).We created a minimalist C microkernel called NanoGEMM designed specifically for sub-microsecond execution on small matrices via direct AVX2+FMA and ARM NEON register tiling.
Micro-benchmark on 8x8 & 16x16 covariance update:
@: ~3.2 µs per iterationIt can be integrated as an optional drop-in acceleration with a zero-friction fallback:
This would give robotics and sensor fusion pipelines a 2-3x speedup on CPU without adding any mandatory dependencies.
Would you be open to a lightweight PR introducing an optional fast path?
All reactions