Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
86 changes: 85 additions & 1 deletion BENCHMARK.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# moon Benchmark Report

**Last Updated:** 2026-07-04 (**new §6.6: multi-shard multi-connection WIN — reply-convoy fix + slot-unified replies take s4 c8 P1 from 0.44–0.64× to 1.57–2.0× Redis and c64 to 2.5×, both arches**; supersedes §6.5's c≥8 rows. Prior 2026-07-03: **§2.10: p=1 single-op WIN on both arches via `--io-busy-poll-us` poll-mode park** — ARM c4a 1.19–1.21×, x86 c3 1.65–1.66× vs Redis, n=3 instances/arch; supersedes the "loses p=1" rows in §2.7–2.9 when busy-poll is on. **New §6.5: shards × busy-poll sweep** — busy-poll recovers +22–42% of the cross-shard hop but shards>1 still loses non-pipelined; `--shards 1 --io-busy-poll-us 40` is the p=1 config. Prior: 2026-06-16 4-feature concurrent-vs-competitor pass — §10.5 vector vs RediSearch, §11.4 graph vs FalkorDB, §12 Full-Text Search; §2.8 = v2-1/PR #189 `db61973` K=1024 re-measurement; v0.1.6 in §2.1–2.6; §2.7 on perf/shard-dispatch-hot-path)
**Last Updated:** 2026-07-08 (**new §7.3: durability write-path campaign (PRs #238–#242)** — closes the AOF-on deficits vs Redis. `appendfsync always` P16 SET 0.12×→**0.91×** (per-command fsync → per-batch group commit + one coalesced write), `everysec` P1 SET 0.80×→**0.99× parity** (park-free writer poll kills a 150k/run futex-wake storm on the shard thread), `everysec` P16 SET **1.32× WIN**, and pub/sub fan-out delivery 438 msg/s (near-total drops) → **5.09M msg/s, zero drops, ~1.04× Redis**. All GCE c3-standard-8, Redis 7.0.15, 3 alternated reps, provenance-probed. Prior 2026-07-04 §6.6: multi-shard multi-connection WIN — reply-convoy fix + slot-unified replies take s4 c8 P1 from 0.44–0.64× to 1.57–2.0× Redis and c64 to 2.5×, both arches; supersedes §6.5's c≥8 rows. Prior 2026-07-03: **§2.10: p=1 single-op WIN on both arches via `--io-busy-poll-us` poll-mode park** — ARM c4a 1.19–1.21×, x86 c3 1.65–1.66× vs Redis, n=3 instances/arch; supersedes the "loses p=1" rows in §2.7–2.9 when busy-poll is on. **§6.5: shards × busy-poll sweep** — busy-poll recovers +22–42% of the cross-shard hop but shards>1 still loses non-pipelined; `--shards 1 --io-busy-poll-us 40` is the p=1 config. Prior: 2026-06-16 4-feature concurrent-vs-competitor pass — §10.5 vector vs RediSearch, §11.4 graph vs FalkorDB, §12 Full-Text Search; §2.8 = v2-1/PR #189 `db61973` K=1024 re-measurement; v0.1.6 in §2.1–2.6; §2.7 on perf/shard-dispatch-hot-path)
**Platforms:** Linux (GCloud x86_64 + ARM64), macOS (Apple M4 Pro)
**Redis:** 8.6.1 in §2.1–2.6; 7.0.15 in §2.7
**moon:** v0.1.6 in §2.1–2.6; perf/shard-dispatch-hot-path HEAD (commit `6582fa9`) in §2.7. Monoio runtime (io_uring on Linux, kqueue on macOS), fat LTO, codegen-units=1, target-cpu=native
Expand Down Expand Up @@ -49,6 +49,10 @@ The SET absolute number can differ 3-4× between methodologies. Only strict-vs-s
| Baseline RSS (empty) | **Identical (7.0 MB)** | 1-shard |
| CPU efficiency at P=64 | **45x better** | 1.9% vs 43.9% CPU for similar RPS |
| With AOF persistence | **2.75x Redis** | SET, P=64, per-shard WAL |
| AOF everysec SET P16 | **1.32x Redis** | group commit + coalesced write, §7.3 |
| AOF everysec SET P1 | **0.99x (parity)** | park-free writer poll, §7.3 (was 0.80x) |
| AOF always SET P16 | **0.91x Redis** | per-batch group commit, §7.3 (was 0.12x) |
| Pub/sub fan-out delivery | **5.09M msg/s, 1.04x Redis, 0 drops** | 8 subs, coalesced writes, §7.3 (was 438 msg/s) |
| Multi-shard (8s P=16) | **1.84-1.99x Redis** | GET / SET |
| p50 latency (8-shard) | **8-10x lower** | 0.031ms vs 0.26ms |
| Data correctness | **2613+ tests pass** | All types, 1/4/12 shards |
Expand Down Expand Up @@ -622,6 +626,46 @@ Readings:
| Fsync | Dedicated bio thread | Separate timer, every 1 second |
| Under P=64 | Global AOF becomes serialization point | Per-shard WAL scales linearly |

### 7.3 2026-07-08 Durability write-path campaign (PRs #238–#242)

**Goal:** close every remaining Moon-vs-Redis deficit on the AOF-on write path so Moon is at parity-or-better under all durability policies, not just `--appendonly no`.

**Setup:** GCE `c3-standard-8` (Sapphire Rapids, 8 vCPU, pd-ssd), Ubuntu 24.04, Redis 7.0.15, Moon `--shards 2` (its default multi-core posture on 8 vCPU) vs Redis stock single-threaded. `redis-benchmark -t set -c 8 -n 300000 -r 100000 -d 64 [-P 16]`. Fresh server + data dir per scenario, **3 alternated reps** (engine order flips each rep to cancel warm-up/thermal drift), each run **provenance-probed** (thread-name scan + `strace -c` confirming the code path under test). Numbers are 3-rep means.

#### 7.3.1 Results — the four deficits, before → after

| Policy / workload | Before campaign | After (this branch) | Redis | vs Redis after |
|---|---:|---:|---:|:---:|
| **`always` SET P16** | 5,680 | **39,621 → 40,084** | 43,898 | **0.91×** |
| `always` SET P1 (control) | ~3,180 | 3,114 | 3,153 | parity (fsync-device-bound) |
| **`everysec` SET P16** | 605,317 | **788,717** | 597,563 | **1.32× WIN** |
| **`everysec` SET P1** | 116,619 (0.80×) | **134,158–135,280** | 135,888 | **0.99× parity** |
| `nodur` SET P1 (control) | 135,795 | 135,319 | 132,185 | 1.02× (unchanged) |
| **Pub/sub fan-out delivery** | 438 msg/s (drops ≈100%) | **5.09M msg/s (0 drops)** | 4.89M msg/s | **1.04× WIN** |

`always` P16 moved in two steps: PR #239 (per-command awaited fsync → per-batch group commit) took it 0.12× → 0.85×; PR #242 (coalesced batch write) took it 0.85× → 0.91×.

#### 7.3.2 What each fix changed

| PR | Deficit | Root cause (measured) | Fix |
|----|---------|----------------------|-----|
| **#238** | `always` shard stalls | WAL v3 `fdatasync` ran on the shard event-loop thread — every fsync froze SPSC drain + conn I/O + CDC | Off-loop `WalSyncAgent` (per-shard `std::thread`, fd-dup requests, durable-LSN watermark); +18% RPS, p99/p999 −40–60% under `always` |
| **#239** | `always` P16 **0.12×** | Handler awaited **one fsync ack per pipelined command**; a 16-deep pipeline paid 16 serialized fsync RTTs/conn while Redis fsyncs once per event-loop iter | Local writes join the per-batch group commit (fire-and-forget `send_append_group` + ONE `fsync_barrier` per batch); 0.12× → 0.85× (6.98×) |
| **#240** | Pub/sub delivery **438 msg/s** | One `write(2)` per delivered message; under fan-out flood the 256-slot subscriber queue stayed full → near-total drops | Coalesce the queued burst (`try_recv` drain, 64 KB cap) into ONE `write_all`; 438 → 5.09M msg/s, zero drops |
| **#241** | `everysec` P1 **0.80×** | AOF writer parked in `flume::recv_timeout`, so **every producer `try_send` paid a futex WAKE on the shard thread** — `strace -c`: **149,718 futex calls (63% of shard-thread syscall time)** in an 8 s SET run vs 2 under `nodur` | Under `everysec`/`no` the writer polls park-free (`try_recv` + adaptive sleep, `wait/16` clamped 500 µs–50 ms) so producer sends stay pure userspace atomics; **149,718 → 96 futex calls**, +15% RPS → 0.99× parity. `always` keeps the parked recv (ack latency is client-visible RTT) |
| **#242** | `always` P16 **0.85×** residual | Writer was **write-syscall-bound, not fsync-bound**: `strace -c` on `aof-writer-0` showed **127,812 `write(2)` calls / 1.25 s** vs 2,144 `fdatasync` / 0.20 s (one write per record on the raw unbuffered file; per-shard framed path used a header+body **pair**). Redis batches ~120 records/write via `aof_buf` | Coalesce each group-commit batch into ONE contiguous `write_all` before its single fsync (reusable buffer, 1 MB high-water cap); 0.85× → 0.91×, and `everysec` P16 (shares the write path) → 1.32× |

**Durability invariant held throughout.** Every change preserves fsync-before-ack (H1): the AOF writer channel is ordered, so an acked zero-length `AppendSync` barrier proves every prior `Append` is on disk; a failed batch acks every waiter an error (never a silent `+OK`) and latches the torn stream. Re-verified after each change with `crash_matrix_per_shard_aof --ignored` (SIGKILL under `appendfsync always` recovers 100% of acked writes): **3/3 green**.

#### 7.3.3 Reading the numbers

- **`always` P1 is a hardware floor, not a Moon limit** — ~3.1–3.4k for *all three* engines. Each write is one fsync RTT to pd-ssd; the disk, not the server, sets the rate. Parity here is the best any correct implementation can do.
- **`always` P16 residual 0.91×** — both engines are fsync-device-bound at pipeline depth; the last ~9% is the per-batch barrier ack round-trip, and two shards' fsyncs serialize at the device rather than parallelizing. Diminishing returns.
- **`everysec` P16 1.32× WIN** — no per-batch fsync ceiling here, so the coalesced-write + park-free-poll savings surface directly as throughput.
- **Base column caveat** — the `everysec`/`always` "before" figures come from the pre-campaign tip; the P1 `everysec` base (116,619) predates PR #241, which is why the after-figure jumps despite the write-path work.

Full method notes and per-rep tables: `tmp/MOON-VS-REDIS-DURABILITY.md`, `tmp/WALV3-OFFLOOP-FSYNC.md`.

---

## 8. Production Workload Patterns
Expand Down Expand Up @@ -1009,6 +1053,46 @@ redis-server --port 6399 --save "" --appendonly yes --appendfsync everysec --dae
redis-benchmark -p PORT -c 50 -n 200000 -t SET,INCR,LPUSH,HSET -P 16 -q
```

### Durability write-path A/B (§7.3)

Same-instance, alternated-rep harness that produced the §7.3 matrix. Run inside a
Linux VM (fsync numbers are meaningless on shared-core / non-pinned macOS).

```bash
# Fresh server + dir per scenario, 3 reps, engine order flips each rep.
# For EACH (durability, workload) pair, compare Moon --shards 2 vs Redis stock:
for dur in nodur:'--appendonly no' \
esec:'--appendonly yes --appendfsync everysec' \
always:'--appendonly yes --appendfsync always'; do
moon_args=${dur#*:}
MOON_DISK_FREE_MIN_PCT=0 ./target/release/moon --port 6399 --shards 2 --dir "$(mktemp -d)" $moon_args &
redis-server --port 6399 --dir "$(mktemp -d)" --save '' $moon_args & # matching policy
# p1 (no pipeline) AND p16:
redis-benchmark -p 6399 -t set -c 8 -n 300000 -r 100000 -d 64 # p1
redis-benchmark -p 6399 -t set -c 8 -n 300000 -r 100000 -d 64 -P 16 # p16
done

# Provenance probe — confirm the code path is actually engaged:
# park-free writer (#241): strace -c the shard-N thread; futex calls should be
# ~10s (esec), NOT ~150k. (parked recv = the old path; a regression.)
# batch-coalesced write (#242): strace -c aof-writer-N under always -P 16;
# write(2) count should be << record count (one write per batch, not per record).
sudo strace -c -p "$(ps -T -p <moon_pid> -o tid,comm | awk '/shard-0/{print $1}')"
```

### Pub/sub fan-out delivery A/B (§7.3, PR #240)

```bash
# N subscribers on one channel + one pipelined publisher; metric = messages
# DELIVERED per second across all subscribers (wire-frame count on each socket).
# Drops are allowed by the slow-subscriber policy, so delivered-throughput IS
# the ceiling under test. Harness: scratchpad pubsub_delivery_bench.py.
./target/release/moon --port 6399 --shards 2 --appendonly no &
python3 pubsub_delivery_bench.py 6399 8 5 # 8 subs, 5s
# Coalesced build delivers ~5M msg/s at ratio 1.000; a per-message-write build
# drops ~100% under flood (subscriber queue cap = 256).
```

### GCloud Linux Benchmark

```bash
Expand Down
13 changes: 13 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,19 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

### Docs — durability write-path benchmark results (PR #TBD)

- `BENCHMARK.md` §7.3: new section recording the 2026-07-08 durability
write-path campaign (PRs #238–#242) — the vs-Redis matrix (`always` P16
0.12×→0.91×, `everysec` P16 1.32× win, `everysec` P1 0.80×→0.99× parity,
pub/sub fan-out 438→5.09M msg/s), per-PR root-cause table (strace
diagnostics), durability-invariant note, and A/B reproduction steps.
Executive summary + header updated.
- `docs/production-guide.md`, `docs/benchmarks.md`, `docs/configuration.md`:
AOF `appendfsync` tuning guidance updated with the measured ratios and a
plain-language explanation of group commit / coalesced writes / park-free
writer poll (all automatic; durability unchanged).

### Fixed — Windows build: `accept_backoff` used unix-only `libc` errnos (PR #TBD)

- `is_resource_exhaustion` (accept-loop backoff, PR #230) referenced
Expand Down
14 changes: 14 additions & 0 deletions docs/benchmarks.md
Original file line number Diff line number Diff line change
Expand Up @@ -94,6 +94,20 @@ At pipeline=64, Moon delivers **1.71x the throughput of Redis while using 23x le
!!! note
Moon's per-shard WAL avoids the global serialization point that Redis's single AOF file introduces. The advantage grows with pipeline depth because per-shard WAL scales linearly with shards.

### Max-durability (`appendfsync always`) and `everysec` write path

A 2026-07 write-path campaign (GCE c3-standard-8, Redis 7.0.15, `--shards 2`, 3 alternated reps) closed Moon's remaining AOF-on deficits:

| Policy / workload | Before | After | vs Redis |
|:---|---:|---:|:---:|
| `always` SET p16 | 5.7K (0.12x) | **40.1K** | **0.91x** |
| `always` SET p1 | ~3.2K | 3.1K | parity (fsync-device-bound) |
| `everysec` SET p16 | 605K | **789K** | **1.32x** |
| `everysec` SET p1 | 117K (0.80x) | **135K** | **0.99x parity** |
| Pub/sub fan-out delivery | 438 msg/s (drops) | **5.09M msg/s** | **1.04x, 0 drops** |

The wins come from per-batch group commit (one fsync per pipeline batch, not per command), a single coalesced `write_all` per batch, a park-free AOF-writer poll that removes a ~150K/sec futex-wake storm on the shard thread under `everysec`, and coalesced pub/sub delivery writes. Durability is unchanged (`always` = RPO 0, `everysec` = RPO ≤ 1s), verified by the SIGKILL crash-recovery matrix. Full detail: `BENCHMARK.md` §7.3.

Comment on lines +109 to +110

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Fix the futex-rate wording.

The benchmark data is 149,718 futex calls over an 8 s run, so this is ~19K/sec, not ~150K/sec. The current phrasing overstates the overhead by about 8×.

♻️ Suggested text fix
-    a park-free AOF-writer poll that removes a ~150K/sec futex-wake storm on the shard thread under `everysec`,
+    a park-free AOF-writer poll that removes a ~19K/sec futex-wake storm on the shard thread under `everysec`,
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
The wins come from per-batch group commit (one fsync per pipeline batch, not per command), a single coalesced `write_all` per batch, a park-free AOF-writer poll that removes a ~150K/sec futex-wake storm on the shard thread under `everysec`, and coalesced pub/sub delivery writes. Durability is unchanged (`always` = RPO 0, `everysec` = RPO ≤ 1s), verified by the SIGKILL crash-recovery matrix. Full detail: `BENCHMARK.md` §7.3.
The wins come from per-batch group commit (one fsync per pipeline batch, not per command), a single coalesced `write_all` per batch, a park-free AOF-writer poll that removes a ~19K/sec futex-wake storm on the shard thread under `everysec`, and coalesced pub/sub delivery writes. Durability is unchanged (`always` = RPO 0, `everysec` = RPO ≤ 1s), verified by the SIGKILL crash-recovery matrix. Full detail: `BENCHMARK.md` §7.3.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/benchmarks.md` around lines 109 - 110, The benchmark wording overstates
the futex-wake rate in the performance summary. Update the futex-storm
description in the benchmarks text to reflect the measured rate from the data
(~19K/sec over the 8 s run, based on 149,718 futex calls) and keep the rest of
the AOF-writer/pipeline batching explanation unchanged; use the existing
benchmark narrative around the AOF-writer poll and everysec behavior to locate
the sentence.

## Latency

| Metric | Redis | Moon | Improvement |
Expand Down
2 changes: 1 addition & 1 deletion docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,7 @@ All options are available as command-line flags. Run `moon --help` for the full
| Flag | Default | Description |
|------|---------|-------------|
| `--appendonly` | `yes` | Enable AOF persistence (`yes`/`no`) — Moon is durable by default |
| `--appendfsync` | `everysec` | AOF fsync policy (`always`/`everysec`/`no`) |
| `--appendfsync` | `everysec` | AOF fsync policy (`always`/`everysec`/`no`). `everysec` SET is ~1.32× Redis at pipeline depth and at parity non-pipelined; `always` (RPO 0) is fsync-device-bound — parity non-pipelined, ~0.91× Redis at depth. See `BENCHMARK.md` §7.3 |
| `--aof-fsync-timeout-ms` | `2000` | Bound on a write's wait for durability — the fsync ack under `always`, writer-queue backpressure under `everysec` (0 = unbounded) |
| `--wal-kv-log` | `auto` | KV logging into the per-shard WAL. `auto`: skipped while the AOF is the recovery authority and no CDC subscriber is attached (halves write volume at `--shards >= 2`); `on`: always log (needed for PITR / full CDC history with AOF on); `off`: never |
| `--appendfilename` | `appendonly.aof` | AOF filename |
Expand Down
28 changes: 26 additions & 2 deletions docs/production-guide.md
Original file line number Diff line number Diff line change
Expand Up @@ -297,12 +297,36 @@ moon --appendonly yes --appendfsync everysec --dir /data

| `appendfsync` | Durability | Performance Impact |
|---|---|---|
| `always` | RPO = 0 (zero data loss) | Highest latency; suitable for financial data |
| `everysec` | RPO <= 1 second | Recommended default; negligible throughput impact |
| `always` | RPO = 0 (zero data loss) | fsync-per-batch under pipelining; P1 is fsync-device-bound (parity with any engine), P16 ~0.91x Redis |
| `everysec` | RPO <= 1 second | Recommended default; **P16 SET ~1.32x Redis**, P1 at parity |
| `no` | OS flush window (minutes) | Cache-mode only; do not use for primary storage |

Moon's per-shard WAL avoids the global serialization bottleneck that Redis's single AOF file creates. The AOF advantage over Redis **grows** with pipeline depth (2.75x at p=64).

**Write-path internals (how the AOF writer keeps up).** These are automatic — no
tuning knobs — but understanding them explains the durability/throughput tradeoff:

- **Group commit.** Under `appendfsync always`, pipelined writes are enqueued
fire-and-forget and the whole batch is made durable by ONE `fsync_barrier` (not
one fsync per command). A 16-deep pipeline therefore pays ~1 fsync, not 16.
The fsync still runs strictly before the client's `+OK` (RPO = 0 is preserved);
a failed batch returns an error to every write in it, never a silent success.
- **Coalesced batch write.** Each group-commit batch is written to the AOF with a
single contiguous `write_all`, not one `write(2)` per record — so the writer
thread is not syscall-bound at high pipeline depth (this is what makes
`everysec` P16 beat Redis rather than trail it).
- **Park-free writer poll (`everysec`/`no`).** The AOF writer thread does not park
in a blocking channel receive; it polls. This matters because a parked receiver
forces every shard thread to issue a futex wake on each write — at non-pipelined
`everysec` load that was ~150k wakes/sec of pure overhead on the hot path.
Polling keeps producer enqueues as plain userspace atomics, restoring P1 parity
with Redis. Under `always` the writer still parks (the client is already blocked
on the fsync ack, so receive latency there is client-visible RTT anyway).
Comment on lines +318 to +324

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Fix the futex-rate wording here too.

This repeats the same benchmark figure as docs/benchmarks.md: 149,718 futex calls in an 8 s run, so the rate is ~19K/sec, not ~150K/sec. The current text inflates the non-pipelined overhead and weakens the explanation.

♻️ Suggested text fix
-  a parked receiver forces every shard thread to issue a futex wake on each write — at non-pipelined
-  `everysec` load that was ~150k wakes/sec of pure overhead on the hot path.
+  a parked receiver forces every shard thread to issue a futex wake on each write — at non-pipelined
+  `everysec` load that was ~19k wakes/sec (149,718 total over an 8 s run) of pure overhead on the hot path.
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
- **Park-free writer poll (`everysec`/`no`).** The AOF writer thread does not park
in a blocking channel receive; it polls. This matters because a parked receiver
forces every shard thread to issue a futex wake on each write — at non-pipelined
`everysec` load that was ~150k wakes/sec of pure overhead on the hot path.
Polling keeps producer enqueues as plain userspace atomics, restoring P1 parity
with Redis. Under `always` the writer still parks (the client is already blocked
on the fsync ack, so receive latency there is client-visible RTT anyway).
- **Park-free writer poll (`everysec`/`no`).** The AOF writer thread does not park
in a blocking channel receive; it polls. This matters because a parked receiver
forces every shard thread to issue a futex wake on each write — at non-pipelined
`everysec` load that was ~19k wakes/sec (149,718 total over an 8 s run) of pure overhead on the hot path.
Polling keeps producer enqueues as plain userspace atomics, restoring P1 parity
with Redis. Under `always` the writer still parks (the client is already blocked
on the fsync ack, so receive latency there is client-visible RTT anyway).
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/production-guide.md` around lines 318 - 324, The futex-rate wording in
the Park-free writer poll explanation is overstated and should match the
benchmark figure already used elsewhere. Update the numeric claim in the
relevant prose near the AOF writer polling discussion to reflect about 149,718
futex calls over 8 seconds, i.e. roughly 19K/sec, and keep the rest of the
explanation intact. Use the existing "Park-free writer poll (`everysec`/`no`)"
section in docs/production-guide.md as the anchor for the edit.


None of these weaken durability: `always` remains RPO = 0, `everysec` remains
RPO ≤ 1 s, validated by the SIGKILL crash-recovery matrix (100% of acked writes
recovered). See `BENCHMARK.md` §7.3 for the measured before/after matrix.

### RDB snapshots

```bash
Expand Down
Loading