Skip to content

docker: make the shard count configurable, defaulting to conway's 8 - #39

Merged
ndr-ds merged 2 commits into
mainfrom
ndr-ds/configurable-shard-count
Aug 31, 2026
Merged

ndr-ds merged 2 commits into
mainfrom
ndr-ds/configurable-shard-count

Conversation

@ndr-ds

@ndr-ds ndr-ds commented Aug 31, 2026

Copy link
Copy Markdown
Collaborator

Motivation

The compose stack hardcoded four shards; testnet-conway runs eight (linera-infra argo/app-values/validator/base-values.yaml:18, numShards: 8). An external operator following our docs was therefore running half our shard count against a network we compare them to on latency — which is exactly the situation that prompted this.

It is a capacity gap, not a correctness one. Shard assignment is per-validator by construction (linera-rpc/src/config.rs:309-315):

pub fn get_shard_id(&self, chain_id: ChainId) -> ShardId {
    let mut s = DefaultHasher::new();
    self.public_key.hash(&mut s);   // deliberately salted per validator
    chain_id.hash(&mut s);
    (s.finish() as ShardId) % self.shards.len()
}

It lives in ValidatorInternalNetworkConfig. Nothing outside the validator knows or cares how many shards it has.

The trap this had to avoid

generate_validator_keys deliberately never regenerates an existing server.json (deploy-validator.sh:247-249) — doing so would rotate the signing key and drop the validator from the committee. But server.json is also where the shard list lives, and it is what the proxy and every shard actually read.

So on an existing validator, a re-run that picked up the new default would write validator-config.toml with 8 shards, leave server.json at 4, and start containers with --shard 4..7. ValidatorInternalNetworkPreConfig::shard is a plain Vec index (config.rs:319-321), so all four panic on boot.

Resolution order is therefore:

  1. an existing server.json pins the count,
  2. an explicit --num-shards that disagrees with it is refused, not applied,
  3. the new default of 8 applies only to a fresh deployment.

A missing jq is fatal in that path rather than a silent fallthrough, since falling through means guessing, and guessing here migrates a running validator.

Changes

  • Eight shard services in compose; shard-4..7 behind per-shard profiles. Shards 0–3 carry no profile, so a deployment that predates this has no COMPOSE_PROFILES and keeps exactly the four it already ran. --num-shards accepts 4–8 and deploy-validator.sh writes the matching COMPOSE_PROFILES into .env.
  • Prometheus and Alloy discover shards through the Docker daemon (docker_sd_configs / discovery.docker) and read the index off each container's new linera.shard label. Neither file names a shard any more. This is load-bearing: with static target lists, the LineraValidatorDown alert fixed in docker: make the shipped alert rules able to fire #38 would fire on every shard a smaller deployment does not run. Prometheus gets the Docker socket read-only; Alloy already had it.
  • docker-compose.remote-scylla.yaml maps scylla for the new shards too — without it an 8-shard remote-Scylla deployment would have four shards unable to resolve the database.
  • Docs: hardware table, a "Changing the shard count" section, --num-shards help, .env template.

Sizing

The reference box does not change. 8 × 6 GiB is the same 48 GiB as 4 × 12 GiB, so the memory column still totals ≈110 GiB on 16 cores / 128 GB. A smaller shard count buys a bigger per-shard cache rather than a smaller host.

Verification

tests/deploy-validator-test.sh gains three cases. The regression test was checked against the bug it exists to catch — with the server.json pinning disabled:

== a re-run never changes the shard count in server.json
  FAIL: NUM_SHARDS pinned by server.json: want '4', got '8'
  FAIL: no extra shard profiles enabled: want '', got 'shard-4,shard-5,shard-6,shard-7'
  FAIL: validator-config.toml stays at 4: want '4', got '8'
  FAIL: --num-shards 8 was accepted against a 4-shard server.json
  FAIL: NUM_SHARDS unchanged after refusal: want '4', got '8'
5 assertion(s) failed

and green with it restored. Also verified locally:

  • docker compose config --services yields 4 shards with no profiles, 8 with all four, 6 with two — and every overlay combination renders, including remote-scylla where all 8 shards plus proxy and shard-init get the scylla= mapping.
  • The generated .env drives a real compose render at 8 shards (the CI step's path).
  • promtool check config --syntax-only accepts the docker_sd_configs block; check-rule-job-selectors.py, yamllint --strict, shellcheck all clean.

Out of scope

helm/linera-validator/values.yaml still defaults shards.replicas: 4 — the same drift on the other deployment path. Changing it needs the same care about existing StatefulSets and their server.json, so it is deliberately not in this PR.

ndr-ds added 2 commits August 31, 2026 21:44
The compose stack hardcoded four shards while testnet-conway runs eight
(linera-infra argo/app-values/validator/base-values.yaml numShards: 8), so an
external validator following our own docs ran half our capacity. Shard
assignment is hash(validator_public_key, chain_id) % num_shards and lives in
ValidatorInternalNetworkConfig, so the count is per-validator capacity, not
something the network agrees on - which is why this was a sizing gap rather
than a correctness one.

Compose cannot template a variable number of services, so define eight and put
shard-4..7 behind per-shard profiles. Shards 0..3 carry no profile: a
deployment that predates this has no COMPOSE_PROFILES in .env and keeps exactly
the four it already ran.

The trap is that server.json holds the shard list and generate_validator_keys
deliberately never regenerates it, because that would rotate the signing key
and drop the validator from the committee. A re-run that took the new default
would therefore start shards indexing past the end of that list, and
ValidatorInternalNetworkPreConfig::shard is a plain Vec index - every one of
them panics on boot. So the resolution order is: an existing server.json pins
the count, an explicit --num-shards that disagrees is refused, and 8 applies
only to a fresh deployment. Missing jq is fatal there rather than a silent
fallthrough, for the same reason.

Prometheus and Alloy now discover shards through the Docker daemon and take the
index from each container's linera.shard label, so neither names a shard and
neither alerts on shards a smaller deployment does not run. Without that the
LineraValidatorDown alert fixed in #38 would fire on every phantom target.

The hardware budget is unchanged: 8 x 6 GiB is the same 48 GiB as 4 x 12 GiB,
so the reference box stays 16 cores / 128 GB.

Note helm/linera-validator still defaults shards.replicas to 4. Same drift,
different deployment path, and changing it needs the same care about existing
StatefulSets - left alone here.
…able

The jq guard checked that jq exists, then ran it with 2>/dev/null || true.
A jq that is present but fails - wrong version, broken install, anything
non-zero - therefore produced an empty count, fell through to
DEFAULT_NUM_SHARDS, and migrated a running 4-shard validator to 8. Exit 0, no
warning. That is exactly the migration the pinning exists to prevent.

Found by running the script with a jq stub that exits 127: NUM_SHARDS went from
4 to 8 silently. Now the presence check and the parse are separate, neither
swallows failure, and an unreadable server.json aborts. Regression test covers
the parse failure; the jq-absent branch was verified by hand against a PATH
with no jq on it.
@ndr-ds
ndr-ds marked this pull request as ready for review August 31, 2026 22:14
@ndr-ds
ndr-ds merged commit def9855 into main Aug 31, 2026
11 checks passed
@ndr-ds
ndr-ds deleted the ndr-ds/configurable-shard-count branch August 31, 2026 22:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant