Repository navigation
docker: make the shard count configurable, defaulting to conway's 8 - #39
Merged
Merged
Conversation
The compose stack hardcoded four shards while testnet-conway runs eight (linera-infra argo/app-values/validator/base-values.yaml numShards: 8), so an external validator following our own docs ran half our capacity. Shard assignment is hash(validator_public_key, chain_id) % num_shards and lives in ValidatorInternalNetworkConfig, so the count is per-validator capacity, not something the network agrees on - which is why this was a sizing gap rather than a correctness one. Compose cannot template a variable number of services, so define eight and put shard-4..7 behind per-shard profiles. Shards 0..3 carry no profile: a deployment that predates this has no COMPOSE_PROFILES in .env and keeps exactly the four it already ran. The trap is that server.json holds the shard list and generate_validator_keys deliberately never regenerates it, because that would rotate the signing key and drop the validator from the committee. A re-run that took the new default would therefore start shards indexing past the end of that list, and ValidatorInternalNetworkPreConfig::shard is a plain Vec index - every one of them panics on boot. So the resolution order is: an existing server.json pins the count, an explicit --num-shards that disagrees is refused, and 8 applies only to a fresh deployment. Missing jq is fatal there rather than a silent fallthrough, for the same reason. Prometheus and Alloy now discover shards through the Docker daemon and take the index from each container's linera.shard label, so neither names a shard and neither alerts on shards a smaller deployment does not run. Without that the LineraValidatorDown alert fixed in #38 would fire on every phantom target. The hardware budget is unchanged: 8 x 6 GiB is the same 48 GiB as 4 x 12 GiB, so the reference box stays 16 cores / 128 GB. Note helm/linera-validator still defaults shards.replicas to 4. Same drift, different deployment path, and changing it needs the same care about existing StatefulSets - left alone here.
…able The jq guard checked that jq exists, then ran it with 2>/dev/null || true. A jq that is present but fails - wrong version, broken install, anything non-zero - therefore produced an empty count, fell through to DEFAULT_NUM_SHARDS, and migrated a running 4-shard validator to 8. Exit 0, no warning. That is exactly the migration the pinning exists to prevent. Found by running the script with a jq stub that exits 127: NUM_SHARDS went from 4 to 8 silently. Now the presence check and the parse are separate, neither swallows failure, and an unreadable server.json aborts. Regression test covers the parse failure; the jq-absent branch was verified by hand against a PATH with no jq on it.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
The compose stack hardcoded four shards; testnet-conway runs eight (
linera-infraargo/app-values/validator/base-values.yaml:18,numShards: 8). An external operator following our docs was therefore running half our shard count against a network we compare them to on latency — which is exactly the situation that prompted this.It is a capacity gap, not a correctness one. Shard assignment is per-validator by construction (
linera-rpc/src/config.rs:309-315):It lives in
ValidatorInternalNetworkConfig. Nothing outside the validator knows or cares how many shards it has.The trap this had to avoid
generate_validator_keysdeliberately never regenerates an existingserver.json(deploy-validator.sh:247-249) — doing so would rotate the signing key and drop the validator from the committee. Butserver.jsonis also where the shard list lives, and it is what the proxy and every shard actually read.So on an existing validator, a re-run that picked up the new default would write
validator-config.tomlwith 8 shards, leaveserver.jsonat 4, and start containers with--shard 4..7.ValidatorInternalNetworkPreConfig::shardis a plainVecindex (config.rs:319-321), so all four panic on boot.Resolution order is therefore:
server.jsonpins the count,--num-shardsthat disagrees with it is refused, not applied,A missing
jqis fatal in that path rather than a silent fallthrough, since falling through means guessing, and guessing here migrates a running validator.Changes
shard-4..7behind per-shard profiles. Shards 0–3 carry no profile, so a deployment that predates this has noCOMPOSE_PROFILESand keeps exactly the four it already ran.--num-shardsaccepts 4–8 anddeploy-validator.shwrites the matchingCOMPOSE_PROFILESinto.env.docker_sd_configs/discovery.docker) and read the index off each container's newlinera.shardlabel. Neither file names a shard any more. This is load-bearing: with static target lists, theLineraValidatorDownalert fixed in docker: make the shipped alert rules able to fire #38 would fire on every shard a smaller deployment does not run. Prometheus gets the Docker socket read-only; Alloy already had it.docker-compose.remote-scylla.yamlmapsscyllafor the new shards too — without it an 8-shard remote-Scylla deployment would have four shards unable to resolve the database.--num-shardshelp,.envtemplate.Sizing
The reference box does not change. 8 × 6 GiB is the same 48 GiB as 4 × 12 GiB, so the memory column still totals ≈110 GiB on 16 cores / 128 GB. A smaller shard count buys a bigger per-shard cache rather than a smaller host.
Verification
tests/deploy-validator-test.shgains three cases. The regression test was checked against the bug it exists to catch — with theserver.jsonpinning disabled:and green with it restored. Also verified locally:
docker compose config --servicesyields 4 shards with no profiles, 8 with all four, 6 with two — and every overlay combination renders, includingremote-scyllawhere all 8 shards plus proxy and shard-init get thescylla=mapping..envdrives a real compose render at 8 shards (the CI step's path).promtool check config --syntax-onlyaccepts thedocker_sd_configsblock;check-rule-job-selectors.py,yamllint --strict,shellcheckall clean.Out of scope
helm/linera-validator/values.yamlstill defaultsshards.replicas: 4— the same drift on the other deployment path. Changing it needs the same care about existing StatefulSets and theirserver.json, so it is deliberately not in this PR.