Estimate how much load AuthServer can handle before adopting it in-house, using k6.
- Endpoint under test:
POST /user/login - Real AWS KMS was replaced with an in-process mock signer (
scripts/run_mock_kms.py) to isolate server/DB performance from KMS network latency, and to avoid KMS cost/rate-limit noise during repeated runs. Test-only, never use against dev/live. - Load pattern: a staircase ramp (
k6/login_stress.js) stepping VUs 100 → 300 → 600 → 1000, holding at each step, so the breaking point could be located rather than just pass/fail at one VU count.
All numbers below are from the same staircase test against the mocked-KMS server.
| Config | req/s | error rate | median latency | p95 latency |
|---|---|---|---|---|
| 1 worker, default pools | ~180-190 | 0% | 335-760ms | 5.8-7.0s |
| 1 worker, DB pools enlarged (see below) | ~180-190 | 0% | ~same | ~same |
| 4 workers, oversized per-worker pool (20/30) | 190 | 36.2% | 8.7ms (success only) | 3.5s |
| 4 workers, right-sized per-worker pool (10/10) | 379.5 | 0.87% | 14.5ms | 1.52s |
- Hypothesis: DB connection pool too small.
main_app.pycreated a rawasyncpgpool with the library default (min_size=10, max_size=10). Bumping it to 50 had zero effect — turned out this pool is dead code, never read anywhere else in the codebase. The actual ORM queries go through a separate SQLAlchemy engine (src/core/database.py). - Hypothesis: SQLAlchemy pool too small. Its default (
pool_size=5, max_overflow=10= 15 total) was raised to 50. Still no effect on throughput — ruled out. - Real cause: single-process GIL bottleneck. With one worker process,
htopshowed the process pinned around ~72% of one core (of 24 available) while throughput stayed flat regardless of pool size. FastAPI/asyncio concurrency only removes I/O-wait time from the critical path; it does not parallelize the CPU-bound segment of each request (Pydantic validation, SQLAlchemy query building/row mapping, JSON serialization, middleware overhead). That segment runs on a single Python thread (GIL), capping single-process throughput at roughly1 / (CPU time per request)≈ 188 req/s, i.e. ~5.3ms of CPU-bound work per login request. (Note: JWT signing itself was not part of this bottleneck — it already ran in a thread pool viarun_in_executor, and the underlying crypto call releases the GIL.) - Confirmed via multi-process scaling. Running 4 worker processes (
scripts/run_mock_kms.py <workers>, usinguvicorn.run("scripts.run_mock_kms:app", workers=N)) gave each process its own interpreter/GIL, so the CPU-bound segments run truly in parallel across cores. - New bottleneck surfaced: Postgres
max_connections. 4 workers × the enlarged per-worker pool (20+30=50) requested up to 200 connections against Postgres's defaultmax_connections=100, causing a 36.2% failure rate (fast rejections, not timeouts). Shrinking the per-worker pool topool_size=10, max_overflow=10(80 total across 4 workers, under the 100 limit) fixed this: throughput roughly doubled (188 → 379.5 req/s) and errors dropped to 0.87%.
src/core/database.py:echo=False,pool_size=10,max_overflow=10(wasecho=Truewith library defaults5/10).scripts/run_mock_kms.py(new, test-only): mocksJwtLogic.initializewith an in-process RSA signer; supportsworkersargument for multi-process runs.k6/*.js,k6/users.json(new): smoke, login, multi-user, and staircase stress test scripts.
- Find the actual max sustainable throughput (p95 under a target SLA) with the tuned 4-worker config — VU 1000 still breaks the
p95<300msthreshold. - Re-run the tuned config against real AWS KMS to get a production-realistic number (mocking removed real KMS network latency from the picture).
- Worker scaling curve (2 / 4 / 8 workers) to see how far throughput scales linearly before Postgres/Redis (both single shared instances) becomes the ceiling.
- A soak test (sustained moderate load for 10-20 min) to catch connection/memory leaks that a short stress test can't.