Skip to content

Latest commit

 

History

History
47 lines (32 loc) · 4.31 KB

File metadata and controls

47 lines (32 loc) · 4.31 KB

한국어 버전은 아래에서 확인할 수 있습니다.

한국어 버전

Load Testing Report: AuthServer Capacity

Goal

Estimate how much load AuthServer can handle before adopting it in-house, using k6.

Setup

  • Endpoint under test: POST /user/login
  • Real AWS KMS was replaced with an in-process mock signer (scripts/run_mock_kms.py) to isolate server/DB performance from KMS network latency, and to avoid KMS cost/rate-limit noise during repeated runs. Test-only, never use against dev/live.
  • Load pattern: a staircase ramp (k6/login_stress.js) stepping VUs 100 → 300 → 600 → 1000, holding at each step, so the breaking point could be located rather than just pass/fail at one VU count.

Findings

All numbers below are from the same staircase test against the mocked-KMS server.

Config req/s error rate median latency p95 latency
1 worker, default pools ~180-190 0% 335-760ms 5.8-7.0s
1 worker, DB pools enlarged (see below) ~180-190 0% ~same ~same
4 workers, oversized per-worker pool (20/30) 190 36.2% 8.7ms (success only) 3.5s
4 workers, right-sized per-worker pool (10/10) 379.5 0.87% 14.5ms 1.52s

Root cause chain

  1. Hypothesis: DB connection pool too small. main_app.py created a raw asyncpg pool with the library default (min_size=10, max_size=10). Bumping it to 50 had zero effect — turned out this pool is dead code, never read anywhere else in the codebase. The actual ORM queries go through a separate SQLAlchemy engine (src/core/database.py).
  2. Hypothesis: SQLAlchemy pool too small. Its default (pool_size=5, max_overflow=10 = 15 total) was raised to 50. Still no effect on throughput — ruled out.
  3. Real cause: single-process GIL bottleneck. With one worker process, htop showed the process pinned around ~72% of one core (of 24 available) while throughput stayed flat regardless of pool size. FastAPI/asyncio concurrency only removes I/O-wait time from the critical path; it does not parallelize the CPU-bound segment of each request (Pydantic validation, SQLAlchemy query building/row mapping, JSON serialization, middleware overhead). That segment runs on a single Python thread (GIL), capping single-process throughput at roughly 1 / (CPU time per request) ≈ 188 req/s, i.e. ~5.3ms of CPU-bound work per login request. (Note: JWT signing itself was not part of this bottleneck — it already ran in a thread pool via run_in_executor, and the underlying crypto call releases the GIL.)
  4. Confirmed via multi-process scaling. Running 4 worker processes (scripts/run_mock_kms.py <workers>, using uvicorn.run("scripts.run_mock_kms:app", workers=N)) gave each process its own interpreter/GIL, so the CPU-bound segments run truly in parallel across cores.
  5. New bottleneck surfaced: Postgres max_connections. 4 workers × the enlarged per-worker pool (20+30=50) requested up to 200 connections against Postgres's default max_connections=100, causing a 36.2% failure rate (fast rejections, not timeouts). Shrinking the per-worker pool to pool_size=10, max_overflow=10 (80 total across 4 workers, under the 100 limit) fixed this: throughput roughly doubled (188 → 379.5 req/s) and errors dropped to 0.87%.

Code changes made during testing

  • src/core/database.py: echo=False, pool_size=10, max_overflow=10 (was echo=True with library defaults 5/10).
  • scripts/run_mock_kms.py (new, test-only): mocks JwtLogic.initialize with an in-process RSA signer; supports workers argument for multi-process runs.
  • k6/*.js, k6/users.json (new): smoke, login, multi-user, and staircase stress test scripts.

Open questions / suggested next steps

  • Find the actual max sustainable throughput (p95 under a target SLA) with the tuned 4-worker config — VU 1000 still breaks the p95<300ms threshold.
  • Re-run the tuned config against real AWS KMS to get a production-realistic number (mocking removed real KMS network latency from the picture).
  • Worker scaling curve (2 / 4 / 8 workers) to see how far throughput scales linearly before Postgres/Redis (both single shared instances) becomes the ceiling.
  • A soak test (sustained moderate load for 10-20 min) to catch connection/memory leaks that a short stress test can't.