A practical collection of English-language notes for understanding, building, testing, debugging, and profiling ClickHouse. The guides focus on source-code paths, storage internals, query execution, and reproducible operational workflows.
This is a community-maintained learning resource and is not official ClickHouse documentation. Details may vary by ClickHouse version.
- MergeTree interfaces
- MergeTree read path
- MergeTree read modes
- MergeTree Granules, Object Storage (S3/Azure), and Zero-Copy Replication
- Broken parts
- Merge implementation notes
- Disk abstractions
- MinIO setup
- SIMD Vectorization (AVX2/AVX-512/ARM Neon), Memory Infrastructure (PODArray/Arena) & FlameGraphs
- CPU and memory flame graphs
- Profiling release builds with perf
- Dynamic and JSON storage models
- Map column design and storage
- Expression trees and actions
- Function return types
- Utility classes
- Distributed query send/receive path
- Hash join internals
- Text skip index internals
apply_mutations_on_flyinternals- Python client and concurrent insert example
Start with the topic closest to the subsystem you are investigating. Most pages assume familiarity with Linux, C++, SQL, and the ClickHouse source tree. Commands and source paths should be checked against the ClickHouse version you are using.
Corrections, clearer explanations, and version-specific updates are welcome. Keep examples reproducible, identify the ClickHouse version when behavior is version-dependent, and link to relevant source files where possible.