diff --git a/.agents/skills/add-transformer/SKILL.md b/.agents/skills/add-transformer/SKILL.md new file mode 100644 index 00000000..a9fc368a --- /dev/null +++ b/.agents/skills/add-transformer/SKILL.md @@ -0,0 +1,209 @@ +--- +name: add-transformer +description: Implement a new Spark transformer and paired Keras layer with tests. Includes optional estimator path for transforms needing fit(). +--- + +# Add Transformer + +Use this skill when adding a new preprocessing operation to Kamae. Every operation requires a paired Keras layer and Spark transformer. If the operation needs to learn from data (e.g., computing mean/stddev), it also needs an estimator. + +## Before you start + +Read the reference implementation that matches your case: +- **Simple numeric op:** `src/kamae/spark/transformers/multiply.py` + `src/kamae/keras/core/layers/multiply.py` +- **With estimator:** `src/kamae/spark/estimators/standard_scale.py` + `src/kamae/spark/transformers/standard_scale.py` +- **TF-only (string/datetime):** `src/kamae/spark/transformers/string_index.py` + `src/kamae/keras/tensorflow/layers/string_index.py` + +For the full guide with code examples and checklists, read `docs/adding_transformer.md`. + +## Step 1: Decide placement + +Does your layer need TensorFlow-specific ops (string manipulation, datetime parsing)? +- **Yes** -> `src/kamae/keras/tensorflow/layers/.py`, import `TENSORFLOW_ONLY` from `kamae.keras.core.backend` +- **No** -> `src/kamae/keras/core/layers/.py`, use only `keras.ops`, import `ALL_BACKENDS` from `kamae.keras.core.backend` + +## Step 2: Implement the Keras layer + +Create `src/kamae/keras/{core or tensorflow}/layers/.py`: + +```python +from typing import Any, Dict, List, Optional + +import keras +from keras import KerasTensor, ops + +import kamae +from kamae.keras.core.backend import ALL_BACKENDS +from kamae.keras.core.base import BaseLayer + + +@keras.saving.register_keras_serializable(package=kamae.__name__) +class Layer(BaseLayer): + supported_backends = ALL_BACKENDS # or TENSORFLOW_ONLY + jit_compatible = True # or False + + def __init__( + self, + name: Optional[str] = None, + input_dtype: Optional[str] = None, + output_dtype: Optional[str] = None, + # add custom params here + **kwargs: Any, + ) -> None: + super().__init__(name=name, input_dtype=input_dtype, output_dtype=output_dtype, **kwargs) + # store custom params as instance attributes + + @property + def compatible_dtypes(self) -> Optional[List[str]]: + return ["float32", "float64"] # or None for all dtypes + + def _call(self, inputs: KerasTensor, **kwargs: Any) -> KerasTensor: + # transform logic using keras.ops + ... + + def get_config(self) -> Dict[str, Any]: + config = super().get_config() + config.update({...}) # add custom params + return config +``` + +Key points: +- Constructor must accept `name`, `input_dtype`, `output_dtype` and pass them to `super().__init__` +- `compatible_dtypes` returns dtype *strings* (e.g., `"float32"`), not Spark types +- Use `@allow_single_or_multiple_tensor_input` decorator from `kamae.keras.core.utils.input_utils` if the layer accepts single or multiple tensor inputs + +## Step 3: Implement the Spark transformer + +Create `src/kamae/spark/transformers/.py`: + +```python +# pylint: disable=unused-argument +# pylint: disable=invalid-name +# pylint: disable=too-many-ancestors +# pylint: disable=no-member +from typing import List, Optional + +import keras +from pyspark import keyword_only +from pyspark.sql import DataFrame +from pyspark.sql.types import DataType, DoubleType, FloatType + +from kamae.keras.core.backend import ALL_BACKENDS +from kamae.keras.{core or tensorflow}.layers import Layer +from kamae.spark.params import SingleInputSingleOutputParams # pick the right mixin +from .base import BaseTransformer + + +class Transformer( + BaseTransformer, + SingleInputSingleOutputParams, + # add shared param classes if needed +): + supported_backends = ALL_BACKENDS + jit_compatible = True + + @keyword_only + def __init__( + self, + inputCol: Optional[str] = None, + outputCol: Optional[str] = None, + inputDtype: Optional[str] = None, + outputDtype: Optional[str] = None, + layerName: Optional[str] = None, + # add custom params here + ) -> None: + super().__init__() + # self._setDefault(param=default) for custom params with defaults + kwargs = self._input_kwargs + self.setParams(**kwargs) + + @property + def compatible_dtypes(self) -> Optional[List[DataType]]: + return [FloatType(), DoubleType()] # Spark types, not strings + + def _transform(self, dataset: DataFrame) -> DataFrame: + # transform logic + return dataset.withColumn(self.getOutputCol(), ...) + + def get_keras_layer(self) -> keras.layers.Layer: + return Layer( + name=self.getLayerName(), + input_dtype=self.getInputKerasDtype(), + output_dtype=self.getOutputKerasDtype(), + # pass custom params + ) +``` + +Key points: +- Always include the four pylint disable comments at the top +- `@keyword_only` on `__init__`, then `super().__init__()` then `self.setParams(**kwargs)` +- Custom params need setter methods named `set` — either via a shared params class or defined inline +- `compatible_dtypes` returns Spark `DataType` instances, not strings +- Use transform utilities from `kamae.spark.utils.transform_utils` for nested array support + +## Step 4 (if fit needed): Implement the Spark estimator + +Create `src/kamae/spark/estimators/.py` (separate file from the transformer): + +```python +from kamae.spark.estimators.base import BaseEstimator + +class Estimator(BaseEstimator, SingleInputSingleOutputParams, ...): + # same constructor pattern as transformer + + def _fit(self, dataset: DataFrame) -> "Transformer": + # compute statistics from dataset + return Transformer( + inputCol=self.getInputCol(), + outputCol=self.getOutputCol(), + layerName=self.getLayerName(), + inputDtype=self.getInputDtype(), + outputDtype=self.getOutputDtype(), + # pass computed statistics + ) +``` + +Reference: `src/kamae/spark/estimators/standard_scale.py` + +## Step 5: Register + +Add imports (alphabetical order) to: + +1. **Layer `__init__.py`** — `src/kamae/keras/{core or tensorflow}/layers/__init__.py` + - Add the import line + - Add to the `__all__` list +2. **Transformer `__init__.py`** — `src/kamae/spark/transformers/__init__.py` +3. **Estimator `__init__.py`** (if applicable) — `src/kamae/spark/estimators/__init__.py` + +## Step 6: Write tests + +### Layer tests +Create `tests/kamae/keras/{core or tensorflow}/layers/test_.py`: +- Use `@pytest.mark.parametrize` with multiple input shapes, dtypes, and edge cases +- Assert on output dtype, shape, and values +- Use `tf.debugging.assert_near` for floating-point comparisons + +### Transformer tests +Create `tests/kamae/spark/transformers/test_.py`: +- Use the `spark_session` and `example_dataframe` fixtures from `conftest.py` +- Test transform output against expected DataFrames + +### Parity tests +Verify Spark transformer and Keras layer produce identical results: +- Transform with Spark, convert output to numpy +- Run Keras layer on same input, convert to numpy +- `np.testing.assert_almost_equal(spark_result, keras_result)` + +### Serialisation test +Add an entry to `tests/kamae/keras/test_layer_serialisation.py`: +- Import your layer +- Add a tuple to the `@pytest.mark.parametrize` list: `(YourLayer, [input_tensors], {kwargs}, skip_predict)` +- `skip_predict=True` only for non-deterministic layers (datetime/timestamp) + +## Step 7: Verify + +```bash +make format && make lint && make test +``` + +If parity fails, read `docs/achieving_type_parity.md` and `docs/achieving_shape_parity.md`. diff --git a/.agents/skills/verify/SKILL.md b/.agents/skills/verify/SKILL.md new file mode 100644 index 00000000..4424765e --- /dev/null +++ b/.agents/skills/verify/SKILL.md @@ -0,0 +1,55 @@ +--- +name: verify +description: Run the standard quality pipeline — formatting, linting, tests, and coverage checks. +--- + +# Verify + +Run this before declaring any task complete. + +## Quick check + +```bash +make format && make lint && make test +``` + +This runs: +- `black --check --diff` and `isort --check` (formatting) +- `flake8` with annotation checks and `pylint` with fail-under 5 (linting) +- `pytest -n auto` (all tests in parallel) + +## Full CI-equivalent check + +```bash +make test-cov +``` + +Runs tests with 80% coverage threshold and branch coverage enabled. + +## Fix formatting and license headers + +```bash +pre-commit run --all-files +``` + +This auto-fixes: +- Black formatting +- isort import ordering +- Apache 2.0 license headers on all `.py` files + +## Common failures + +| Failure | Fix | +|---|---| +| Missing license header | Run `pre-commit run --all-files` | +| flake8 ANN error | Add type annotations to function signatures | +| pylint score < 5 | Address pylint-reported issues; common disables for transformers: `unused-argument`, `invalid-name`, `too-many-ancestors`, `no-member` | +| Coverage < 80% | Add tests for uncovered branches — check `--cov-report term-missing` output | +| isort error | Run `pre-commit run --all-files` or `uv run isort src/kamae --profile black` | + +## CI matrix + +CI tests across a matrix — local `make test` runs one combination. The full matrix: +- Python: 3.10, 3.11, 3.12 +- PySpark: 3.4.1, 3.5.0 +- Keras: 3.3.0, 3.7.0, 3.10.0, 3.12.0 diff --git a/AGENTS.md b/AGENTS.md new file mode 100644 index 00000000..9cfb90ea --- /dev/null +++ b/AGENTS.md @@ -0,0 +1,131 @@ +# Kamae + +Kamae is a Python library that bridges Apache Spark preprocessing and Keras 3 model serving, eliminating training/serving skew by providing paired Spark transformers and Keras layers. + +## Architecture + +``` +src/kamae/ + keras/ + core/ + base.py # BaseLayer — all layers extend this + layers/ # Multi-backend layers (Keras ops only, all backends) + utils/ + tensorflow/ + layers/ # TF-only layers (string, datetime ops — TensorFlow backend only) + spark/ + transformers/ + base.py # BaseTransformer — all transformers extend this + estimators/ + base.py # BaseEstimator — for transforms needing fit() + pipeline/ # KamaeSparkPipeline — chains transformers for export + params/ + base.py # I/O param mixins (SingleInput*, MultiInput*) + shared.py # Reusable param classes (MathFloatConstantParams, etc.) + utils/ # Transform helpers, UDFs, array utils + graph/ # Pipeline DAG construction (NetworkX) + discovery.py # get_compatible_layers(), get_compatible_transformers() +tests/kamae/ # Mirrors src/ structure +examples/spark/ # Example pipelines +docs/ # Contributor guides +``` + +## Naming conventions + +| Concept | Class name | File | +|---|---|---| +| Keras layer | `Layer` | `keras/core/layers/.py` or `keras/tensorflow/layers/.py` | +| Spark transformer | `Transformer` | `spark/transformers/.py` | +| Spark estimator | `Estimator` | `spark/estimators/.py` | +| Params class | `Params` | `spark/params/shared.py` or inline | + +Use the verb stem: `StringIndex` not `StringIndexer`. File names are `snake_case`. + +## Backend placement + +- Uses only Keras 3 ops (`keras.ops`) -> `src/kamae/keras/core/layers/` +- Requires TensorFlow ops (strings, datetimes) -> `src/kamae/keras/tensorflow/layers/` + +## Param mixins + +From `src/kamae/spark/params/base.py`: +- `SingleInputSingleOutputParams` — one input col, one output col +- `SingleInputMultiOutputParams` — one input col, multiple output cols +- `MultiInputSingleOutputParams` — multiple input cols, one output col +- `MultiInputMultiOutputParams` — multiple input cols, multiple output cols + +Shared param classes in `src/kamae/spark/params/shared.py` (e.g., `MathFloatConstantParams`, `StandardScaleParams`). + +## Registration + +When adding new classes, update these `__init__.py` files (alphabetical order): + +| What | File | +|---|---| +| Core layer | `src/kamae/keras/core/layers/__init__.py` + `__all__` list | +| TF-only layer | `src/kamae/keras/tensorflow/layers/__init__.py` + `__all__` list | +| Transformer | `src/kamae/spark/transformers/__init__.py` | +| Estimator | `src/kamae/spark/estimators/__init__.py` | + +Also add a serialisation test entry in `tests/kamae/keras/test_layer_serialisation.py`. + +## Commands + +| Command | Purpose | +|---|---| +| `make install` | Install dependencies (uses uv) | +| `make test` | Run tests in parallel (`pytest -n auto`) | +| `make test-cov` | Run tests with 80% coverage threshold (branch coverage) | +| `make lint` | flake8 (annotation checks) + pylint (fail-under 5) | +| `make format` | Check formatting (Black + isort, check-only) | +| `make build` | Build wheel | +| `make run-example` | Run example Spark pipeline | +| `make docs` | Build and serve docs locally | +| `pre-commit run --all-files` | Fix formatting, imports, license headers | + +## Code style + +- **Formatter:** Black (88 char line length) +- **Imports:** isort with black profile +- **Linting:** flake8 with ANN checks, pylint >= 5.0 +- **License:** Apache 2.0 headers on all `.py` files (enforced by pre-commit) +- **Commits:** Conventional commits — `fix:` (patch), `feat:` (minor), `BREAKING CHANGE:` (major) +- **Coverage:** 80% minimum with branch coverage +- **Python:** 3.10+ for development +- **Deps:** managed with uv (`uv sync`) + +Common pylint disables at top of transformer/estimator files: +```python +# pylint: disable=unused-argument +# pylint: disable=invalid-name +# pylint: disable=too-many-ancestors +# pylint: disable=no-member +``` + +## Reference implementations + +| Complexity | Example | Files | +|---|---|---| +| Simple (no fit) | Multiply | `spark/transformers/multiply.py`, `keras/core/layers/multiply.py`, `tests/.../test_multiply.py` | +| With estimator | StandardScale | `spark/estimators/standard_scale.py`, `spark/transformers/standard_scale.py`, `keras/core/layers/standard_scale.py` | +| TF-only | StringIndex | `spark/transformers/string_index.py`, `keras/tensorflow/layers/string_index.py` | + +## Key documentation + +- `docs/adding_transformer.md` — full guide with code examples and checklists +- `docs/achieving_type_parity.md` — ensuring consistent dtypes between Spark and Keras +- `docs/achieving_shape_parity.md` — ensuring consistent shapes between Spark and Keras +- `docs/testing_inference.md` — validating outputs with TensorFlow Serving +- `.github/PULL_REQUEST_TEMPLATE.md` — PR checklist + +## Agent skills + +This repository uses agent-agnostic skills in `.agents/skills/`. + +Before starting a task, inspect the available skills by reading the `name` and `description` +frontmatter in each `.agents/skills/*/SKILL.md`. + +If a skill matches the task, read the full `SKILL.md` before acting. Use any referenced +files under that skill only when needed. + +Project skills take precedence over user/global skills.