Skip to content

New benchmark: Atomic-AffQRF for matbench_steels - #371

Open
ToukoUrsin wants to merge 1 commit into
materialsproject:mainfrom
ToukoUrsin:atomic-affqrf-steels
Open

ToukoUrsin wants to merge 1 commit into
materialsproject:mainfrom
ToukoUrsin:atomic-affqrf-steels

Conversation

@ToukoUrsin

Copy link
Copy Markdown

Benchmark submission: Atomic-AffQRF on matbench_steels

Subset submission (one task): matbench_v0.1 / matbench_steels.

Brief description of the algorithm

Composition-only model. Atomic-fraction vectors (four fixed transforms) feed 16 fixed ExtraTrees configurations; each gives a leaf-distribution conditional mean and lower median, i.e. 32 predictors. Inside each official outer training fold, five inner folds produce out-of-fold predictions. 128 Bayesian-bootstrap weighted-MAE linear programs fit nonnegative coefficients (sum <= 2) plus a free intercept on those OOF predictions, and the averaged coefficients/intercept combine the final 900-tree forests. Outer-test labels are not used for fitting or calibration.

Results (reference seed 417, declared before scoring)

fold MAE (MPa)
0 90.3882
1 67.5943
2 73.3462
3 77.1833
4 69.2008
mean 75.5426 (std 8.14, mean RMSE 107.48)

Current best on the matbench_steels leaderboard is TPOT-Mat at 79.9468 MPa. Seed repeats 0-6 give 75.77 +/- 0.33 MPa (sample SD); all run seeds and two ablation aggregators are listed in all_seed_scores.csv.

Checked on current main: MatbenchBenchmark.from_file(...) reports is_valid=True, is_recorded=True for matbench_steels; fold MAEs recomputed directly from the stored predictions and the official test sets match the recorded scores; python -m unittest scripts/test_submission.py passes.

Limitations

  • The official folds were used repeatedly during method development. Each run fits only on outer-training labels, but choosing the model family across runs may be optimistic; this is not an untouched holdout.
  • On a separate composition-grouped stress test (not the official task), seed 417 scores 99.56 MPa versus 99.09 MPa for the released TPOT pipeline refit with the same seed, so no chemistry-transfer advantage is claimed.

Included files (benchmarks/matbench_v0.1_Atomic_Affine_QRF/)

  • results.json.gz, info.json, notebook.ipynb: required files. The notebook rebuilds all forests from raw inputs and asserts the fold dictionaries equal results.json.gz exactly.
  • run.py, inputs/base_qrf.py, protocol.json, inputs/protocol.json: fitting source and frozen settings.
  • inputs/matbench_steels.json.gz, inputs/steels_splits.json: snapshots of the official dataset and folds (~50 KB).
  • fresh.py, export.py, reproduce_entry.py: command-line rebuild/export (python -B reproduce_entry.py --output fresh_script_review --workers 3).
  • preflight.py, environment_lock.json, manifest.json, upstream_source_hashes.json: refuse to run on mismatched package versions, changed inputs or existing fit caches. Exact replay needs CPython 3.13.16 with requirements.txt.
  • README.md, all_seed_scores.csv: method summary and all seed scores.

Please add the new_benchmark label (I cannot set labels on this repository).

Prepared with AI assistance; reviewed and tested by me.

Signed-off-by: Touko Ursin <touko@heliosone.fi>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant