Repository navigation
New benchmark: Atomic-AffQRF for matbench_steels - #371
Open
ToukoUrsin wants to merge 1 commit into
Open
ToukoUrsin wants to merge 1 commit into
ToukoUrsin wants to merge 1 commit into
Conversation
Signed-off-by: Touko Ursin <touko@heliosone.fi>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Benchmark submission: Atomic-AffQRF on
matbench_steelsSubset submission (one task):
matbench_v0.1/matbench_steels.Brief description of the algorithm
Composition-only model. Atomic-fraction vectors (four fixed transforms) feed 16 fixed ExtraTrees configurations; each gives a leaf-distribution conditional mean and lower median, i.e. 32 predictors. Inside each official outer training fold, five inner folds produce out-of-fold predictions. 128 Bayesian-bootstrap weighted-MAE linear programs fit nonnegative coefficients (sum <= 2) plus a free intercept on those OOF predictions, and the averaged coefficients/intercept combine the final 900-tree forests. Outer-test labels are not used for fitting or calibration.
Results (reference seed 417, declared before scoring)
Current best on the
matbench_steelsleaderboard is TPOT-Mat at 79.9468 MPa. Seed repeats 0-6 give 75.77 +/- 0.33 MPa (sample SD); all run seeds and two ablation aggregators are listed inall_seed_scores.csv.Checked on current
main:MatbenchBenchmark.from_file(...)reportsis_valid=True,is_recorded=Trueformatbench_steels; fold MAEs recomputed directly from the stored predictions and the official test sets match the recorded scores;python -m unittest scripts/test_submission.pypasses.Limitations
Included files (
benchmarks/matbench_v0.1_Atomic_Affine_QRF/)results.json.gz,info.json,notebook.ipynb: required files. The notebook rebuilds all forests from raw inputs and asserts the fold dictionaries equalresults.json.gzexactly.run.py,inputs/base_qrf.py,protocol.json,inputs/protocol.json: fitting source and frozen settings.inputs/matbench_steels.json.gz,inputs/steels_splits.json: snapshots of the official dataset and folds (~50 KB).fresh.py,export.py,reproduce_entry.py: command-line rebuild/export (python -B reproduce_entry.py --output fresh_script_review --workers 3).preflight.py,environment_lock.json,manifest.json,upstream_source_hashes.json: refuse to run on mismatched package versions, changed inputs or existing fit caches. Exact replay needs CPython 3.13.16 withrequirements.txt.README.md,all_seed_scores.csv: method summary and all seed scores.Please add the
new_benchmarklabel (I cannot set labels on this repository).Prepared with AI assistance; reviewed and tested by me.