Skip to content

fix(sparse_probing): add std_floor to standardization - #1827

Merged
jlarson4 merged 4 commits into
TransformerLensOrg:devfrom
VictorAraujopy:fix/sparse-probe-std-floor
Sep 28, 2026
Merged

jlarson4 merged 4 commits into
TransformerLensOrg:devfrom
VictorAraujopy:fix/sparse-probe-std-floor

Conversation

@VictorAraujopy

@VictorAraujopy VictorAraujopy commented Sep 26, 2026 •

Copy link
Copy Markdown

Fixes #1816

Under preprocess="standardize", only an exactly-zero std was guarded, so a near-constant column was divided by its own tiny std and reached the fit as noise at unit scale. This adds a std_floor parameter to fit_sparse_probe and sweep_sparse_probe (default 1e-3, matching the reference): each non-constant column's scale is now max(std, std_floor).

  • Zero-variance columns keep scale one, as the guide already documents. Their coefficient is always zero, so the scale never reaches a prediction; scale one just avoids 0/0 when std_floor=0. This is the one remaining divergence from the reference, and it's noted in the guide's Reference section.
  • std_floor=0 disables the floor and reproduces the previous behavior.
  • std_floor is validated (finite, non-negative, not bool) and recorded on SparseProbeResult.
  • The max_iter docstrings and the guide now state the LBFGS evaluation budget, max_iter * 5 // 4 (equal to max_iter when max_iter <= 3); the line search can go one evaluation past it.

Note: with the new default, preprocess="standardize" results change for columns whose training std is below 1e-3.

make unit-test and uv run mypy . pass locally. Happy to change the constant-column handling if you'd prefer it to follow the reference exactly.

Type of change

  • Bug fix (non-breaking change which fixes an issue)
  • This change requires a documentation update

Checklist:

  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes
  • I have not rewritten tests relating to key interfaces which would affect backward compatibility

Near-constant columns were being divided by their own tiny std,
blowing noise up to unit scale. The scale now has a floor
(default 1e-3, same as the reference). Zero-variance columns
still get scale 1, and std_floor=0 keeps the old behavior.

Also document the max_eval budget in the max_iter docstrings.

@jlarson4 jlarson4 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Welcome to working on TransformerLens @VictorAraujopy! Thanks for putting this solution together. A couple small feedback items below before we get this merged

Comment thread transformer_lens/tools/analysis/sparse_probing.py Outdated
Comment thread transformer_lens/tools/analysis/sparse_probing.py
Comment thread docs/source/content/sparse_probing.md Outdated
Test that sweep_sparse_probe passes std_floor to its main and
control fits, fix the zero-variance guard comment, and note that
the line search can go one evaluation past the max_eval budget.
@VictorAraujopy

VictorAraujopy commented Sep 28, 2026 •

Copy link
Copy Markdown
Author

Thanks for the review @jlarson4 , all three addressed in e1e13e3 ;)
The sweep test records the std_floor reaching every _selected_data call, so it also covers the control fits.

@jlarson4

Copy link
Copy Markdown
Collaborator

Thank you for resolving those for me @VictorAraujopy! Looks great, approved and merging now.

@jlarson4
jlarson4 merged commit 0c24826 into TransformerLensOrg:dev Sep 28, 2026
27 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Proposal] Sparse probing: no standard-deviation floor under standardize, and max_eval is documented only as a stop reason

2 participants