Skip to content

Improve performance of aggregate_region(), filtering and index creation - #1000

Open
pmussak wants to merge 5 commits into
IAMconsortium:mainfrom
pmussak:perf/aggregate-region
Open

pmussak wants to merge 5 commits into
IAMconsortium:mainfrom
pmussak:perf/aggregate-region

Conversation

@pmussak

@pmussak pmussak commented Oct 1, 2026

Copy link
Copy Markdown

Please confirm that this PR has done the following:

  • Tests Added
  • Documentation Added
  • Name of contributors Added to AUTHORS.rst
  • Description in RELEASE_NOTES.md Added

Description of PR

Region processing in nomenclature calls aggregate_region() once per common region for every variable with weights or components. That is up to about 13,000 calls in the example below (601 variables × 22 common regions). Every call deep-copies (almost) the entire IamDataFrame, which leads to massive memory usage for dataframes with millions of data points.

This PR makes two changes:

  1. make_index() builds the index with MultiIndex.from_arrays() instead of creating a Python tuple per data row and then using MultiIndex.from_tuples(). from_arrays() takes column-wise data as input and we already hold the data like that. from_tuples takes row-wise input and then transforms row- to column-wise data under the hood. So using from_arrays() directly saves two conversions: arrays -> tuples -> arrays. The old code ensured that unused categories of categorical columns don't become index levels. To keep the behavior the same, remove_unused_levels() was added,. This affects IamDataFrame initialization, filter() and everything else that uses make_index(), and it mostly saves runtime.
  2. aggregate_region() selects rows with boolean masks instead of filter(). filter() deep-copies the IamDataFrame (data, meta, exclude), rebuilds the index and sorts the data, although the aggregation only needs the selected rows. Using filter() with inplace=True is not an option, since it would change the callers dataframe. This is the change that reduces the memory usage. components=True (auto-detection) still uses filter(), because the auto-detection relies on IamDataFrame._variable_components(). This might be changed in a dedicated follow-up PR.

Neither of these changes the time or space order of complexity (big O), but both reduce constant factors of said complexity.

Behavior changes

The results are unchanged: the region processing below returns identical output (pd.testing.assert_series_equal), and the test suite gives the same results as on main.

One minor difference: aggregate_region() with subregions that have no data no longer logs "Filtered IamDataFrame is empty!". That warning came from the internal filter() call. The result is the same (empty, or only the components at the region level). I think this is acceptable.

Tests

test_init_df_with_unused_categories: an IamDataFrame from a categorical column with unused categories (covers remove_unused_levels()).

Benchmark

All numbers are from a ~15MiB IAMC spreadsheet with 2,076,660 data points, 2,226 variables and 33 native regions. I processed it with the common-definitions variables and region mappings: 28 common regions (22 real aggregations and 6 renames), and 601 variables with weights or components. Python 3.12, pandas 2.3.3.

Region processing, with nomenclature v0.32.0 and the commits of this PR added one at a time. "Peak memory increase" is the maximum RSS during nomenclature RegionProcessor.apply() minus the RSS right before the call (after loading data and definitions), sampled every 20 ms.

pyam runtime peak memory increase
main 2,979 s +2,190 MiB
+ make_index() (change 1) 910 s* +1,842 MiB
+ boolean masks in aggregate_region() (change 2) = this PR 268 s +698 MiB

* ran in parallel with other benchmarks, so the runtime is overestimated.

The remaining memory is mostly used by nomenclature's own processing. I will open a PR with changes to nomenclature to reduce that as well.

Benchmark code

Region processing (each variant in a fresh process, on Linux):

import os, threading, time
from pathlib import Path

import pyam
from nomenclature import DataStructureDefinition, RegionProcessor

def rss_mib():
    with open("/proc/self/statm") as f:
        return int(f.read().split()[1]) * os.sysconf("SC_PAGE_SIZE") / 2**20

df = pyam.IamDataFrame(data)  # the submission, 2,076,660 data points
dsd = DataStructureDefinition(Path("common-definitions") / "definitions")
processor = RegionProcessor.from_directory(Path("common-definitions") / "mappings", dsd)

# sample the RSS every 20 ms while the region processing runs
before, peak, running = rss_mib(), 0.0, True

def sample():
    global peak
    while running:
        peak = max(peak, rss_mib())
        time.sleep(0.02)

thread = threading.Thread(target=sample)
thread.start()
start = time.perf_counter()
result = processor.apply(df)
runtime = time.perf_counter() - start
running = False
thread.join()

print(f"{runtime:.0f} s, peak memory increase +{peak - before:.0f} MiB")

make_index() microbenchmarks (runtime is the best of 3 runs after a warm-up):

import time, tracemalloc

df = pyam.IamDataFrame(series)
cases = {
    "IamDataFrame(series)": lambda: pyam.IamDataFrame(series),
    "filter(region=<all but one>)": lambda: df.filter(region=df.region[1:]),
    "filter(region=<one region>)": lambda: df.filter(region=df.region[0]),
}
for name, func in cases.items():
    func()  # warm-up
    runtimes = []
    for _ in range(3):
        start = time.perf_counter()
        func()
        runtimes.append(time.perf_counter() - start)
    tracemalloc.start()
    func()
    peak = tracemalloc.get_traced_memory()[1]
    tracemalloc.stop()
    print(f"{name}: {min(runtimes) * 1e3:.0f} ms, peak allocation {peak / 2**20:.0f} MiB")

`make_index()` created a Python tuple per data row, which dominated the
runtime of `filter()` (and every method using it) on large data.
`filter()` copies the full IamDataFrame, rebuilds the index and sorts the
data, which dominated the runtime when aggregating many variables one by
one. Use boolean masks on the data instead.
`MultiIndex.from_arrays()` keeps unused categories of categorical columns
as levels, so they showed up in `IamDataFrame.model` and similar.
@pmussak pmussak changed the title Perf/aggregate region Improve performance of aggregate_region(), filtering and index creation Oct 1, 2026
@pmussak
pmussak requested a review from danielhuppmann October 1, 2026 09:36
@codecov

codecov Bot commented Oct 1, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 94.7%. Comparing base (7dddc85) to head (e57de39).

Additional details and impacted files
@@          Coverage Diff          @@
##            main   #1000   +/-   ##
=====================================
  Coverage   94.7%   94.7%           
=====================================
  Files         70      70           
  Lines       6762    6771    +9     
=====================================
+ Hits        6406    6415    +9     
  Misses       356     356           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@pmussak

pmussak commented Oct 1, 2026

Copy link
Copy Markdown
Author

Some tests where failing with errors like these
FAILED tests/test_iiasa.py::test_conn_invalid_creds_file[creds0-InvalidCredentials-rejected] - httpx.ReadTimeout: The read operation timed out
this seemed like a temporary issue to me, so I did rerun failed job of the pytest and pytest-legacy actions. Now we are all green.

Are these intermittent test run failures something you have seen happen for pyam before @danielhuppmann ?

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant