Repository navigation
Sync upstream/main: compile uncached where numba can write no cache (f8a6fe6) - #116
Conversation
…t died at its first decorated function
numba sets a cached function up when it is decorated and raises there when no cache location can be written: an import from an .egg, .whl or .pyz archive, or a read-only install whose user cache directory cannot be written either. A .zip is given the user's cache directory from numba 0.61 on without a check that it can be written, and died with OSError at the first save instead. Every module decorates under the one jit_options, so the question is put once, when configurations is imported, and answered for the package: a function whose file is one of the package's is given to CompileResultCacheImpl, the way numba asks at decoration, then its source stamp read and its cache path ensured, the writability check of the first save; nothing is compiled. One module of each directory of the package is asked, since numba's in-tree cache is a __pycache__ beside each source, found from this module's __file__, each real directory once; a module that survives as .pyc alone, or a .pyc member of a .zip that zipimport would run, is asked by the file its code was compiled from, which is what numba looks up. Where any answer is no, jit_options comes back with cache off and one RuntimeWarning names the remedy for the placement: NUMBA_CACHE_DIR for a source file on disk; the user's cache directory made writable, or put at a shorter path, for a .zip or a frozen application, which numba caches there whatever the variable says; the source files on disk, or a .zip holding them, for any other archive or a .pyc without its source; and NUMBOX_JIT_OPTIONS='{"cache": false}' to turn caching off and silence it. An error that is not the cache's is raised as it was. A NUMBOX_JIT_OPTIONS value that is not a JSON object, or whose cache is not true or false, is refused by name; one without a cache key leaves njit to its default, off, but the sqlite callbacks cache under it, so the question is put for those too. Thirty tests place the tree in each archive and each tree where numba can cache nothing, or not everything, and import it in a child; a docs page covers the options, where the cache lands, the fallback and each remedy.
…ut 93 characters overflowed them numba names its cache files after the anchor's stem and the generated function's qualified name, and both carried the struct's name, so a struct named with about 93 characters died with OSError: File name too long in numba's own files, past the anchor's write. bounded_stem keeps a name of 40 bytes or fewer as it is, so nearly every struct keeps the file names it had, and cuts a longer one to the whole characters within 31 bytes and a digest of the whole; the measure is the name's UTF-8, which is the file system's, so a name of 40 accented or CJK characters is bounded too. make_structref defines the generated class, the field getters, the method thunks and the make_ and ol_ functions under bounded names, a long field's getter handed to the field's name and kept clear of the other fields' and the methods', and the class takes the struct's full __name__ and __qualname__ back once its body is compiled. The longest file numba writes for any struct, a method thunk's, stays under 230 bytes, with numba's temporary name at the write 21 bytes longer. Six tests build structs named with 40 to 300 ASCII, accented and CJK characters, each with a field of the struct's length, a field named with that one's bounded name, a method of the bounded length and one of 200 characters, cache them and load them again in a second process.
…the generated code uncached where there is none; it died at the write or at numba's set-up The code numbox generates at run time, make_structref's, compile_kernel's, the work builder's derives and the sqlite registrations', is anchored to a file under NUMBA_CACHE_DIR or the user's cache directory, which can be unwritable where the package's own files cache beside their sources; make_graph's kernel is anchored to the builder's own file and cached beside it. The anchor is written whenever it can be, as before, since numba quotes the source from it in its messages, and each anchor now puts the package's question for its own file when caching (_anchored_or_uncached), so the code it names compiles without a cache after a warning of the same shape where the answer is no, instead of dying at the write or at numba's set-up; the builder's derive falls back without a warning, as it did, and an uncached kernel has no anchor, as before. make_structref, compile_kernel and the builder take jit options of the caller's, which NUMBOX_JIT_OPTIONS does not reach, and compile_kernel's cache argument overrides those too, so the warning's silence is cache off in the options the code was given, the argument where it takes one, or the variable where the options are the package's; make_graph puts the package's question for the builder's file under those options, for its kernel alone, since that kernel took a caller's cache to numba past the package's answer and died from an archive. A path too long for the file system is one such failure, and the warning says so, offering NUMBA_CACHE_DIR at a shorter path, which moves the anchor out of a long user cache directory too, the names being bounded; a directory within about 230 bytes of the path limit passes the check and overflows at numba's first save instead, which the docs say. Twenty-four tests run each anchor writer with its cache directory unwritable, a file, or a component too long for a path, warm in a directory that stopped being writable, and under a NUMBA_CACHE_DIR or a home too deep for the anchor's name; make_graph under a caller's cache from an archive; the warning for options the caller gave, compile_kernel's cache argument among them; and a typing error quoting its line with caching off.
…mit whatever tmp_path's length; from a base directory of 122 bytes it ended past the limit and its mkdir died before the test ran
…es; a location within their length of the path limit passed numba's temporary file and the import died at the first save numba's writability check makes a temporary file, one without a name on Linux, and the files it saves are named after the module's stem, the function's qualified name, a line number, the interpreter tag and an index number, under a 21-byte temporary name: 105 bytes for the longest-named function of the package, 117 for the builder's generated kernel. A NUMBA_CACHE_DIR whose location for a module came within that of the path limit passed the probe, as it passes numba, and the import died at numba's first save with File name too long. The probe makes and removes a file of 128 bytes in each location it asks about, the package's bound on numba's names, which a test holds every function of the package under, and the warning's remedy for a source on disk names the length: a shorter NUMBA_CACHE_DIR, or none, since each location numba picks for a source on disk but the one beside it appends the source's directory path, else the package at a shorter path. The anchors keep numba's own check, their names being the struct's; a NUMBA_CACHE_DIR deep enough to overflow an anchor's name is too deep for the package's files first, so the package answers, which the anchors' path-limit test pins now, its home case as before.
…mbox builds its other generated functions, rather than replacing a function's co_filename; numba's locators read the same file, cache path, source stamp and errors either way, numba 0.60 to 0.67
…bility check numba's .zip locator is missing
…cache-can-be-written Compile uncached where numba can write no cache for numbox, with one warning naming the remedy
MiniMax Code ReviewThe user wants me to review a pull request. Let me analyze the changes carefully.This is a significant PR that adds functionality for handling numba caching in various edge cases. Let me understand what's being done:
Let me provide a code review: Strengths:
Potential Issues:
The implementation appears sound overall—the test suite seems thorough and the approach is reasonable. Code Review: Numba Caching Support PRThis is a substantial PR that adds robust handling for numba caching in various edge cases (read-only installs, zip archives, path length limits, etc.). Here's my analysis: SummaryThe PR adds the ability for numbox to gracefully degrade when numba cannot write a cache, instead of crashing at the first decorated function. It also bounds long names to prevent filesystem limits. Strengths
Minor Issues/Suggestions1. Potential encoding issue in
|
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The synchronized upstream change is coherent, thoroughly tested, documented, and introduces no unresolved correctness issues.
Review effort: Balanced
Findings: None
What changed in this PR
Synchronizes the fork with upstream’s graceful fallback when Numba caches are unavailable.
Changes:
- Detects unusable cache locations and compiles uncached with actionable warnings.
- Safely bounds generated cache filenames.
- Adds extensive tests and cache-behavior documentation.
| File | Description |
|---|---|
numbox/core/configurations.py |
Adds cache probing, fallback, and option validation. |
numbox/utils/preprocessing.py |
Adds anchor fallback and bounded names. |
numbox/utils/highlevel.py |
Applies safe naming and fallback to structrefs. |
numbox/core/work/builder.py |
Protects generated graph and derive caching. |
numbox/core/variable/compile_kernel.py |
Uses shared cache fallback. |
numbox/core/bindings/sqlite/udf_helpers.py |
Protects generated UDF callbacks. |
numbox/core/bindings/sqlite/tvf.py |
Protects generated TVF callbacks. |
test/core/test_no_cache_location.py |
Covers cache-location and filename edge cases. |
test/core/test_jit_options.py |
Tests environment-option validation. |
test/core/test_compile_kernel.py |
Updates fallback warning assertion. |
docs/numbox.core.configurations.rst |
Documents cache configuration and remedies. |
docs/numbox.core.variable.rst |
Documents kernel fallback behavior. |
docs/numbox.utils.rst |
Documents anchors and bounded names. |
docs/modules.rst |
Adds configuration documentation to the index. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Merges upstream's Goykhman@f8a6fe6 into fork main. It carries the seven commits of Goykhman#42: where numba can write no cache (an archive import, a read-only install), numbox compiles uncached after one warning naming the remedy instead of dying at import.
Clean merge, no conflicts. Outside the fork-only files the merged tree is upstream's, and it is identical to the head of #115.