Repository navigation
Compile uncached where numba can write no cache for numbox, with one warning naming the remedy - #115
nelson2005 wants to merge 69 commits into
Conversation
…t died at its first decorated function
numba raises at decoration when no cache location can be written: a read-only install whose user cache directory cannot be written either, or an import from an .egg, .whl or .pyz archive. Every module decorates under the one jit_options, so the question is put once, to a function of configurations.py compiled and saved at import, and answered for the package: cache off and one warning naming the remedy, NUMBA_CACHE_DIR for a source file on disk, an unpacked install or a .zip for an archive, NUMBOX_JIT_OPTIONS='{"cache": false}' to silence it. Four tests place the tree in an .egg, a .whl, a .zip with and without a writable home, and a read-only directory; a docs page covers the options and where the cache lands.
… kernel tests count what a child writes The compiled probe left an index and a data file in NUMBA_CACHE_DIR on every import with caching on, which test_compile_kernel.py's cache tests count as a kernel's. CompileResultCacheImpl runs the set-up decoration runs, picking the locator or raising, and ensure_cache_path the writability check the first save runs and numba skips for a .zip; nothing is compiled and nothing written but the cache directory.
…p; the archive text offered the .zip they had A .zip is the one archive numba finds a cache location for, the user's cache directory, and the OSError from the writability check names it. The warning for that case now says to make it writable; the unpacked-install-or-.zip text stays for the archives numba has no location for. The zip test with a read-only home asserts the new text from numba 0.61 on.
…piles uncached A probe that only loaded its own entry would have said the cache works and left the import to die at numba's first write for anything not there. The writability check numba runs for every other placement runs for the .zip too, so the second import warns once and works.
… was The fallback answers no-locator and a directory that cannot be written. A RuntimeError of another kind at the same step is re-raised, which no test showed: with the re-raise removed the file still passed.
…ry that stopped being writable falls back too
Valid JSON of another shape, null or a list, went to njit as keyword arguments and died there, or at the fallback's read of "cache", with a message naming neither the variable nor the shape it takes.
…make_vector's create is bare and uncacheable
…t the zip-cached test pins array_data_p is a module-level bare @njit, not a closure. The test that shows numba caching a .zip from 0.61 pins the remedy the warning names, not the fallback, and its comment now says so.
…yc-only install was offered NUMBA_CACHE_DIR numba finds a locator by the code's co_filename, which a sourceless install still names after the .py that is gone, while __file__ names the .pyc that is there. The remedy for that placement is the source files on disk, said in the archive branch's text, and a test compiles the tree to .pyc, removes the sources and reads the warning.
…ere its directory cannot be written make_structref and the sqlite aggregate, window and table-valued function registrations wrote their generated code's anchor under NUMBA_CACHE_DIR or the user's cache directory whatever the cache option said, and died at the write where that directory cannot be written, which the package's own probe does not see when its functions cache beside their sources. The anchor is now written when the options say cache, its directory checked as numba checks a cache directory, and where either fails the code compiles without a cache after one warning naming NUMBA_CACHE_DIR. Three subprocess tests, one per API, run against a read-only user cache directory and again cured by NUMBA_CACHE_DIR.
…tree cache is a __pycache__ beside each source One directory can be writable while another is not, and the check on configurations.py's directory alone passed where libm's died at its first binding. A probe of the same code is given each directory's first module as its file, which is all a locator reads of it; an archive shows no directories to walk and a .pyc-only install no modules, so there the probe as it is answers. Pinned by a tree with one read-only directory among writable ones and a read-only home, cured by NUMBA_CACHE_DIR.
…and for each anchor check_cache_location(py_file) puts the question the way numba does, the set-up decoration runs and the writability check the save runs, for a probe given that file, and the package asks it for a file in each of its directories. The anchors ask it too, in place of a write to the anchor's own directory: numba caches the code in the user's cache directory when that directory cannot be written, which the write refused, and a warm anchor there with no user cache left the first decorated function to die, which the write let through. Pinned by a warm anchor in a NUMBA_CACHE_DIR that stopped being writable, first with the user's cache directory writable and then without, for each of the three APIs. The docstring no longer says numba skips the check at save time for a .zip; it skips it at decoration.
… looks up A walk that named directories by their .py files skipped one left .pyc-only beside sourced ones, and an import from it died at its first cached function. Pinned by a tree whose sqlite directory is compiled and stripped of its sources.
…ched from It said the generated code is written to a content-addressed file without condition; with caching off nothing is written and the path is the code's filename.
… directory numba finds a location by each module's own source, so a module left as .pyc alone beside sourced ones, libm beside a kept __init__.py, failed where its directory's first module passed, and importing it died at its first cached function. One module still stands for each directory's location; every module without its .py answers for itself. Pinned by a tree with libm.py compiled and removed.
…uestion Both wrote their anchor and caught an OSError from the write alone, so a warm anchor in a directory that stopped being writable, with no user cache to fall back to, reached numba's set-up and died there with no locator. compile_kernel now warns and compiles uncached through _anchored_or_uncached, with its text, and the builder's derive returns None as it did for an anchor it could not write. The warm-anchor test runs both; the kernel test that read the old text reads the new.
…oo long was told to set NUMBA_CACHE_DIR Permission denied, operation not permitted and a read-only file system are the cache's; any other OSError at the anchor write, a name too long for the file system from a 300-character struct name, is raised as it was.
…fall back, in both docs pages
… is the location's, a name too long is numbox's own Narrowing the cache's errors to three errnos took away the fallback compile_kernel and the builder's derive had for any other NUMBA_CACHE_DIR numba cannot use, a path into a file or a component too long, which numba itself passes over on any OSError. _anchor_or_error makes the directory, writes the anchor and asks numba, in three steps: any OSError at the directory is the location's, an ENAMETOOLONG at the file is the name numbox made and is raised, and the answer's errors are numba's. is_a_cache_error counts every OSError again, the no-locator RuntimeError, and the ValueError numba raises for an archive under a directory whose name holds .zip, which its locator takes by substring. Pinned by the five APIs under a NUMBA_CACHE_DIR that is a file and one with a 300-character component, and an .egg under my.zip.dir.
…en the path is too long Re-raising ENAMETOOLONG at the file took the base's fallback from compile_kernel and the builder's derive under a NUMBA_CACHE_DIR deep enough that their fixed-length anchor names overflow, the location's fault. Any OSError from the directory or the file means no cache here; the warning names the length when that is the cause, NUMBA_CACHE_DIR or the struct's or function's name being the long part, and offers a writable directory otherwise. The docs say a directory of numbox that numba cannot cache in turns caching off whether or not its modules cache anything, the answer erring toward uncached.
…nce, and name two edges the check does not reach The docstring said nothing is written where numba cannot cache; the anchor is written first, since the question needs the file, and left where it is. It promised a warning wherever caching was asked for; a method without a canonical fingerprint turns caching off without one, as it always did. The derive's cases now assert no warning. The docs name numba's own cache file names, which repeat a struct's name and overflow for one around a hundred characters, and a per-directory location of a .zip made unwritable on its own inside the user's cache directory, as edges outside the question.
…ut 93 characters overflowed them The anchor's stem carried the struct's name and numba's index file name carries the stem and the qualified name of the jitted function, which carried the struct's again through the class, so a struct named with about 93 characters died in numba's own files past the anchor's check, on the branch and on main alike. bounded_stem keeps a name of 40 characters or fewer as it is and makes a longer one its first 31 characters and a digest of the whole; the anchor stems, the generated make_ and ol_ names, the method thunks and the class itself take it, and the class takes the full name back once compiled. Pinned by structs of 40, 41, 150 and 300 characters with a method, cached and loaded again; the length warning is pinned by a NUMBA_CACHE_DIR deep enough for a fixed-length anchor name to overflow, with the remedy asserted.
…er's cache directory for a frozen application numba caches a .zip per directory of the archive, each in a location of its own under the user's cache directory, and the walk saw none of them, so one such location made unwritable passed the check; the archive's directories are listed from the .zip now. os.walk did not follow a symlinked directory of the package, so a read-only one behind the link passed too. A frozen application is cached in the user's cache directory without its sources on disk, and where that directory cannot be written the warning offered the sources or a .zip; it names the directory now, as for a .zip. Each pinned by a child in that placement.
…ba 0.60 never makes
…a name of 40 accented characters overflowed numba's cache file names
…asured for the longest, a method thunk's, and test it
…ield of 250 characters overflowed numba's cache file names
…ce; NUMBOX_JIT_OPTIONS reaches only the package's, and the bounded names leave the directory as the long part
…d the methods' names; a field named with it lost its property
…nning it; a member it could not decompress raised past the exceptions listed
… say a symlinked directory is walked wherever it points
…ave the walk an exponential number of paths and the import did not finish
… remedy; the warning named the user's cache directory where a location under it had lost writability
…_kernel's overrides the jit options and the variable alike
…y, as a .py member does; a directory of such members was never asked for
…e system to put it at a shorter path; the warning said to make it writable
…ptions the anchors' under them
…oved .zip whose .pyc members name its old path passed the check and died there
…where numba can write no cache for its own files
…lies under it; a location that cannot be made under a dangling link is ENOENT too
…here numba can write no cache for numbox's own files
…eft the thunk's bound unpinned
…e path limit passes the package check and overflows at numba's first save
…the source from it; a typing error in a struct method pointed at a file that was not there
MiniMax Code ReviewThis is a substantial pull request that adds comprehensive handling for numba caching scenarios in the numbox library. Let me analyze the changes:Summary of ChangesThe main goal appears to be handling cases where numba cannot write its cache (e.g., read-only installs, archive imports like .egg/.whl/.pyz, zip imports). Previously, the import would fail with cryptic errors. Now it gracefully falls back to non-cached compilation with warnings. Key Components:
I'll focus on the testing and documentation updates. The modifications include new test files to validate cache handling in various scenarios like invalid JIT options and different caching locations. Documentation has been expanded to explain the caching mechanism comprehensively, with updates across multiple RST files to provide clear guidance on how caching works in numbox. The testing approach looks thorough, covering edge cases such as:
Documentation changes appear equally meticulous, ensuring users understand the new caching behavior and its implications for different installation environments. The implementation seems well-thought-out, with robust error handling and flexible cache management strategies. The approach of providing a single warning and defaulting to non-cached compilation appears pragmatic. Key improvements include handling complex scenarios like read-only installations, archive imports, and potential cache directory issues. The fallback mechanisms seem comprehensive, addressing various potential failure modes in caching strategies. The code demonstrates careful consideration of different edge cases, particularly around file system interactions and compilation constraints. Validation of environment variables appears thorough, preventing potential configuration errors. Implementation details show nuanced handling of different import scenarios, ensuring graceful degradation when standard caching mechanisms fail. The approach provides a robust solution that maintains functionality across diverse system configurations. Code Review: Numba Caching Fallback PRThis is a well-designed and comprehensive implementation addressing a real pain point: numbox imports failing in various read-only or archive-based deployment scenarios. Here's my review: Strengths1. Robust Cache Detection StrategyThe approach of probing each directory of the package is sound. The code correctly:
2. Comprehensive Error HandlingThe def is_a_cache_error(error):
if isinstance(error, OSError):
return True
if isinstance(error, ValueError):
return "No zip file found" in str(error)
return isinstance(error, RuntimeError) and "no locator available" in str(error)3. Good API DesignThe
4. Proper Warning HierarchyThe warning messages include specific remedies based on the failure mode (NUMBA_CACHE_DIR, path length, source on disk vs archive, etc.). Minor Suggestions1. Test Coverage for Edge CasesThe tests cover many scenarios, but consider adding:
2. Documentation ClarityThe new documentation in # If you see a RuntimeWarning about caching, set:
export NUMBOX_JIT_OPTIONS='{"cache": false}'3. Error Message ConsistencyIn Potential Concerns1. Performance on Large PackagesThe 2. Windows Path HandlingThe code handles Windows-specific cases (e.g., 3. Race ConditionsThe checks happen at import time, but the cache directory could become unwritable later (e.g., disk full). The anchor-level checks in ConclusionThis is a solid, production-ready implementation. The code is well-documented, thoroughly tested, and handles the edge cases gracefully. The fallback to non-cached compilation is the right behavior - it's slower but correct, with clear warnings explaining the situation and remedies. Recommendation: Approve with confidence. The changes are backward-compatible (existing code that worked continues to work), and the new behavior for problematic environments is an improvement. |
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
It changes cross-platform cache discovery and dynamic code generation using private Numba internals, requiring final human validation.
Review effort: Balanced
Findings: 1
Open (1)
What changed in this PR
Adds graceful uncached compilation when Numba cannot write cache files, including actionable warnings and bounded generated names.
Changes:
- Probes package and generated-code cache locations before enabling caching.
- Bounds generated StructRef names to avoid filesystem limits.
- Adds extensive archive, permission, path-length, and configuration tests and documentation.
| File | Description |
|---|---|
numbox/core/configurations.py |
Adds cache probing, fallback, warnings, and option validation. |
numbox/utils/preprocessing.py |
Adds anchor validation and bounded stems. |
numbox/utils/highlevel.py |
Applies fallback and bounded StructRef names. |
numbox/core/work/builder.py |
Guards generated work caches. |
numbox/core/variable/compile_kernel.py |
Gracefully disables unusable kernel caches. |
numbox/core/bindings/sqlite/udf_helpers.py |
Guards generated UDF caches. |
numbox/core/bindings/sqlite/tvf.py |
Guards generated TVF caches. |
test/core/test_no_cache_location.py |
Covers cache-location and generated-code edge cases. |
test/core/test_jit_options.py |
Tests environment-option validation. |
test/core/test_compile_kernel.py |
Updates fallback warning expectation. |
docs/numbox.core.configurations.rst |
Documents cache configuration and remedies. |
docs/numbox.core.variable.rst |
Documents kernel cache fallback. |
docs/numbox.utils.rst |
Documents anchors and bounded names. |
docs/modules.rst |
Adds configuration documentation to the index. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| A content-addressed file under numba's cache directory, not | ||
| ``highlevel.py``, is the ``compile()`` anchor of the generated | ||
| ``code_txt``. With caching off nothing is written and the path | ||
| serves as the code's filename. With caching on the anchor is written | ||
| and numba asked whether it can cache a function of it, which needs |
…e directory is under Library/Caches there and the shell's known folder here, a too-long component is a syntax error on Windows, and %r doubles its backslashes
… it can be, caching off or on
…mit whatever tmp_path's length; from a base directory of 122 bytes it ended past the limit and its mkdir died before the test ran
…es; a location within their length of the path limit passed numba's temporary file and the import died at the first save numba's writability check makes a temporary file, one without a name on Linux, and the files it saves are named after the module's stem, the function's qualified name, a line number, the interpreter tag and an index number, under a 21-byte temporary name: 105 bytes for the longest-named function of the package, 117 for the builder's generated kernel. A NUMBA_CACHE_DIR whose location for a module came within that of the path limit passed the probe, as it passes numba, and the import died at numba's first save with File name too long. The probe makes and removes a file of 128 bytes in each location it asks about, the package's bound on numba's names, which a test holds every function of the package under, and the warning's remedy for a source on disk names the length: a shorter NUMBA_CACHE_DIR, or none, since each location numba picks for a source on disk but the one beside it appends the source's directory path, else the package at a shorter path. The anchors keep numba's own check, their names being the struct's; a NUMBA_CACHE_DIR deep enough to overflow an anchor's name is too deep for the package's files first, so the package answers, which the anchors' path-limit test pins now, its home case as before.
|
|
||
| A ``.zip`` whose cache directory holds every entry but can no longer be written takes the fallback too, where | ||
| numba alone would have loaded the entries: the writability check is the rule numba applies to every other | ||
| placement, and the one the ``.zip`` locator is missing. |
There was a problem hiding this comment.
Do you want to add a link to the issue?
…mbox builds its other generated functions, rather than replacing a function's co_filename; numba's locators read the same file, cache path, source stamp and errors either way, numba 0.60 to 0.67
…bility check numba's .zip locator is missing

import numboxdies at its first decorated function where numba can write no cache for it:RuntimeError: cannot cache function 'cos': no locator available for file '.../numbox-0.0.0-py3-none-any.whl/numbox/core/bindings/libm.py'. numba sets a cached function up at decoration, and every module here decorates underjit_options,{"cache": True}unlessNUMBOX_JIT_OPTIONSsays otherwise. Two placements have no cache location: an import from an.egg,.whlor.pyzarchive, which Spark's--py-filesships, and a read-only install whose user cache directory cannot be written either. A.zipis given the user's cache directory from numba 0.61 on without a check that it can be written, and dies withOSErrorat the first save instead. Nothing in the errors named the way out.uncached_where_no_cache_can_be_writtenanswers the question once for the package whenconfigurationsis imported, the way numba asks it at decoration plus the writability check of the first save, which numba skips for a.zip, made with a file named as long as numba's longest for the package, since numba's own temporary file is short and a location within their length of the path limit passed it and died at the first save; nothing is compiled, and nothing written but the cache directory, which numba makes at the first decoration anyway. numba's in-tree cache is a__pycache__beside each source, so one module of each directory of the package is asked, each real directory once; a module that survives as.pycalone, on disk or in a.zip, is asked by the file its code was compiled from, which is what numba looks up, and a.pyczipimport would not run is passed over as zipimport passes it. Where any answer is no,jit_optionscomes back withcacheoff and oneRuntimeWarningnames the remedy:NUMBA_CACHE_DIRfor a source file on disk; the user's cache directory made writable, or put at a shorter path, for a.zipor a frozen application, which numba caches there whatever the variable says; the sources on disk, or a.zipholding them, for any other archive or a.pycwithout its source; andNUMBOX_JIT_OPTIONS='{"cache": false}'turns caching off and silences the warning. An error that is not the cache's is raised as it was.One probe rather than a guard in every decorator, which is what Goykhman/numbarrow#12 did for numbarrow's three: numbox has 105
njit(sites and about two hundred decorated functions, and a placement that gives numba no cache for one of them gives it none for the rest. The answer errs toward uncached, which is never wrong: a directory numba cannot cache in turns caching off for the package even where every cached function's own directory is fine. The probe compiles nothing, so an import costs what it did, measured within noise of main, cold and warm.The code numbox generates at run time,
make_structref's,compile_kernel's, the work builder's derives and the sqlite registrations', is anchored to a file underNUMBA_CACHE_DIRor the user's cache directory, which can be unwritable where the package's own files cache beside their sources. The anchor is written whenever it can be, as before, since numba quotes the source from it in its messages, and each anchor now puts the same question for its own file when caching (_anchored_or_uncached): the code it names compiles without a cache after a warning of the same shape where the answer is no, instead of dying at the write or at numba's set-up; the builder's derive falls back without a warning, as it did. Those three take jit options of the caller's, whichNUMBOX_JIT_OPTIONSdoes not reach, so that warning's silence iscacheoff in the options the code was given,compile_kernel'scacheargument where it has one, or the variable where the options are the package's;make_graph's kernel, anchored to the builder's own file, puts the package's question for that file under the options it was given, since a caller'scachereached numba past the package's answer and died from an archive. A path too long for the file system is one such failure, and the warning says so, offeringNUMBA_CACHE_DIRat a shorter path.numba names its cache files after the anchor's stem and the generated function's qualified name, and both carried the struct's name, so a struct named with about 93 characters died with
OSError: File name too longin numba's own files, on main as on this branch.bounded_stemkeeps a name of 40 bytes or fewer as it is, so nearly every struct keeps the file names it had, and cuts a longer one to the whole characters within 31 bytes and a digest of the whole;make_structrefdefines the generated class, the field getters, the method thunks and themake_andol_functions under bounded names, and the class takes the struct's full__name__and__qualname__back once its body is compiled. The longest file numba writes for any struct stays under 230 bytes.Sixty tests place the tree in each archive and each tree where numba can cache nothing, or not everything, import it in a child, and assert exactly one warning with the remedy for the placement, or none, and that a lazily compiled helper, an eagerly compiled one and a proxied binding work uncached; the generated code's tests run each anchor writer where its anchor cannot be written, and build structs named with 40 to 300 ASCII, accented and CJK characters that cache and load again in a second process. The tests that take write access from a directory skip on Windows and as root. On main, 49 of the 60 fail with the errors above; the eleven that pass are controls.
A docs page covers
NUMBOX_JIT_OPTIONS, which nothing documented before, where the cache lands, the fallback and each remedy; the cache-anchor section ofnumbox.utilscovers the anchors' fallback and the bounded names.