Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions docs/modules.rst
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@ numbox
numbox.core.any
numbox.core.bindings
numbox.core.bindings.sqlite
numbox.core.configurations
numbox.core.proxy
numbox.core.variable
numbox.core.vector
Expand Down
92 changes: 92 additions & 0 deletions docs/numbox.core.configurations.rst
Original file line number Diff line number Diff line change
@@ -0,0 +1,92 @@
numbox.core.configurations
==========================

Overview
++++++++

Every function numbox caches is decorated under one set of numba options, ``jit_options``, read once from
the ``NUMBOX_JIT_OPTIONS`` environment variable when this module is first imported; a bare ``@njit`` in
numbox, as in ``lowlevel.py``, ``meminfo.py`` and ``make_vector``, is not cached and takes none. The value is a JSON
object passed to ``@njit`` as keyword arguments, and any other shape is refused by name, as is a ``cache``
that is not ``true`` or ``false`` (the string ``"false"`` is true to numba, which reads the option's truth);
unset means ``{"cache": true}``, so numbox compiles into numba's on-disk cache by default, and
``export NUMBOX_JIT_OPTIONS='{"cache": false}'`` turns that off.

Where the cache lands
+++++++++++++++++++++

numba writes a cached function's entries under ``NUMBA_CACHE_DIR`` when that is set, else into the
``__pycache__`` directory beside the function's source file, else into the user's cache directory
(``~/.cache/numba`` on Linux), taking the first of those it can write. It sets the cache up when the function is
decorated, so with caching on the question of where it lands is settled at import.

Where no cache can be written, numba raises at decoration, and numbox's import used to die on its first module
with ``RuntimeError: cannot cache function ...: no locator available for file ...``. The two placements that do
that are a read-only install whose user cache directory cannot be written either, and an import from an
``.egg``, ``.whl`` or ``.pyz`` archive, which Spark's ``--py-files`` ships. Since every function numbox caches
decorates under the one ``jit_options``, numbox puts the question once, when this module is imported, and
answers it for the package: for a probe in each directory of the package that holds a module, since numba's
in-tree cache is a ``__pycache__`` beside each source, it runs the cache set-up numba runs at decoration and
the writability check numba runs at the first save, compiling nothing, and where either fails for any
directory numbox compiles without a cache and one ``RuntimeWarning`` names the remedy. Every directory
counts, whether or not its modules cache anything, so the answer errs toward uncached, which is never wrong.
A module that survives as ``.pyc`` alone is asked by the file it was compiled from, which its code keeps and
numba looks up: the ``.py`` that is gone where it was compiled in place, or a tree elsewhere, on disk or not.
For a ``.zip``, which numba caches per directory of the archive, each in a location of its own under the
user's cache directory, the archive's directories are listed and the question put for one module of each; a
``.pyc`` in the archive that zipimport would run, which it takes before the ``.py`` beside it unless it is
stale against it or of another interpreter, and whose code keeps the file it was compiled from, is asked by
that file, since that is what numba looks up for it, on disk or gone, and one zipimport would pass over is
passed over; any other archive has no location at all. A directory of the package reached through a symlink
is walked like the rest, wherever the link points. The check makes the cache directories it asks about, as
numba would at the first decoration in each; with caching beside the sources that is an empty ``__pycache__``
per directory of the package, a linked one included. numba's own writability check makes a temporary file,
one without a name on Linux, and the files it saves have names of a hundred bytes and more, so a location
within their length of the path limit, 4096 on Linux, passes numba's check and the first save overflows;
the check here makes a file named as long as the longest numba writes for the package's files (128 bytes,
which a test holds every function of the package under), so that location turns caching off with the
warning instead.

- For a source file on disk the remedy is ``NUMBA_CACHE_DIR`` pointed at a writable directory; where its
location is too long for the file system, a shorter ``NUMBA_CACHE_DIR`` or none, since each location numba
picks for a source on disk but the one beside it appends the source's directory path, else the package
installed at a shorter path.
- For a ``.zip``, or a frozen application, it is the user's cache directory made writable: a ``.zip`` is the
one archive numba caches, from 0.61 on, and it caches it there, taking the directory without checking that
it can be written; a frozen application (``sys.frozen``) is cached there too, its sources not being on disk.
numba reads ``NUMBA_CACHE_DIR`` only for a source file on disk, so the variable changes nothing for either.
A ``.zip`` whose cache directory holds every entry but can no longer be written falls back too, where numba
alone would have loaded the entries: the writability check is the rule numba applies to every other
placement.
- For an ``.egg``, ``.whl`` or ``.pyz``, a ``.pyc``-only install or a ``.pyc`` in a ``.zip``, it is the source
files on disk or a ``.zip`` holding them.
- ``NUMBOX_JIT_OPTIONS='{"cache": false}'`` turns caching off and silences the warning in every case, the
package's options being what it sets; the anchors' warning under a caller's own options, below, is
silenced by those.

A function numba cannot cache is compiled in every process that uses it, never wrong; that is the cost the
warning reports. An error at decoration that is not the cache's is raised as it was.

The code numbox generates at run time, ``make_structref``'s, ``compile_kernel``'s, the work builder's derives
and the sqlite registrations', is anchored to a file under ``NUMBA_CACHE_DIR`` or the user's cache directory
so that numba can cache it. That directory can be unwritable where the package's own functions cache, since
those cache beside their sources, so each anchor puts the same question for its own file when it is written,
and the code it names compiles without a cache where the answer is no, after a warning of the same shape (the
builder's derive falls back without one, as it did). ``make_structref``, ``compile_kernel`` and the builder
take jit options of the caller's, which the variable does not reach, and ``compile_kernel``'s ``cache``
argument overrides those too, so that warning's silence is ``cache`` off in the options the code was given,
the argument where it takes one, or the variable where the options are the package's. ``make_graph``'s
kernel is anchored to the builder's own file and cached beside it, so it puts the question for that file
under the options it was given, and falls back the same way with the package's remedy for the placement.
See the cache-anchor section of :doc:`numbox.utils`.

Modules
++++++++

numbox.core.configurations
--------------------------

.. automodule:: numbox.core.configurations
:members:
:show-inheritance:
:undoc-members:
7 changes: 5 additions & 2 deletions docs/numbox.core.variable.rst
Original file line number Diff line number Diff line change
Expand Up @@ -265,8 +265,11 @@ state cannot be fingerprinted -- a ``cres``-compiled callable, or a value with
no canonical form -- make that one kernel uncacheable: always recompiled per
process, never wrong. The
``cache`` keyword is tri-state: ``None`` (the default) defers to
``jit_options["cache"]``, then the ``NUMBOX_JIT_OPTIONS`` environment
default, then ``True``; an explicit ``True``/``False`` wins. Two costs are
``jit_options["cache"]``, then numbox's default, the ``NUMBOX_JIT_OPTIONS``
environment value or ``True`` unless numba can write no cache for numbox's
own files, where it is ``False`` (see :doc:`numbox.core.configurations`);
an explicit ``True``/``False`` wins, and the kernel's own anchor still turns
caching off where it cannot be cached. Two costs are
worth knowing: a formula that references or closes over a **large array**
pays a per-compile ``sha256`` over that array's bytes (proportional to its
size) on every ``compile_kernel`` call; and numba itself declines to
Expand Down
53 changes: 53 additions & 0 deletions docs/numbox.utils.rst
Original file line number Diff line number Diff line change
Expand Up @@ -162,6 +162,59 @@ the failure mode on that version to constants outside the inline
range. Earlier supported versions (3.10--3.13) collide on any
constant edit.

The anchor is written whenever it can be: numba caches from it, and
quotes the source from it in its messages, a typing error's among them,
so with caching off a write that fails is nothing and the path serves
as the code's name, as it did. With caching on numba is then asked
whether it can cache a function of that file, the question the package
puts for its own modules; where the write fails, for whatever reason,
or the answer is
no, an unwritable user cache directory or ``NUMBA_CACHE_DIR`` with no
other location left, the generated code compiles without a cache after
one warning naming the remedy, instead of dying at the write or at
numba's set-up. A path too long for the file system is one such
failure, and the warning says so where the file system does (Windows
reports a component too long as a syntax error, and the warning then
offers the directory, quoting the error); with the names bounded, as below,
the directory is the only part that can make it so, ``NUMBA_CACHE_DIR``
or the user's cache directory, and ``NUMBA_CACHE_DIR`` at a shorter
path moves the anchor out of either. The anchor's check writes names
shorter than the longest numba writes for a struct, which carry its name
and its methods', so a user cache directory within about 230 bytes of
the path limit, 4096 on Linux, passes it and overflows at numba's first
save instead, where numba's own error names the length; the remedy is
the same. A ``NUMBA_CACHE_DIR`` that deep is too deep for the package's
own files first, whose location under it appends their directory's path
and whose check reserves the length of numba's files, so the package
answers, with its remedy, and the generated code compiles under its
answer.

The anchor's name carries the struct's or the function's, and numba
names its cache files after the anchor and the qualified name of the
function it caches, which carries the struct's again through the class
whose body defines the jitted getters and method thunks, so a struct
named with about 93 characters overflowed the 255 bytes a file system
allows a name, in numba's own files past the anchor's check. Those
names are bounded now (``bounded_stem``): a name of 40 bytes or fewer
is used as it is, so nearly every struct keeps the file names it had,
and a longer one becomes the whole characters of it that fit in 31
bytes and a digest of the whole. The measure is the name's UTF-8,
which is the file system's: a character of another script takes up to
four bytes there. The generated class is defined under the bounded name and
takes the struct's full name back once its body is compiled, so
``__name__``, ``__qualname__`` and ``repr`` show the name the caller
gave, whatever its length, and the struct caches. A field's jitted
getter is defined under the bounded name the same way and the
property takes the field's, so a field's name of any length caches
too; a method's name is bounded in its thunk and its overload.
``compile_kernel``, the work builder's derives and the sqlite
aggregate, window and table-valued function registrations anchor their
generated code the same way and fall back the same way, the derive
without a warning, and the kernel and the derive write no anchor when
not caching, as before; ``make_graph``'s kernel, anchored to the builder's
own file, asks for that file under the options it was given. See :doc:`numbox.core.configurations` for the
package-wide rule the anchors follow.

See also ``numba.core.caching.Cache._index_key`` and
``numba.core.caching._SourceFileBackedLocatorMixin.get_source_stamp``
in numba's source for the cache key construction and source-stamp
Expand Down
6 changes: 3 additions & 3 deletions numbox/core/bindings/sqlite/tvf.py
Original file line number Diff line number Diff line change
Expand Up @@ -35,7 +35,7 @@
)
from numbox.utils.digest import digest
from numbox.utils.preprocessing import (
_anchor_path, _materialize_anchor, _orphan_anchor_sweep,
_anchor_path, _anchored_or_uncached, _orphan_anchor_sweep,
)

# Names referenced by the GENERATED source; importing them here puts them in
Expand Down Expand Up @@ -361,9 +361,9 @@ def _compile_xfilter(stem, arg_tags, out_dtype, fn):
src = _XFILTER_SRC.format(arg_decode=arg_decode, fn_call=fn_call)
tvf_digest = digest((out_dtype, tuple(arg_tags)), [fn])
code_txt = "# tvf-digest: %s\n%s" % (tvf_digest, src)
ns = {**globals(), "_fn": fn, "_N_HIDDEN": n_hidden}
anchor = _anchor_path(_ANCHOR_SUBDIR, stem, code_txt)
_materialize_anchor(anchor, code_txt)
ns = {**globals(), "_fn": fn, "_N_HIDDEN": n_hidden,
"jit_options": _anchored_or_uncached(anchor, code_txt, jit_options)}
code = compile(code_txt, str(anchor), mode="exec")
exec(code, ns) # nosec B102 - JIT codegen of internal source
return ns["_tvf_xfilter_impl"]
Expand Down
6 changes: 3 additions & 3 deletions numbox/core/bindings/sqlite/udf_helpers.py
Original file line number Diff line number Diff line change
Expand Up @@ -54,7 +54,7 @@
from numbox.utils.digest import digest
from numbox.utils.preprocessing import (
_anchor_path,
_materialize_anchor,
_anchored_or_uncached,
_orphan_anchor_sweep,
)

Expand Down Expand Up @@ -226,9 +226,9 @@ def _compile_callbacks(stem, srcs, state_type, fns):
code_txt = "# udaf-digest: %s\n%s" % (udaf_digest, "".join(srcs))
# globals() is this module's __dict__; it carries __name__, which numba's
# warm-cache Environment rebuild requires when reloading the cached impls.
ns = {**globals(), "_state_type": state_type, **fns}
anchor = _anchor_path(_ANCHOR_SUBDIR, stem, code_txt)
_materialize_anchor(anchor, code_txt)
ns = {**globals(), "_state_type": state_type, **fns,
"jit_options": _anchored_or_uncached(anchor, code_txt, jit_options)}
code = compile(code_txt, str(anchor), mode="exec")
exec(code, ns) # nosec B102 - JIT codegen of internal source
return ns
Expand Down
Loading
Loading