Repository navigation
Nest struct fields under their column, carry nulls out of a UDF, build output columns as declared, and give each viewer its own cache files - #11
Conversation
data_dict[col] and bitmap_dict[col] are now a dict of fields for a struct or list-of-struct column, so a field never shares a namespace with another column or with another struct's fields. Uniform columns are unchanged. The flat layout is what made a field collide with a same-named column at all. The issue 7 remediation turned that collision into an error; the nested layout removes the class outright, so the collision check goes with it. The zero-field struct refusal goes too: its reason was that such a column vanished from both dicts without a word, and nested it is present as an empty dict. The struct-level validity is still folded into each field's bitmap, so one is_null call per field sees both layers, exactly as before.
pa.array([]) infers null where the same column with rows infers string, so a UDF that filtered a whole batch away, which is the ordinary reason to use mapInArrow, yielded a null column next to string siblings and Spark's stream writer refused the second schema. A |U output is now built as string from its tolist(), which also keeps the NULs that C string semantics cut. A |S output took the plain pa.array path and got exactly that cut, so b"a\x00b" arrived as b"a". It now takes the same tolist() route as binary.
With output_schema every column was inferred first and cast afterwards. That is why a list of dicts keyed Amount/Label against a declared struct<amount, label> came back as two columns of nulls under an identical schema, since Arrow matches struct fields by exact name and nulls a missing one; why a map column could not be produced from any documented value, since cast(struct -> map) does not exist; and why a string column declared large_string went through an intermediate string. Each column is now built as its declared type from the start, a dict key that no declared field has raises, and a list of dicts or of pairs declared map becomes one. pa.array iterates a mapping, so a dict of arrays returned under one key became a string column of the dict's keys, with a different row count and nothing raised, output_schema included. It is refused naming the column. A numpy record array, the one ndarray shape that means struct and what an @njit function returns for a numba record type, died on "Unsupported numpy type 20" naming neither the column nor the dtype. It now becomes a struct column, one child per field, each child converted like a column of its own. A key the schema does not name was dropped silently; it raises. A result that is not a dict died on "'NoneType' has no attribute 'items'"; it raises a TypeError that says so. And every failure on the output side now names the column it happened in, chaining the original error. The output_schema docstring promised ArrowInvalid on any lossy conversion. Measured, pa.array refuses an integer out of range, a float with a fraction into an integer type and a timestamp unit change that drops digits, but a timestamp into date32 or date64 floors to the day and float64 into float32 overflows to inf, both silently. The docstring now says that.
…option
numba names a function's cache files after its qualname and source line.
The five viewers all come from one factory, so they shared one index file
and one set of data files, which numba writes without a lock: processes
importing together on a cold cache read one index and picked the same
data-file name for different viewers, and every process started afterwards
loaded the wrong machine code. Eight importers on an empty NUMBA_CACHE_DIR
poisoned it in five rounds out of five; the consumer then died on "In
'NRT_adapt_ndarray_to_python', 'descr' is NULL" or read an int32 column as
float64. A Spark executor starting its Python workers on a node with an
empty cache is exactly that shape, under the default options.
Each viewer now carries a qualname of its own, view_<dtype>, so it has its
own index and data files, and with one entry per index two writers have
nothing to disagree about. Five rounds of 8 importers and three of 24 pass
with five index files; the checked-in test runs one round of 8 and counts
the files. The subprocess tests count those files under a private cache
directory with caching on and off, which is also the first end-to-end test
that NUMBARROW_JIT_OPTIONS reaches the decorators at all.
"cache" must now be a JSON boolean: numba reads the string "false" as true,
so '{"cache": "false"}' left caching on. The docstring says that an explicit
object replaces the default, so '{}' turns caching off, and that numba's
index carries no compile flags, so a changed option needs a fresh cache
directory.
data.flags.writeable = True succeeded on a view over a locally built array, and a store through it then changed the source Arrow array; pyarrow's own to_numpy(zero_copy_only=True), which the README names as the equivalent, refuses the flip. The view is now built over a read-only memoryview of the buffer, so numpy refuses to set WRITEABLE on it, for a local array and an IPC one alike, and the flag no longer needs clearing afterwards.
…ypes Every refusal here was the right refusal with the wrong message. A Table column was refused as "an array of N elements of type int64", blaming a supported type and never saying chunked; it is now named as a ChunkedArray with the way out. A struct child the dispatcher could not adapt named its type but not which field it was, and through the factory no column at all; the struct adapter now prefixes the field and the factory the column, chaining the original error. Invalid UTF-8 escaped as a bare UnicodeDecodeError whose "position 0" was a byte offset inside the element; it is a ValueError naming the element. input_columns="value" was iterated character by character; a str is refused. A column name Spark's case-insensitive projection had rewritten died on KeyError: 'value' from the README's own example; the KeyError now lists the batch's columns. The two isinstance asserts guarding the struct and list adapters vanished under -O and let any array through to an unrelated error; they are TypeErrors. The adapter messages interpolate the pyarrow type, which grows with the schema: a thousand-field struct is 13,000 characters at the dispatcher and twice that in the list adapter, where the element type was interpolated a second time. type_repr cuts a type at 120 characters with a note of how many fields and characters were cut, and the list message no longer repeats the element type.
… views Spark binds a UDF's output columns by position and, measured on pyspark 3.5.7, checks nothing about their types: it reads each Arrow vector through the accessor its declared type expects, so an int64 column declared as a timestamp reads as one and any two columns sharing an accessor family swap silently; the docstring claimed it "compares only their types". A bitmap is None when the batch carries no validity buffer, not when there are no nulls: Spark's transport drops the buffer at zero nulls while slice, take, filter and fill_null keep an all-valid one, so an eager signature naming a bitmap array type compiles on one batch and raises on the next. The README named the string width as the one exception to declaring a signature; this is the second, and the repo's own Spark test already used Optional for it without saying why. The width exception itself fires on a narrower batch as well as a wider one, and the fixed-width result costs the widest value times the row count times four bytes, which the README now says with the measured figure. The timestamp handler reads the unit and never the zone, as pyarrow's to_numpy does, and nothing said so; on the way out a date64 comes back timestamp[ms] and a zoned timestamp naive unless output_schema says otherwise. The read-only paragraph cited to_numpy(zero_copy_only=True) for strings, which are copies and for which that call raises; it now says which results are views and which are copies, and that neither can be made writable. The MapArray sentence named raggedness as the only condition; a uniform map still raises for values of any type the table does not name. is_null and unpack_booleans document their index range and the boundscheck option that turns an overrun into IndexError, and NUMBA_DISABLE_JIT=1 is documented as unsupported, since the viewers' intrinsic has no pure-Python form.
pyspark 3.3 bundles cloudpickle 2.0.0, which predates the co_qualname argument Python 3.11 added to code(), so on the declared Python every UDF dies in the worker with "TypeError: code() argument 13 must be str, not int": a bare identity mapInArrow on 3.3.4, no numbarrow involved. The test extra and the README both declared 3.3 as the floor; pip accepted it because pyspark declares >=3.7, and the suite reported the failure as skips. The floor is 3.4.0, the first release that runs one. pandas 2.1.1, the floor the README named, installs next to numpy 2 and then fails to import with "numpy.dtype size changed"; 2.2.2 is the first release built against numpy 2, and the README row says so. The CI sentence described a matrix this repository does not run; it now describes the one it does.
pyarrow<=24.0.0 was a release behind PyPI. The suite passes on 25.0.1 with numba 0.67.0, numpy 2.5.3, pandas 2.3.2 and pyspark 3.5.7, so the cap, the README row and the ceiling cell of build-pyarrow-range move to it.
…eg is lost The four tests that run a child interpreter let it import whichever numbarrow it found: from any cwd but the checkout root, with a distribution installed, they validated a different copy. Each child now gets PYTHONPATH set to the tree. conftest turned any Spark or JVM failure into a skip, so losing every end-to-end test was a green run. NUMBARROW_REQUIRE_SPARK=1 turns that skip into a failure that says why, and the CI cells whose JVM runs the leg, the linux ones and macos, set it; windows never runs it and is left alone. The round-trip test compared values only, so a column could change type on the way through unseen, and it covered no temporal type. It now asserts the type that comes back, covers date32, date64 and naive and zoned timestamps, pins the three documented drifts (large_string to string, date64 to timestamp[ms], a zoned timestamp to naive) and checks that output_schema restores each. input_columns gets the end-to-end test it never had. The README type-table test skipped a scalar row whose Copy cell was not yes or no, so rewording a cell disarmed its row; it fails instead, and holds a per-field row to its label. mutation_guard_check prints the suite's last lines when the baseline fails, where before a red job named nothing.
…e fields differ The key check covered a plain list of dicts against a struct field and nothing else. A list-of-struct column fed lists of dicts, a struct field holding a dict, an object-dtype array of dicts and a generator of dicts all came back null on the same typo, and a ready-built StructArray whose field names differed from the declared ones was cast by name and came back all null too, since a cast matches struct fields by name as well. The check now follows the declared type through lists, maps and nested structs, over any iterable of Python objects, and a ready-built array is cast only once its field names, at every depth, are found among the declared ones. Output columns of unequal length raised pyarrow's bare "Arrays were not all the same length: 2 vs 3", naming neither column; the refusal now names each column with its row count. An OverflowError from a Python int too large for the declared type escaped without the column's name; it is caught and renamed like the rest.
…a column twice KeyError.__str__ reprs its single argument, so both new KeyErrors arrived wrapped in a second pair of quotes, and a plain KeyError renamed through `renamed`, such as pyarrow's own "Field ... does not exist in schema", came back doubly mangled. MissingKeyError keeps `except KeyError` working and prints the sentence as given; `renamed` builds one from a plain KeyError's argument rather than its repr. A batch with two columns of one name, which an unaliased join produces, died on pyarrow's own KeyError from batch.column, outside the code that names the column and without a remedy. It is refused naming the count and the alias that fixes it.
… index Three sentences claimed more than the code does. Only the views refuse data.flags.writeable = True; the copies, booleans, date32 and strings, start read-only but a caller can flip the flag and then writes into memory that is their own. Under NUMBA_DISABLE_JIT=1 a boolean or string column and any column carrying a validity buffer raise, since the viewers' intrinsic has no pure-Python form, but a null-free numeric column adapts through np.frombuffer and never reaches it. And bounds checking turns an index past the bitmap into IndexError but not a negative one, which reads from the bitmap's end like any numpy index, in bounds and wrong. The compatibility paragraph no longer says the numba floor is not swept, which is true of one CI and false of another that carries the same README.
…d align the pandas floors The key check followed a declared map into its values but not its keys, so a struct-keyed map fed a misspelled key field from pairs reached pa.array and came back null. Both sides are traversed now. A one-element "pair" raised IndexError inside the check, outside the wrapper that names the column; anything that is not a pair is left for pa.array to refuse, which is named. The output_schema docstring lumped an integer beyond int64 in with the ArrowInvalid cases; pa.array raises OverflowError for it, and the docstring says so. Both extras still declared pandas>=1.5.0 next to a README row that says 2.2.2; they now agree.
type_repr's docstring promises a note of what was cut, and the note said how long the whole type was. It now counts the characters past the width, and the test pins the arithmetic for the struct and the list forms.
A bare array returned from main_func hands pa.array no validity, so every null the UDF received came back as whatever sat under it, and a value the caller had masked out with pc.if_else came back verbatim. The README example did exactly this. A (data, bitmap) pair per output column now carries the validity out, the bitmap being the packed uint8 array bitmap_dict hands out, or None. A fixed-width or string column with no nulls of its own takes the bitmap as its validity buffer and keeps its data buffer; anything else is masked through if_else, which keeps the nulls it had. A bitmap of another length or dtype is refused naming the column. The README and rst examples return the pair, and the round-trip test covers every supported type with a null.
The bool and float64 viewers were compiled at import, used by nothing, and cost an index and a data file each in the numba cache. A view from the dict is a bare carray over the address it was given, with no owner and no read-only flag, and the utils page called it the key abstraction without a lifetime rule. The two entries are gone, the page says what the three viewers are and points at arrow_array_adapter as the supported way in, and the viewer factory's docstring carries the lifetime rule.
… bitmap on a resized column A bare 2-tuple was already a two-row column, so reading one whose second element is None or an ndarray as a (data, bitmap) pair changed the meaning of ([1, 2], None) from list<int64> [[1, 2], None] to int64 [1, 2] with nothing said, and gave (5, None) an error about a bitmap the caller never meant. The pair is now an explicit Nullable(data, bitmap), a named tuple no existing shape collides with, and a bare tuple means exactly what it meant before. A packed bitmap carries no row count, so the byte check could not see a UDF that dropped a row or two and passed the batch's bitmap through: the dropped row's bit landed on the wrong row. A bitmap the batch handed out is now accepted only on a column of the batch's row count, by identity; a bitmap the UDF builds for its own rows still goes through the byte check. The fold's docstring also stops claiming a no-copy fold for a non-contiguous bitmap, which every handed-out bitmap is not.
…s rows A list-of-struct column's bitmaps cover the flattened struct elements rather than the outer rows, so the documented pass-through Nullable(data_dict[col][field] * 2, bitmap_dict[col][field]) returned a six-element column from a two-row batch and was refused as a resize, while a byte-identical copy of the same bitmap was accepted and gave the right column. The refusal also asserted that the bitmap covered the batch's two rows, which a field bitmap does not. Every bitmap the batch hands out is now recorded with the count it covers, read off the data handed out beside it, and a returned column is refused only when its length differs from that count, naming both counts. A UDF that drops rows and hands the batch's own bitmap back is refused exactly as before.
pandas marks a missing row with NaN or pd.NA and never with None, so the key pre-pass over a declared list or map column walked into the marker itself and died on "'float' object is not iterable", refusing a whole batch whose column pa.array converts with that row as a null. The pre-pass exists only to catch a struct key no declared field has, so it now looks inside the rows it can iterate and leaves the rest to pa.array, which converts what it can and names the column when it cannot. The typo is still caught in exactly that shape.
The fold reached for bitmap.dtype before validating anything, and AttributeError is outside the classes the output wrapper catches, so a bitmap handed back as a list, as bytes, as a pyarrow array or as an int escaped as "'list' object has no attribute 'dtype'", naming neither the column nor what a bitmap must be, against the promise that every failure converting an output column names the column. Anything that is not an ndarray now raises the same named TypeError the wrong dtype raises, with the kind it was given.
The output_schema paragraph promised pyarrow.ArrowInvalid for a float with a fraction into an integer type and for a timestamp unit change that drops digits, and a Python list is a documented output shape. pa.array has two converters, though: the typed one an ndarray and a pyarrow Array take raises on both, and the sequence converter a list takes truncates both without a word, so [1.5] into int64 comes back 1 and a microsecond datetime into timestamp[s] loses its digits. The paragraph now names the shape each refusal holds for and lists the two truncations among the lossy conversions that pass silently. The four cases are pinned beside the type table, where a claim and its measurement sit together.
The binding paragraph said Spark checks nothing about an output column's type and that only a width mismatch fails, which is what int32 under LongType had shown. float64 under LongType and int64 under DoubleType are both 64 bits wide and both die in the JVM with java.lang.UnsupportedOperationException, so an operator chasing that exception was told the types could not be the cause. What swaps silently is two columns whose types share an accessor family; a column read through the accessor of another family fails whatever its width, and so does a width mismatch inside one family. The Spark leg measures all three refusals and the matching control.
Lazily typed it took one entry per signature in a single index file, which numba writes without a lock and names by counting the entries it just read, so two processes warming a cold cache could pick one data file for two signatures and leave a directory a later process dies on. An int64 index and an optional read-only bitmap per layer is a signature every caller shape resolves to.
Under a declared timestamp pa.array reads the 8-byte values of a datetime64[D] array as the 4-byte days of a date32, so three dates came back as the first, the epoch and the second. Seconds are exact for whole days, and every other conversion a declared type makes is unchanged, the refusal on a unit change that drops digits included.
| # booleans, int32 and int64 for string offsets. Each is compiled at import and | ||
| # carries an index and a data file of its own in the numba cache, so an entry | ||
| # nothing reads is not free. | ||
| arrays_viewers = { |
There was a problem hiding this comment.
Maybe we can have arrays_viewers that generates viewers on an as-needed basis, and stashes them. When it's asked for a viewer for a given numpy dtype, it will retrieve one if it's already available, or create, store, and return one otherwise.
There was a problem hiding this comment.
From the fake Slim Shady:
Done in 26f9a54: arrays_viewers builds a viewer through the factory the first time a dtype is asked for and keeps it, so nothing is compiled that nothing asks for, and a dtype the adapters do not use is one lookup away. The cache test asks for one viewer, then all three, and counts one index file, then three.
A str, a bytes and an ndarray of a non-object dtype iterate, but over scalars that carry no keys, and pa.array refuses such a row at its first element. The key pre-pass spread the whole row into a list before that refusal: 7 s and 161 MB for a 20M-char str, 19 s and 1.4 GB when its characters are outside latin-1, 18 s and 641 MB for a 20M-element int array, measured on pyarrow 24.0.0. The test pins it with rows whose __iter__ raises, and the mutation catalogue holds each new term.
The compatibility paragraph still explained a 2.2.2 row against a 1.5.0 mapinarrow extra after both extras had moved to 2.2.2, so it described a discrepancy the project no longer had. The test reads every pandas floor out of pyproject.toml and holds the README's row to it.
The check on "cache" stopped one string short of the rule its own
message states: numba reads any non-empty string as true, so
'{"boundscheck": "false"}' turned bounds checking on without a word,
and nogil, looplift, no_rewrites, debug, forceinline and _nrt take the
string the same way. Only error_model and inline take a string by
design; parallel and fastmath numba refuses itself. The mutation
catalogue holds the new term.
The key check reads the rows before pa.array does, and a generator read once has nothing left for the second reader, so the guard reads it into a list first. Nothing tested the guard: with it deleted the suite stayed green while a generator of well-keyed dicts came back as a column of no rows. The test pins the rows and the mutation catalogue holds the guard.
pa.StructArray.from_arrays([], names=[]) has no child to take a length from, so a record array with no fields came back as a struct column of no rows, with or without a zero-field struct declared, and as the only output column dropped the batch without a word; before this branch pa.array refused the shape outright. The childless struct is now built with the record array's row count, pinned for both shapes, and the mutation catalogue holds the term.
arrays_viewers held three viewers compiled at import. As suggested in review it is now a dict that compiles a viewer through the factory on the first request for a dtype and returns the same one after, so nothing is compiled that nothing asks for, and a dtype the adapters do not use is one lookup away instead of gone. The cache test asks for one viewer, then all three, and counts one index file, then three; the mutation catalogue holds the keeping.
The recursive call that reads a struct inside a struct was covered by no test: deleted, a nested field the declared type does not name was filled with nulls by the cast while the suite stayed green. One case per branch now, a struct inside a struct and a struct inside a map beside the list case, and one catalogue entry per branch.
Widened to seconds under a declared timestamp alone, a day-unit array declared int32 still had its 8-byte days read as 4-byte ones: four dates came back as two day numbers and two zeros. Inferred first it is a date32, and the cast to the declared type gives the day numbers under int32, the midnights under every timestamp unit, ISO dates under string, and pyarrow's own refusal elsewhere, each naming the column.
Without a schema the keys of the dict main_func returns are the column names, and a non-string one died inside pyarrow on "expected bytes, int found", naming neither the key nor the rule, where every other refusal on the output side says what to return.
A dictionary-encoded struct array under a declared dictionary-of-struct with other field names was cast, and a cast decodes the dictionary and matches the value structs by name: the column came back all null, which is what the check exists to refuse. A dictionary on both sides is now compared by its value types.
The docstring named four shapes; a bare tuple and an object ndarray holding None come back with a null as well. The docs test runs every shape the sentence names and a bare array.
A record array under a non-struct declared type is refused as such, a child that fails conversion names its field, a ChunkedArray output is combined before the field check reads it, a malformed pair in a map of structs is left to pa.array, and an exception that cannot be rebuilt from one string is renamed as a ValueError: each survived deletion with the suite green. One test and one catalogue entry each.
| ``` | ||
|
|
||
| A bare array returned under a column name carries no nulls out: every null the | ||
| UDF received comes back as whatever sat under it, `0`, `0.0` or `''`. |
There was a problem hiding this comment.
"sat under it", sat - in the sense of as it was brought in from the input? If yes, then why is it relevant for the discussion of nullable output?
There was a problem hiding this comment.
From the fake Slim Shady:
Yes, the value Arrow keeps in the data buffer under an input null. It matters on the way out because a bare output array has no validity, so that row goes out valid, holding whatever the UDF computed from the placeholder (0.0 in the README example) instead of a null. Reworded to say that in 569e91d.
| return {"output_col": result} | ||
| # A bare array carries no nulls out; Nullable(data, bitmap) keeps them, | ||
| # bitmap being the packed uint8 array bitmap_dict hands out, or None | ||
| return {"output_col": Nullable(result, bitmap_dict["input_col"])} |
There was a problem hiding this comment.
Can you please elaborate a bit why is bitmap for 'input_col' bundled up in the same tuple with the calculated 'result' array? I understand a more motivated use-case it probably to include the bitmap of 'result', and unless the latter happens to be identical with the bitmap of 'input_col' - which is not intuitive to be a general case - we probably shouldn't be conflating them.
There was a problem hiding this comment.
From the fake Slim Shady:
Agreed, the example didn't motivate it. Passing the input's bitmap through is only right when the result is null exactly where that input is, like an elementwise function of it. In 2a7d362 the example computes data_dict["input_col"] * broadcasts["scale"], so the pairing holds, and the docs show the other case: a result with nulls of its own packs them with np.packbits(valid, bitorder="little").
| array = _convert(value, arrow_type) | ||
| if bitmap is None: | ||
| return array | ||
| covers = handed.get(id(bitmap)) |
There was a problem hiding this comment.
No need to change anything, just curious if bitmap.ctypes.data was considered and ruled as inferior or equivalent (most likely) choice here.
There was a problem hiding this comment.
From the fake Slim Shady:
I didn't compare them at the time, but they aren't quite equivalent. A ctypes.data key would also catch a view of a handed-out bitmap, but it would refuse a correct prefix slice like Nullable(result[:5], bitmap_dict["v"][:1]), since the slice starts at the same address and the check would compare 5 rows against the batch's 10. id() only matches the object that was handed out, and everything else gets the byte-length check, so I'd keep it. The one hole it had, a bitmap the UDF dropped from bitmap_dict being freed and its id reused by one the UDF built, is closed in b53a383: each entry keeps its bitmap alive for the batch.
| The bitmap has the layout ``bitmap_dict`` hands out and | ||
| :func:`~numbarrow.core.is_null.is_null` reads: one bit per row, LSB first, | ||
| set for a valid row, ``(rows + 7) // 8`` bytes of uint8. A fixed-width or | ||
| string column with no nulls of its own takes the bitmap as its validity |
There was a problem hiding this comment.
No nulls of its own means bitmap is None, no?
Edit: nm, I understand it's a more general utility that allows to associate any bitmap with the given array.
There was a problem hiding this comment.
From the fake Slim Shady:
Right, any bitmap on any array. "No nulls of its own" meant the array built from the UDF's data before the bitmap goes on, so it says "no nulls yet" now, in 09d7449.
| set for a valid row, ``(rows + 7) // 8`` bytes of uint8. A fixed-width or | ||
| string column with no nulls of its own takes the bitmap as its validity | ||
| buffer and keeps its data buffer, so neither side is copied when the | ||
| bitmap is contiguous, which every bitmap ``bitmap_dict`` hands out is. |
There was a problem hiding this comment.
Btw this sentence reads a bit Kantian ;) Instead of ", which every bitmap bitmap_dict hands out is." it might be easier to comprehend it as ", which is the case for every bitmap bitmap_dict hands out."
There was a problem hiding this comment.
From the fake Slim Shady:
Fair enough. Taken as suggested in 09d7449.
"Comes back as whatever sat under it" left open what sat under what and why it matters on the way out. A bare output array has no validity, so a row that was null on the way in goes out valid, holding what the UDF computed from the placeholder under the null: 0, 0.0 or '' in a batch from Spark, as test_nullable_carries_nulls_back_through_spark records.
The docs example paired an unspecified result with bitmap_dict["input_col"], which read as if a result's validity were the input's in general. It is only when the result is null exactly where that input column is, as for an elementwise function of it, so the example now computes one, the comment says why the bitmap goes with it, and the prose shows the other case: np.packbits(valid, bitorder="little") builds a bitmap in the layout bitmap_dict hands out and is_null reads.
"No nulls of its own" read as a column whose bitmap is None, when it means the array built from the UDF's data before the bitmap goes on, so it says "no nulls yet". The contiguity clause takes the wording offered in review.
_handed_bitmaps keyed the row counts by id() and held nothing else, so a UDF that dropped a bitmap from bitmap_dict freed it, the next one-byte bitmap the UDF built landed on the same id, and a correct resized column with that bitmap was refused as if it carried the handed-out one. Each entry now carries the bitmap beside its count, which keeps it alive for the batch, so no bitmap the UDF builds can share an id with one. The test drops a bitmap and checks through a weakref that it survives; the catalogue entry stores None in its place.
|
Thanks for contributing, merging this. |
…26-09-25 Close what a whole-tree pass found after #11: the output side's row shapes and time units, two aborts, a use-after-free, and a cache that could not be written
Two things you agreed to on #9, the rest of the whole-tree pass at d240a90 that a green suite could not see, and two contract choices the pass left open, made here and easy to reverse: nulls on the way out, and what
arrays_viewersis for.Struct columns now arrive nested:
data_dict[col][field]andbitmap_dict[col][field], with uniform columns unchanged. A field no longer shares a namespace with another column or with another struct's fields, so the collision check from #8 goes, and so does the refusal of a zero-field struct, whose reason was that such a column vanished from both dicts. The struct-level validity is still folded into each field's bitmap.The output side of
make_mapinarrow_funcis where the findings were:pa.arraygot no validity, so every null the UDF received came back as whatever sat under it,(1, None, None, None)as(1, 0, 0.0, '')through Spark, and a value the caller had masked out withpc.if_elsecame back verbatim; the README example did exactly this. ANullable(data, bitmap)per output column now carries the validity out, the bitmap being the packed uint8 arraybitmap_dicthands out, orNone, so passing it through isNullable(result, bitmap_dict["value"])and a kernel that decides its own nulls clears bits in a layout it already reads throughis_null. A fixed-width or string column with no nulls of its own takes the bitmap as its validity buffer and keeps its data buffer; anything else, a dictionary column among them, is masked throughif_else, which keeps the nulls it had. A bitmap of another length or dtype, or one that is not an ndarray at all, is refused naming the column, and so is a bitmap the batch handed out on a column whose length is not the count it was handed out for, the batch's rows for a column and the flattened element count for a struct field's, since a packed bitmap cannot tell counts apart inside one byte. It is a named tuple rather than a bare pair because a 2-tuple was already a two-row column. The README example returns one.|Uoutput is built asstringfrom itstolist()and a|Soutput asbinary. An empty or all-null batch keeps its type, where before a UDF that filtered a whole batch away yielded anullcolumn and Spark's stream writer refused the schema the next batch did not share; and a bytes value keeps its NULs instead of being cut at the first one.output_schemaevery column is built as its declared type rather than inferred and cast, with one exception: adatetime64[D]array is inferred asdate32first and cast to its declared type, sincepa.arrayreads its 8-byte days as the 4-byte days of adate32under a timestamp or anint32and returns the wrong column. A list of dicts keyedAmount/Labelagainststruct<amount, label>used to come back as two columns of nulls under an identical schema, since Arrow matches struct fields by exact name and nulls a missing one; a key no declared field has now raises, however deep the struct sits inside a list, a map or another struct, a row that carries no keys,Noneor a pandas marker, is left topa.array, which makes it null, and a ready-built array whose field names differ from the declared ones, a dictionary-encoded one included, is refused rather than cast by name into nulls. A map column, which no inferred struct casts to, is producible from a list of dicts or of pairs. A key the schema does not name raises instead of being dropped, and a key that is not a string is refused by name.pa.arrayread as its keys, is refused naming the column. A numpy record array, what an@njitfunction returns for a record type, becomes a struct column field by field instead of dying on "Unsupported numpy type 20". A result that is not a dict, two columns of unequal length, and every other failure on the output side name what they are about.output_schemadocstring promisedArrowInvalidon any lossy conversion. Measured, from an ndarray or apyarrow.Array: an integer out of range, a float with a fraction into an integer type and a timestamp unit change that drops digits raise; a timestamp into a date type floors to the day and float64 into float32 overflows to inf, both silently. From a Python list, or any other sequence of objects, the integer out of range still raises, but the fraction and the dropped digits are truncated silently, sincepa.array's sequence converter has no such check. It says that now, by shape.The five
arrays_viewersshared one numba cache index and one set of data files, written without a lock, so processes importing together on a coldNUMBA_CACHE_DIRpoisoned it for every later process: eight importers did it in five rounds of five, and the consumer then died inNRT_adapt_ndarray_to_pythonor read an int32 column as float64. Each viewer now carries a qualname of its own, so it has its own index and data files; five rounds of 8 importers and three of 24 pass.is_null_structwas cached with no signature, so every index type and bitmap combination it was called with was one more entry in one such index, and 32 importers on a cold cache left a directory a later process died in; it now carries one explicit signature, anint64index and an optional read-only bitmap for each layer, so one entry. The subprocess test that counts those files is also the first end-to-end test thatNUMBARROW_JIT_OPTIONSreaches the decorators."cache"must be a JSON boolean, since numba read the string"false"as true, and a string is refused for every other option buterror_modelandinline, the two numba reads as strings.Of the five,
boolandfloat64were compiled at import and used by nothing, and the docs called the dict the key abstraction while a view from it is a barecarrayover an address, with no owner and no read-only flag. Nothing is compiled at import now:arrays_viewersbuilds a viewer the first time a dtype is asked for and keeps it, so a process compiles the three the adapters ask for and any other dtype is one lookup away; the utils page says what a view from it is and points atarrow_array_adapteras the supported way in, and the viewer factory's docstring carries the lifetime rule.data.flags.writeable = Truesucceeded on a view over a locally built array, and a store through it changed the source Arrow array; the view is now over a read-only memoryview, which numpy refuses to flip.Messages: a Table column is named as a ChunkedArray rather than blamed as int64; a struct child's failure names the field and, through the factory, the column; invalid UTF-8 is a
ValueErrornaming the element;input_columns="value"is refused rather than read as five columns; anoutput_schemathat is not apyarrow.Schema, a SparkStructTypebeing the likely one, is refused at factory time naming the parameter and the type it needs, rather than at the first batch; a column name Spark's case-insensitive projection rewrote raises aKeyErrorthat lists the batch's columns, and a name the batch carries twice, as an unaliased join produces, is refused with the alias that fixes it; a KeyError's message reads as written rather than wrapped in a second pair of quotes; the two asserts guarding the struct and list adapters survive-O; and, as you agreed on #9, a type wider than 120 characters is cut with a count of what was cut, and the list message no longer repeats the element type.Docs: Spark binds output columns by position, so two columns whose types share an accessor family swap silently, an int64 declared as a timestamp reads as one, while a column read through the accessor of another family fails in the JVM whatever its width, float64 under
LongTypeas much as int32; adatetime64output comes back a naive timestamp of its unit, except the day unit, which comes backdate32; a bitmap isNonewhen the batch has no validity buffer, not when it has no nulls, so bitmap parameters wantOptionalin an eager signature; the timestamp handler drops the zone asto_numpydoes; the width exception fires on a narrower batch too; the read-only paragraph says which results are views and which are copies, and that only the views cannot be made writable; theMapArraysentence;is_null's index range and whatboundscheckdoes and does not catch;NUMBA_DISABLE_JIT.Floors: pyspark 3.3 bundles cloudpickle 2.0.0, which cannot deserialise a UDF on Python 3.11 or later, so the test extra and the README say 3.4.0; pandas 2.1.1 installs next to numpy 2 and fails to import, so the README row says 2.2.2; pyarrow 25.0.1 is admitted and is the ceiling cell of
build-pyarrow-range.Tests: the four subprocess tests validate the tree under test rather than whichever numbarrow the child found;
NUMBARROW_REQUIRE_SPARK=1turns a lost Spark leg into a failure and the linux and macos cells set it; the round-trip test asserts the type that comes back, covers the temporal types, and carries a null through every supported type; the cold-cache test coversis_null_structbeside the viewers; the README type-table test fails rather than skips on a reworded cell; the mutation catalogue grows from 11 entries to 60.