Skip to content

Close what a whole-tree pass found after #11: the output side's row shapes and time units, two aborts, a use-after-free, and a cache that could not be written - #12

Merged
Goykhman merged 59 commits into
Goykhman:mainfrom
nelson2005:upstream-pr/whole-tree-pass-2026-09-25
Sep 29, 2026
Merged

Goykhman merged 59 commits into
Goykhman:mainfrom
nelson2005:upstream-pr/whole-tree-pass-2026-09-25

Conversation

@nelson2005

@nelson2005 nelson2005 commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

The rest of what a whole-tree pass at 9edb99f turned up, after #11 closed the first one, and what a second pass over the branch's own diff found after that: twelve commits at the end, two of them regressions of the branch, and a thirteenth pinning the extras gate's floor parsing; four more came out of the review. Fifty-nine commits: thirty-five change the library, eight the docs and the packaging floor, and sixteen the tests, the CI, the catalogue and the helpers' layout. Every new guard has its test and, where it gained a term, a catalogue entry; the catalogue stands at 118 entries, all killed, and the suite at 296 with the Spark leg required, on pyarrow 14.0.0 and 25.0.1 as well as 24.0.0.

The output side of make_mapinarrow_func is where most of it was.

On the input side.

Compilation. numba raises at decoration where no cache location can be written, a read-only install or an import from an .egg, .whl or .pyz archive such as spark-submit --py-files ships, and nothing named the way out. The three compiled functions and the viewers decorate through one helper that compiles such a function without a cache and warns with the remedy that works there: NUMBA_CACHE_DIR for a source file on disk, and for an archive, where numba never reads it, an unpacked install or a .zip, which numba 0.61 and later cache; NUMBARROW_JIT_OPTIONS='{"cache": false}' silences it either way. A .zip whose cache directory cannot be written, which numba takes unchecked, compiles uncached too instead of failing the import. The README says how the cache works, that NUMBARROW_JIT_OPTIONS is read when numbarrow is first imported and NUMBA_CACHE_DIR when numba is, and that numba re-reads its environment only when it compiles something, so a directory set between the two imports misses at least the first function numbarrow compiles.

Messages. The unexpected-keys listing, sized by the data, is cut with a count; a nested key refusal names the path to the field; a record array field that fails names the field on the inferred path as on the declared one; the dispatcher describes a scalar as one and leaks no column called type; an output_schema naming a column or field twice is refused when the function is made, and so is an empty input_columns, a generator already used up or a filter that matched nothing, which handed main_func no columns for a bare KeyError or a batch of no rows and no columns; the function names the shape it takes when handed a batch or a table; the options refusal names the rule and shows the value; the 64-bit date view refuses a unit that is not the array's own, and a time64 or duration array, which carries a unit too and came back as instants on 1970-01-01.

Docs. Spark reads a struct column's fields by position too, and output_schema derived from the Spark schema keeps the two orders equal; the list of silent lossy conversions gains the time types, bool and narrower floats and the OverflowError an unsigned type raises, and no longer the decimals it first named, since a float under a decimal type is refused on both routes; the negative-index sentence is bounded; the rendered parameter names keep their underscore; the read-only paragraph names the @njit sub-view route; the mapinarrow extra supplies setuptools; the 3.3 rationale names the driver-side failure. The test extra's pyspark floor is 3.5 on Python 3.13, where 3.4 does not import.

Tests. Timestamp units at value level, the key check's list layouts and null rows, the options reaching is_null's decorators and the boundscheck contract, an empty bytes column, every refusal class of a record field, and the dispatcher's own cut, each of which a mutant had passed; then the four terms the second pass found a mutant still passing: the no-locator narrowing of the cache fallback, pinned; the repeated-name check inside a map and through an extension type, pinned; the union check's extension unwrap, pinned; and the key check's two, pinned, the last three over an extension type the tests define so that pyarrow 14 exercises them; and the extras gate's floor parsing, its ~= clause and its refusal of a set naming no floor, which the workflow's run against this tree's own >=3.12 had never reached, pinned. In CI the Windows install and lint steps run under bash, where a failing pin or a failing syntax gate no longer goes green, and the extras gate parses the requires-python floor, a ~= clause included.

A child that failed to convert was wrapped with its field name only on the declared path. With output_schema left as None the same failure named the output column alone, as in "output column 'out': Unsupported numpy type 15", and the field it choked on had to be guessed from the dtype. Both paths now convert a field through one helper that names it, the docstring says so, and the catalogue pins the inferred path.
The names were read from the argument inside the batch loop, so a generator, map() or filter() handed in was used up by the first batch, every later batch was adapted with no columns at all, and the UDF died on a bare KeyError that said nothing about input_columns. The names are read once now, deduplicated in first-seen order as before, the docstring says a one-shot iterable serves as well as a list, and the catalogue pins the read-once line.
… field

The earlier commit moved the raise that entry mutated into _record_field, so the entry's old text was gone and the catalogue reported it stale instead of running it. The entry now mutates the helper's raise, which drops the field name on the declared and the inferred path alike, and the tests for both kill it.
…r days, before Arrow reads a datetime64 or timedelta64 output

pa.array reads numpy's base unit and ignores a multiplier, so a datetime64[5s] column of five-second bins came back at one-second steps, 2020 read as 1980, and datetime64[2D] slipped past the day-unit inference into the misread it guards against. An hour or minute unit, and a week, month or year unit, raised ArrowNotImplementedError against the docstring's promise of a timestamp of the unit. A multiplier now folds into its base unit, hours and minutes become seconds, weeks, months and years become days and so date32, a unit finer than a nanosecond or a month or year timedelta is refused by name, and the docstring and README say so.
…racter per row

A str or bytes returned as a column went to pa.array, which iterates it, so {'country': 'US'} over a two-row batch was the rows U and S, and a 0-d unicode or bytes array's tolist() is that scalar, which defeated the one-dimensional refusal pa.array gives the array itself. Both are refused naming the column, and the message says how to build a constant column.
A helper's Nullable wrapped once more with the input's bitmap went to pa.array as the 2-tuple it is and came back as a two-row list column, its data as one row and its bitmap's bytes as the other, silently on every batch whose column had no validity buffer and with a message blaming a resize otherwise.
…n, sequences by __getitem__, key/value entry dicts, and refuse a namedtuple in another field order

The key check kept only Mapping rows and 2-tuple map entries, and pa.array accepts more: a struct row given as a tuple, a namedtuple or a pyspark Row binds by position, a sequence by __getitem__ alone iterates without an __iter__ attribute, and a map entry may be a {key, value} dict, the shape Spark's map_entries produces. Behind any of those a mistyped nested key was silently nulled, and a namedtuple or Row naming the declared fields in another order swapped every same-typed field. Tuple rows are now checked by position, iterability is tested with iter, entry dicts are read as pairs, a permuted namedtuple is refused by name, and a pyarrow scalar row is left to pa.array, which from pyarrow 21 made a MapScalar a Mapping whose values is an array and crashed the check.
… KeyError, and combine a chunked array pa.array hands back

pa.array reads a pandas Series row by its index labels, so a sorted or filtered one came back reordered without a word and one from a groupby died on a bare KeyError(0); a DataFrame under one output key died the same way or came back transposed; a KeyError raised inside pa.array escaped _to_arrow and _record_field unnamed although renamed has a branch for it; and a Series over a multi-chunk pyarrow array came back from pa.array as a ChunkedArray that RecordBatch.from_arrays refused naming no column. Series and DataFrame rows and a DataFrame column are refused naming the column and the remedy, KeyError joins both except tuples, and a ChunkedArray is combined.
…, a map against a declared list of entries, and the list view layouts

The ready-built-array field guard compared like kinds only, so an extension array with struct storage cast to a declared struct, a MapArray cast to a declared list of key/value structs, and a list_view or large_list_view container declared for dict rows all came back all-null with no refusal. An extension type is compared and key-checked through its storage, a map is paired with a list of entries through its entries struct, and the view layouts count as list-like where the installed pyarrow has them.
…st, naming the column

Without output_schema every list, tuple or object column's type was inferred from its own batch, so an empty or all-None batch beside a full one, or ints beside floats, carried a second schema into Spark's writer, which refused it with Tried to write record batch with different schema and named nothing. The first batch's schema is now held for the partition and a later batch that differs is refused naming the column, both types and the remedy; the string and bytes pins stand as they were.
_with_validity chose its path by the extension type, which reports no fields and no dictionary whatever its storage, so a Nullable over dictionary storage took the from_buffers path the dictionary term exists to avoid and aborted the interpreter, and one carrying a null went to if_else, which has no extension kernel. The storage is masked and the result rewrapped.
structured_array_adapter called flatten() on every child of a struct with a null row before any child was dispatched, and flatten() hands the struct's validity to each child; a union carries no validity buffer of its own, so Arrow's C++ layer failed its check and aborted the process where the documented NotImplementedError naming the field was due. A union child, extension storage included, is refused by name before anything is flattened.
np.frombuffer handed a memoryview keeps only a wrapper of it as the result's base, and that wrapper's release() drops the memoryview's hold on the source, so a caller who released it, dropped the source array and read the view read freed memory, a segfault at 64 Ki elements, reachable through the public adapters and a UDF's data_dict. The read-only memoryview is now wrapped in a foreign pyarrow buffer, which is what the view keeps as its base: it has no release(), holds the source through the memoryview, and exports read-only, so the flip stays refused. The docs test's view detection accepts a pyarrow buffer at the end of the chain.
…fusal

The one-dimensional guard was meant for the unicode and bytes route alone, where tolist() turns a 0-d array into a scalar and a 2-d one into nested lists and so defeats the refusal pa.array gives the array itself; every other dtype gets that refusal, which names the column, as before.
The generator kept the adapted arrays, the views, the struct triples and the handed-out bitmaps bound in its frame while the consumer wrote the batch out and the next one was adapted, so a string column's |U copy was live twice at the peak, the README's 1.6 GB batch peaking at 3.2 GB, and a handed-out bitmap outlived the batch its docstring says it lives for. Those names are cleared before the yield, and the yielded batch is dropped on resume.
numpy names every structured dtype of one itemsize void<bits>, so two of them compiled under one qualname, landed in one cache index with identical argtypes, and a process that loaded both from the cache ran the first one's code for the second, boxing an int32 field's bytes as float32. A digest of the dtype's description joins the qualname of a structured dtype.
The date32, date64 and timestamp handlers took the bitmap from the result of an integer cast, and pyarrow drops the validity buffer when it casts a zero-length array, so a zero-row column that still carried one adapted to a bitmap of None, against the documented rule that None means no buffer, where every other type gives an empty uint8 array. The bitmap comes from the source array when the column is empty.
The struct key check listed every distinct key no declared field has, and that list is sized by the data rather than the schema: a UDF keying a dict by a row value put every key of a 100,000-row batch, 1.5 MB, into the exception and twice into the executor logs. Ten are shown with a count of the rest, as a wide type is cut.
… type

The dispatcher's fallback read len() and .type of whatever it was handed: a null list scalar died on len(), a struct or map scalar of a supported column was described as an unsupported array of that type, and a pandas frame or record array with a column called type put that column's values into the message or died in type_repr. A scalar is named as one, and only a DataType reaches the message.
A dict holds one value per name, so a name declared twice at any depth was filled twice from the same entry, and the batch died in the JVM with not all nodes and buffers were consumed, naming neither the column nor the repeat; the input side refuses the same shape by name. The schema is checked once, when the function is made.
The generator read .schema of whatever its iterator yielded, so udf(batch) and udf(table) walked the columns and died on an attribute of the first one, and mapInPandas's pandas frames died the same way. An item that is not a RecordBatch is refused naming the shape mapInArrow passes.
One message served both failures, so a value that was valid JSON but a list, a number, null, true or a string was told it must be valid JSON, and neither failure showed the value; the sibling refusals in the same function name the option and repr the value. Both messages now name the object requirement and the value, and the JSON error names its position.
…d say what the function does

cast_64bit_date_arrow_to_numpy_array's docstring offered a cast to various units and its body reinterpreted the int64 payload, so timestamp[ms] asked for as datetime64[s] came back in the year 51971 for a caller who followed the docs page. The unit must be the array's own now, a mismatch is refused naming both, and the docstring says it is a view.
…ze binary type

tolist() drops a trailing NUL, so one digest in 256 came back a byte short and the whole batch was refused under fixed_size_binary(16), and the docstring blamed numpy for a byte the buffer still held. Under a fixed-size binary type pa.array keeps every byte of the array itself, and only there; variable-width binary still cuts at the first NUL and keeps the tolist() route.
Two same-typed sibling fields, a map's key and its value, and different nesting depths all raised a byte-identical message naming the column and the declared struct alone. The check carries the path through its recursion, so the refusal reads output column 's': field 'm': map value: declared ...
…a can write none

numba sets a cached function up when it is decorated and raises RuntimeError there when no cache location can be written: a read-only install, an unwritable site-packages and user cache directory, or an import from an .egg, .whl or .pyz archive, which Spark's --py-files ships. The import then failed with no locator available, naming neither NUMBA_CACHE_DIR nor NUMBARROW_JIT_OPTIONS. The three compiled functions and the viewers now decorate through one helper that compiles such a function without a cache and warns with both remedies, and the README says how the cache works.
…w to derive output_schema from the Spark schema

The positional-binding paragraph covered top-level columns; a record array's dtype order and a dict's key order decided where each struct field landed, batch by batch when some rows left a null field out, and the docstring's remedy did not mention that output_schema binds struct fields by name or that its column order has to match the schema mapInArrow is given, which this function never sees. The docstring and the README say both, and name to_arrow_schema as the way to keep the two orders equal.
… where an unsigned type is declared

The closed list omitted a timestamp into a time type, which drops the date, a float into a decimal, which rounds, a number into bool, and a float into a narrower float than float32; and a negative value under an unsigned declared type raises OverflowError, which the documented except pa.ArrowInvalid does not catch, as does a value beyond the type's own ceiling rather than beyond int64.
…ames, and say which route flips a view writable

is_null's docstring said every negative index stays in bounds and bounds checking cannot catch it, which holds only down to minus eight times the bitmap's length. Sphinx read the trailing underscore of :param index_: and :param dtype_: as a reference marker and rendered index and dtype, so a keyword call copied from the docs raised TypeError. A slice or reshape of a view that an njit function returns is boxed writable by numba and its flag can be flipped, a route the read-only paragraph did not mention; pyarrow's own view has it too.
…ld itself

The test's comment credited the folded bitmap with seeing the null struct row, but Spark's own Arrow writer nulls the child of such a row before the batch arrives, so the test passes with the fold reduced to a no-op; the fold is pinned by the non-Spark tests that build a child carrying no null of its own, and this test is the end-to-end check that the row reaches the UDF at all.
…ruct

The refusal of a Series row named row.to_numpy() and list(row) as the
remedies, and under a declared struct column pa.array refuses both in turn,
wanting a dict or a tuple; the only shape that works there, row.to_dict(),
went unnamed. The remedy now follows the declared type, with a test that
follows it under a struct and a catalogue entry for the term.
…her than rescaling its counts

Taking an hour, minute, week or day unit to seconds or days happened before
the declared type was applied, so counts of [1, 2] days declared int64 came
back as [86400, 172800] with nothing raised, where pa.array had refused the
unit outright. The rescale now waits for a temporal declared type, or none,
and under any other type the array is refused naming the unit and the two
ways out; a multiplier still folds under every type, since pa.array would
otherwise read the count as one of the base unit. The factory docstring says
so, a test covers the three units under int64 and string, the fold under
int64 and the coarse unit under date32, and a catalogue entry pins the term.
…ily, not by a unit's presence

The view refused a type with no unit as not a date64 or timestamp, and a
time64 or a duration carries a unit too, so an hour of the day and a
ninety-second span were viewed as instants on 1970-01-01. The type family is
tested now, with the two shapes in the view's test and a catalogue entry
that puts the unit test back.
…es for every dtype

The cache-name suffix hashed dtype.descr, which numpy refuses to build for a
dtype with out-of-order or overlapping fields, the very dtype its own
multi-field indexing, rec[["b", "a"]], hands back, so the factory raised
ValueError where it had built a viewer before. The repr tells the dtypes
apart as well and exists for all of them; a test builds the viewer over the
reordered dtype and reads through it, and a catalogue entry puts descr back.
… test

The decorator re-raises any RuntimeError at decoration but numba's "no
locator available", and nothing tested that term: with it dropped, every
such failure compiled uncached behind a warning about the cache, and the
suite stayed green. A test hands the decorator another RuntimeError and
expects it back, then the no-locator one and expects the uncached recompile
behind the warning; a catalogue entry drops the term.
The check looks inside a map's entries and through an extension type's
storage, and only its struct and list paths were exercised: deleting either
term left the suite green. The factory-time test now declares the repeated
name inside a map and under an extension type, over an extension type the
tests' common module defines so that every pyarrow the matrix runs exercises
it, and two catalogue entries delete the terms.
The refusal of a union under a struct looks through an extension type over
the union, and nothing exercised that: with the unwrap deleted the suite
stayed green while flatten() would again hand the struct's validity to a
child that cannot take it. The union test wraps its union in an extension
type as well, and a catalogue entry deletes the unwrap.
The key check looks through an extension type over a struct twice, when it
asks whether a type carries keys and when it reads the fields, and nothing
exercised either: with one unwrap deleted the suite stayed green while
pa.array built a column of nulls for dicts keyed Amount under an extension
type over struct<amount>. The any-depth test gains that shape, and two
catalogue entries delete the unwraps one at a time.
…barrow is

The README said both cache variables are read when numbarrow is first
imported. NUMBARROW_JIT_OPTIONS is; NUMBA_CACHE_DIR is numba's, read when
numba is imported, which another library may have done earlier, and again
only when numba compiles something, so a directory set between the two
imports missed the first function numbarrow compiles, and all of them when
the old cache was warm. The paragraph says so now, and a test imports numba
under one directory, switches to another and imports numbarrow, expecting
the first function's index under the first.
The factory docstring listed a float into decimal among the lossy
conversions that pass without a word, rounding to the declared scale. Both
routes refuse it: the sequence converter takes an int or a Decimal, and the
typed one refuses a float dtype. The docstring says so, and a test holds it
to that on both routes and passes a Decimal through.
The gate reads its interpreter from the floor of requires-python, and the
invariants workflow runs it against this tree's own ">=3.12" alone, so the
~= clause it also accepts and the refusal of a set naming neither had no run
behind them. The test loads the script by its path and skips where the tree
carries no .github, as the catalogue's copies do not.
The gate imports tomllib, which arrived in Python 3.11, and the matrix runs
3.10 cells below the declared floor on purpose, so the test errored there
with ModuleNotFoundError and fail-fast cancelled the rest of the matrix. The
test now skips where the script cannot load, as it does where the tree
carries no .github.
Comment thread numbarrow/core/configurations.py Outdated
jit_options = get_jit_options()


def jit_with_options(*signature):

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Trying to understand the underlying motivation. A user can set env var with no cache, thereby achieving the same goal, right? I guess this is more convenient for the user, on top of the environment layer?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From the fake Slim Shady:

Right, NUMBARROW_JIT_OPTIONS='{"cache": false}' gets there too. The helper is for whoever hasn't set it yet: an import from an egg, which Spark's --py-files ships, used to fail at import numbarrow with numba's cannot cache function ...: no locator available, which names neither that variable nor NUMBA_CACHE_DIR. Now the import goes through uncached and the warning names both.

Comment thread numbarrow/core/configurations.py Outdated
jit_options = get_jit_options()


def jit_with_options(*signature):

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will never be used without signature, right?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From the fake Slim Shady:

Right, every caller passes exactly one, so it takes exactly one now: 83713bb. With *signature a second one went on to numba's jit as locals and failed with TypeError: 'Signature' object is not iterable.

def test_input_columns_is_read_once_so_a_generator_serves_every_batch():
# The names were read from the argument inside the batch loop, so a
# generator, map() or filter() was used up by the first batch and every
# later batch was adapted with no columns: the UDF died on a bare KeyError.

@Goykhman Goykhman Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The test is that this does not happen, right? I would make that more explicit.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From the fake Slim Shady:

Right. The comment now leads with what the test checks: 93eb2c9.

Comment thread test/test_cache.py Outdated
zipped.write(path, str(path.relative_to(REPO)))
env = dict(os.environ, PYTHONPATH=str(archive), NUMBA_CACHE_DIR=str(tmp_path / "cache"))
env.pop("NUMBARROW_JIT_OPTIONS", None)
probe = "import numbarrow.core.adapters as a; print(a.__file__)"

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can also set sys.frozen = True here, right?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From the fake Slim Shady:

You can, but then numba takes its frozen-app path and our fallback never runs. Its user-wide locator accepts a source file that isn't on disk when sys.frozen is set, caches outside NUMBA_CACHE_DIR, and stamps the cache with the interpreter instead of the archive. So a rebuilt egg keeps the old machine code: a toy egg whose cached function returned a module constant still returned the old value after I rebuilt it with a new one. That's right for a real frozen app, whose code only changes with its executable, but not for an egg, so I left the test on the fallback.

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, but NUMBA_CACHE_DIR being set will not necessarily silence 'no locator available' error, right? In your test here you silence it because you disable caching altogether.

@Goykhman Goykhman Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also, one can imagine that the package "normally" gets rebuilt (and environment re-established) after any inlined globals and any other lib details have been modified, or any other code changes have been applied.

My question is then, would sys.frozen by a more viable recommendation to use cache (rather than turning off caching outright) for code path'd through a whl?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From the fake Slim Shady:

Right, it doesn't, and the warning shouldn't have offered it. Every numba locator that reads NUMBA_CACHE_DIR needs the source file on disk, so for a wheel or an egg it changes nothing. The warning now names it only for a source file on disk: 105b9ce.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From the fake Slim Shady:

It does cache a wheel, on every numba I tried from 0.60 to 0.67. Two things keep me from recommending it. A rebuild only gets fresh code under a new file name: under the same name an edit to a global, a literal or a callee still ran the old code, since the cache key hashes only the function's own bytecode and the stamp is the interpreter's. And multiprocessing reads sys.frozen too: with it set a spawn child dies on unknown option --multiprocessing-fork, and spawn is the default on macOS and Windows.

A .zip gets a cache without the flag. numba 0.61 and later cache it and stamp it with the zip, so each of those edits got fresh code. A wheel is a zip, so renaming it is enough, and that's what the warning and the README recommend now. Trying it I found the import dies where the user's cache directory isn't writable, since numba doesn't check that one, so the fallback covers it: f45b510.

return None
if inferred is None:
return built.schema
if built.schema != inferred:

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder, how would this realistically happen, considering each partition and batch runs the same code that returns typed output?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From the fake Slim Shady:

With an ndarray it can't, the dtype fixes the type. It takes a column returned as a list or an object array, which is how a None gets out. pyarrow types those from the values, so a batch where every value is None comes out null next to double, and the last batch of a partition can be a single row. I ran a lookup that finds nothing in the first batch through Spark in batches of 4. On main it fails with Tried to write record batch with different schema, which names no column. With output_schema it passes.

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Got it, thanks.

@Goykhman

Copy link
Copy Markdown
Owner

Thanks for contributing, merging this.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants