Conversation
Signed-off-by: Adam Gutglick <adam@spiraldb.com>
Signed-off-by: Adam Gutglick <adam@spiraldb.com>
|
I think |
|
@AdamGS thank you for this, and for the kind words!
It's a little worse than ordering unfortunately, the rows themselves are different. DataFusion 55 folds the So part 1 becomes the 1561 smallest ids instead of the first 1561 rows in scan order (11 of 1561 overlap with the reference). The zone parts are benchmark data that let df = partition.apply_to_dataframe(df)?;
// Materialize the LIMIT before the window step. DataFusion otherwise folds it into a
// TopK over the ORDER BY id sort, which changes which rows land in this part.
let schema = Arc::clone(df.schema().inner());
let batches = df.collect().await?;
let df = ctx.read_table(Arc::new(datafusion::datasource::MemTable::try_new(
schema,
vec![batches],
)?))?;With that, all six Do you happen to know if that TopK rewrite is intentional upstream? It also defeats early termination, so it might be worth an issue. I'll open a separate one here for making zone generation not depend on the optimizer at all, but definitely not for this PR! |
|
There are a few DF issues/PRs around this sort of stuff, like:
Might be worth trying to backport them to 55.2 and try this upgrade again then. |
This PR updates the Apache Arrow and DataFusion to their most recent releases - 59 and 55 (respectively). This change also required updating the rust-toolchain to one that satisfies their MSRV.
Also wanted to say thank you for this crate! Its been extremely useful for benchmarking our geospatial work!