Skip to content

Feat/m4 grouping sets and windows - #32

Merged
9bany merged 5 commits into
masterfrom
feat/m4-grouping-sets-and-windows
Sep 7, 2026
Merged

9bany merged 5 commits into
masterfrom
feat/m4-grouping-sets-and-windows

Conversation

@9bany

@9bany 9bany commented Sep 7, 2026

Copy link
Copy Markdown
Member

No description provided.

Groundwork for GROUPING SETS. A rollup's subtotal row has no value in the
columns its set does not name, and `grouping(x)` says so as a 0/1 - two
spellings of one fact, so one variant carries both rather than two that
could come to disagree about which branch is a subtotal.

Names no plan, so `rebase` leaves it alone, like `Of::Now`.

The renderers are the trap: `Of` is read by three, and two of them end in
`_ => None`, so a constant missing an arm in either would render as its
number in one branch of a statement and as `null` in another with nothing
in the answer able to show it. All three now go through one `const_of`,
which is where that reason is written down.

Also closes a latent gap while here: `Shape::plans` reported no driving
plans for a `Union`, so a key-only branch's `keys` went unnamed. Nothing
called it yet; grouping sets make unions of key-only branches ordinary.
One grouping per set, stacked. Every call a rollup makes is one the planner
already had - `ROLLUP(category, country)` is a `GroupByTuple`, a `GroupBy`
and a `Count` - so this adds no `Plan` variant and no merge arm, and nothing
below `big-sql` changed. A claim test asserts that rather than commenting it.

The three spellings normalise at parse time into subsets of the written
`GROUP BY` list, the way `SELECT DISTINCT` already normalises into one: one
path to the plans instead of three that have to agree.

Deliberately not sharing intermediates between sets. Folding a
`GroupBy(country)` out of an already-computed `GroupByTuple(country, city)`
would be a per-set rollup in the merge, which is the one thing this surface
does not add. `MAX_GROUPING_SETS` is where that price is named instead of
hidden - eight sets, which is `CUBE` over three columns or `ROLLUP` over
seven, and `CUBE` over four is refused by name.

Three refusals, each a decision rather than a gap:

  * `sql_rollup_order` - `ORDER BY`/`LIMIT` over a grouping-sets answer.
    The rows are several groupings rendered one after the other, so there is
    no single list to sort or cut; ordering each set on its own would look
    sorted and not be, and `LIMIT 10` would answer ten rows per set.
  * `sql_with_totals` - the total beside the rows rather than among them. A
    result set here is columns and rows, one definition every format renders
    from. `WITH ROLLUP` is the same number as a row.
  * `sql_too_many_grouping_sets` - the fan-out bound, checked at the text.

With no `ORDER BY` to override it, the row order is a contract: sets longest
first, each branch in its own key order. Pinned in the corpus and in a
render test, which also pins that a tuple branch orders by key string.

The select-list rule is relaxed, not dropped: a bare `b` under
`GROUP BY a, b WITH ROLLUP` is legal because the sets that omit it render it
as null, and that check moved to the written list. A column no set could
name is still refused.
Sixteen functions across three families - ranking, offset, and aggregate
windows - computed at the coordinator over the rows a projection already
materialised. `Plan::Project` and `merge_projected` are untouched: no `Plan`
variant, no merge arm.

The cost is real and is said out loud rather than discovered. A window has
to see every row of its partition before it knows any one row's number, so
the `LIMIT` leaves the plan: `SELECT c, row_number() OVER (ORDER BY c) FROM
t LIMIT 10` reads every record matching the `WHERE` where the same statement
without the window reads ten. The corpus pins that pair side by side, and
what bounds the read is then the record ceiling every unbounded read答s to.

The load-bearing change is that a projection's columns and its plan's fields
are no longer one list. A window reads a column to partition, order or fold
by and that column need not be in the header, so `Selected` now names the
field it reads by position and the render arm indexes rather than zips.
Deduplicated, so `SELECT amount, row_number() OVER (ORDER BY amount)` reads
`amount` once.

Two families, opposite answers about an `ORDER BY`, each with its own
reason. A ranking or an offset needs one - `row_number()` over an unordered
partition is a number nobody can predict or reproduce. An aggregate must not
have one, because an ordering there means the running total, whose frame is
`RANGE UNBOUNDED PRECEDING` and whose answer is a different number from the
partition total. Both earn `sql_window_frame`, as do `ROWS`/`RANGE`/
`GROUPS`/`EXCLUDE`.

`last_value` is the partition's last row, not the current one. With no frame
there is no other reading, and the standard default frame's answer is one
nobody wants and everybody is surprised by.

`Refused::Window` is repurposed rather than retired: it used to mean "window
functions are not supported" and shared `sql_unsupported`. It now means the
shapes a window cannot be over - a grouped answer, a tuple grouping, a join,
`SELECT *` - under its own `sql_window_shape`, and names the rankings those
already carry. Changing a code is client-visible; it is pinned in the corpus.

Also refused by name: `QUALIFY`, and `WINDOW w AS (...)`/`OVER w`.

A window over a view works, because a view is substituted before any of this
is planned - the remap now covers the columns the clause names.
None of these is a gap. Each is a question with an answer here, refused by
name so a client reads the answer instead of a sentence about a missing
function - the same decision `TTL`, `Nullable` and the `bitmap_*` family
were already settled by.

  * sql_sketch - a sketch has nothing to approximate. `count(DISTINCT x)`
    is a bitmap's cardinality, a popcount, exact and paid per container;
    `uniq*`, `approx_count_distinct`, `quantile` and `median` are all
    accepted as written and answered exactly. A `-State`/`-Merge` pair
    carries a partial sketch between queries elsewhere; the partial result
    that travels between nodes here *is* the bitmap. Also the `HLL` and
    `QUANTILE_STATE` column types, separately from `BITMAP`: a bitmap is
    declined because every column already is one, a sketch because there is
    nothing to approximate.
  * sql_no_regex - the pattern language is `LIKE`'s, and it runs over a
    keyed column's dictionary rather than over records, so `LIKE 'G%'` costs
    the field's cardinality once where a row engine pays it per record. The
    decision not to take a regex dependency was already written in
    `big_db::like`; this is where a client hears it.
  * sql_no_analyze - the planner is purely syntactic, so statistics would
    have no cost model to feed. The zone maps that do exist are written as
    facts are, never stale, and already readable in `system.parts`.
    `EXPLAIN ANALYZE` reaches this by recursion, which is the reading that
    makes the sentence worth writing.
  * sql_no_optimize - compaction is whole-file because there are no parts to
    merge, and is a server operation because it needs the file's exclusive
    lock, which a statement inside a read does not hold.

`SIMILAR TO` moves from the shared `sql_unsupported` to `sql_no_regex`, in
both corpora that pinned it.

One list, two callers, for the regex names: the select list reaches it
through `unsupported_call`, and a `WHERE` reaches it directly - a `WHERE`
refuses *known* scalar functions by name, and these are not known ones, so
without the second call site `match(c, '^a')` failed on the bracket.
Through the whole path this time - plan, fan-out to the shards, merge,
render - because that is where the claims live that a translation test
cannot reach.

For grouping sets the claim is arithmetic: each set is a separate question
planned and merged on its own, so the only thing making a subtotal agree
with the detail above it is that both counted the same records. GB's two
cities sum to GB's subtotal, and the three countries sum to the grand
total; a set planned or merged wrongly would show up here as a number that
does not add up.

For windows the claim is that a partition is whole. These records are
spread across shards by record id, so a window that ran per node would
number each shard's rows from one - GB's 100, 300 and 500 come back as 1, 2
and 3 wherever they live.

And the cost, asserted rather than described: `LIMIT 2` over a descending
ranking answers with rank 6 on the first row, which is only possible if the
window saw all six records before the cut.
@9bany
9bany merged commit 055fc50 into master Sep 7, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant