performance optimizations - #218
Conversation
| @generated _concat(a::NTuple{N,Any}, b::NTuple{M,Any}) where {N,M} = | ||
| :(Base.Cartesian.@ntuple $(N + M) i -> i ≤ $N ? a[i] : b[i - $N]) |
There was a problem hiding this comment.
There is no reliably fast way to do this in Base?
There was a problem hiding this comment.
I tried:
map- broadcast
ntuplethe function
All of these suffer from catastrophic performance drop at some point. Seems like it's around 32 elements returned by getall in total, the unrolling threshold for tuples in Base.
Added a comment.
|
Ideally, Accessors should work and be (almost) zero cost in all regimes – from heterogeneous tuples at small sizes, to homogeneous abstractvectors after some threshold. Unfortunately, the seamless transition is notoriously difficult, don't think we ever seriously attempted it. |
Accessorsare known to be hard for the compiler to optimize sometimes, with compiler heuristics to stop optimizing triggering quite early.Here, I make the compiler "try harder": things like avoiding argument splatting, and making small inner functions more optimizable.
I recently used Accessors getall/setall with a few tens of elements – up to ~50. With these optimizations, it compiles down to almost nothing, while before there was a lot of dynamic dispatch inside.
Tests pass both before and after. Maybe I'll think of a short self-contained test demonstrating the improvements and add it... But I confirm in my testing that these changes greatly improve performance in such a regime.