Skip to content

feat(vector): fuse the BBQ rerank distance kernels - #351

Closed
EnRaiha wants to merge 1 commit into
mainfrom
feat/280-bbq-fused
Closed

EnRaiha wants to merge 1 commit into
mainfrom
feat/280-bbq-fused

Conversation

@EnRaiha

@EnRaiha EnRaiha commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Superseded by #391 — fork-based PR at EnRaiha:feat/280-bbq-fused-v2 (commit 42abd92f7). This branch lives upstream and the author's account has no push access to it.

Why

Rerank against 1-bit BBQ candidates materialized a Vec<f32> per candidate (and per query). At an oversample × ef candidate count that dominated the pass and defeated SIMD. This is the fused-kernel work of #280.

What

  • the kernels read the centered query straight from the prepared payload bytes: zero allocation per candidate and per query
  • the kernels join the dispatch table that already exists: l2_bbq is a SimdRuntime field in distance/simd/runtime.rs, selected in SimdRuntime::detect() beside l2_squared_f16 / l2_squared_bf16. Order: avx512 → avx2+fma → neon → wasm-simd128 (wasm32 + simd128 only) → scalar
  • every safe entry point validates both byte slices against dim before any pointer arithmetic. Reading the centered query out of raw payload bytes means a short slice would otherwise be a read past the end, and dispatch guarantees the host's features, never the shape of the data. One shared guard, bbq::assert_payload_shapes, is called by the scalar kernel and all four tiers; saturating_mul keeps dim * 4 from wrapping past the check
  • the wasm simd128 arm is wired into detect() rather than merely compiled: the arm's gate mirrors the tier module's own compile-time gate, since wasm has no runtime feature probe. A build without simd128 compiles both out and keeps the scalar fallback
  • every tier derives its lane mask with the same MSB-first mapping (byte.reverse_bits(), dim k → lane k), and every tail is handled by the scalar accumulator so a vector tail cannot diverge from the head formula
  • decode_payload / the old dequantize path are removed
  • the commits are rebuilt into one, not appended as fix-ups

Validation

Run under cargo nextest, which is what this repo requires.

  • golden parity against an f64 oracle over dims 0..768 — lane boundaries and partial tails included; relative tolerance 1e-4 with an absolute floor of 1e-6
  • cargo nextest run -p nodedb-vector --all-features --cargo-profile ci --profile ci → 466 passed, 0 failed
  • cargo fmt --all --check → clean
  • cargo clippy -p nodedb-vector --profile ci --all-targets -- -D warnings → clean
  • nodedb-preflight.sh <repo> 1ff35512b → pass
  • cargo check -p nodedb-vector --target aarch64-unknown-linux-gnu → clean (NEON tier compiles)
  • cargo check -p nodedb-vector --target wasm32-wasip1 with -C target-feature=+simd128 → clean, and without +simd128 → clean, so the arm and its import gate out and the scalar fallback stands
  • the length guard is pinned by tests checked by mutation, not by green alone: deleting the guard from avx2::l2_bbq makes simd_length_safety::bbq_rejects_short_slices die with SIGSEGV, and shifting the seam to &payload[0..] makes distance_prepared_matches_the_unfused_l2 fail
  • distance_prepared has a value test on a trained codec: it pins the result against an f64 reference built from the documented wire layout, not from the kernel
  • zero allocation is proven, not asserted: with fluxbench's TrackingAllocator installed, both fused benches report 0 bytes / 0 allocations, while the replaced path allocates one Vec<f32> per candidate

Benches against the replaced path (fluxbench, the repo norm): fused is 7.3× (dim 128) and 10.3× (dim 768) faster with allocation tracking installed (13.6× / 16.0× without it). These carry over from the pre-rework build — the rework changed dispatch wiring and module layout, not the kernel bodies or the bench harness — so they were not re-measured.

Known gaps (not covered by this PR)

  • AVX-512 correctness is validated under Intel SDE only; there is no native 512-bit hardware on the development host, so no real-silicon performance numbers. CI runs no SDE job, so on a host without AVX-512 that tier's test returns early and passes
  • the crate's own wasm unit test cannot build yet: tokio (via the fluxbench dev-dependency) does not build for wasip1, so the wasm tier is verified by a WASI-run probe rather than by a cargo test case
  • the wasm simd128 tier is reachable in principle but not in today's shipped artifact: nodedb-lite ships nodedb-lite-wasm and depends on this crate, but wasm32-unknown-unknown enables no simd128 by default and nodedb-lite sets no +simd128 RUSTFLAGS, so the tier compiles out and the browser product runs the scalar fallback
  • no per-tier benches: only the dispatched tier is measured natively
  • differential testing uses fixed seeds; there is no fuzz/property target for the kernel yet

Testing the 512-bit tiers locally (no 512-bit hardware needed)

Intel SDE emulates the 512-bit paths — it is the same release rust-lang pins in its stdarch CI:

curl -LO https://ci-mirrors.rust-lang.org/sde-external-10.8.0-2026-03-15-lin.tar.xz
tar xf sde-external-10.8.0-2026-03-15-lin.tar.xz -C /opt
CARGO_TARGET_X86_64_UNKNOWN_LINUX_GNU_RUNNER="/opt/sde-external-10.8.0-2026-03-15-lin/sde64 -spr --" \
  cargo nextest run -p nodedb-vector --lib distance::simd::bbq

Closes #280

Copilot AI lite review requested due to automatic review settings September 19, 2026 23:48

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@EnRaiha
EnRaiha force-pushed the feat/280-bbq-fused branch 2 times, most recently from f774f5a to 1ad096e Compare September 19, 2026 23:58
@EnRaiha EnRaiha added the run-ci Opt this PR into the full test suite; re-add to force a re-run label Sep 20, 2026

@farhan-syah farhan-syah left a comment •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The kernel work is strong and carries real evidence: a parity check against an f64 reference over dims 0..768, the 512-bit tiers run under Intel SDE, NEON run under qemu, and allocation tracking that shows zero bytes. That part stays.

Blocker: a second SIMD dispatch system. nodedb-vector already selects kernels once, through distance/simd/runtime.rs (SimdRuntime, runtime()). It also already holds fused kernels that decode and compute without an intermediate Vec<f32>: l2_squared_f16 and l2_squared_bf16. That is the exact precedent for this kernel. This PR adds its own OnceLock, its own feature detection, and its own kernel table under rerank/codecs. Two dispatch systems can pick different tiers for one host, and each new kernel must then choose which one to join.

Direction:

  • Add the BBQ distance as a SimdRuntime field, selected in SimdRuntime::detect() next to the f16/bf16 kernels.
  • Move the per-tier kernels beside their siblings in distance/simd/{avx512,avx2,neon}.rs.
  • The rerank codec calls runtime().<field>.

Should-fix: an unreachable tier (inline).

After the change, rebuild the commits rather than appending fix-up commits on top.

Comment on lines +79 to +80
pub(super) fn selected_kernel() -> &'static Kernel {
static KERNEL: OnceLock<Kernel> = OnceLock::new();

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the second dispatch system: its own OnceLock and its own feature detection, beside SimdRuntime in distance/simd/runtime.rs. Register the BBQ kernel as a SimdRuntime field, selected in detect(), the same way l2_squared_f16/l2_squared_bf16 are.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The second dispatch system is gone. Its OnceLock and kernel table are removed, and rerank/codecs/bbq_kernel.rs no longer exists.

l2_bbq is now a SimdRuntime field, pub l2_bbq: BbqFn, selected once per tier arm in detect() beside l2_squared_f16 / l2_squared_bf16. The per-tier kernels moved to distance/simd/{avx512,avx2,neon,wasm_simd128}.rs, the scalar reference to distance/simd/bbq.rs, and the codec calls (crate::distance::simd::runtime().l2_bbq)(…) at rerank/codecs/bbq.rs:178.

grep -rn "OnceLock\|detect_kernel\|bbq_kernel" nodedb-vector/src now matches only the pre-existing static RUNTIME. Rebuilt into one commit, f4b973cfd. Full evidence table in the timeline comment above.

Comment on lines +98 to +102
// Reserved for AVX10/256 hosts: AVX512VL implies AVX512F, so on
// today's feature bits this arm is reachable only when AVX10 detection
// lands in std::arch (the 256-bit masked kernel matches the AVX10.1
// shape). Its correctness is pinned by `avx10_256_tier_matches_the_reference`.
if std::is_x86_feature_detected!("avx512vl") && std::is_x86_feature_detected!("avx512f") {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This arm can never be selected. avx512vl implies avx512f, so any host that passes this check already returned the avx512 tier above. The comment says so too. A tier that production cannot reach is dead code in the dispatch. Remove it, or gate it on a real AVX10 detection once std::arch has one.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed, along with its test. git grep avx10 over nodedb-vector/ returns nothing, and the comment promising a later AVX10 gate went with the arm. The tier returns when std::arch can detect AVX10.

@EnRaiha

EnRaiha commented Sep 23, 2026

Copy link
Copy Markdown
Contributor Author

Reworked and rebuilt into one commit (f4b973cfd), based on main @ 1ff35512b.

Point Landed
Blocker: a second dispatch system The OnceLock, detect_kernel, and the Kernel table are gone; feat/280-bbq-fused now adds pub type BbqFn and pub l2_bbq: BbqFn to SimdRuntime, with one selection line per tier arm in detect(). grep -rn "OnceLock|detect_kernel|bbq_kernel" nodedb-vector/src returns only the pre-existing static RUNTIME
Per-tier kernels beside their siblings distance/simd/{avx512,avx2,neon}.rs carry l2_bbq (safe entry) + l2_bbq_impl (#[target_feature]); the wasm kernel is in distance/simd/wasm_simd128.rs; the scalar reference and the shared helpers are in distance/simd/bbq.rs
The rerank codec calls runtime().<field> rerank/codecs/bbq.rs:178 — (crate::distance::simd::runtime().l2_bbq)(…)
Should-fix: an unreachable tier Removed, with its test. No avx10_256 remains; the comment promising AVX10 detection went with the arm
Rebuild the commits One commit, message rewritten to describe the final shape

Evidence, on the rebased tree:

Check Result
cargo test -p nodedb-vector --profile ci bbq 18 passed, 0 failed (was 19 — the difference is the removed tier's test)
cargo test -p nodedb-vector --profile ci --lib 433 passed, 0 failed
cargo check --all-targets 0 warnings, 0 errors
-C target-feature=+simd128 for wasm32-wasip1 clean — the wasm kernel compiles
Parity unchanged: the f64 oracle over dims 0..768, plus dispatch_selects_a_tier_that_matches_the_reference over the whole table

Two things I fixed in passing that the review would have caught: the whole-range scalar kernel needs .sqrt() (the tier kernels return a distance, the accumulation returns a squared sum), and #[target_feature] belongs on the unsafe impl rather than the safe wrapper — both now match the f16/bf16 shape.

Not run here: the AVX-512 tier under Intel SDE and NEON under qemu (this host is AVX2-only; the avx512 test prints its skip line), and the bench numbers in benches/bbq_kernel.rs were not re-measured.

…table

Reranking against 1-bit BBQ candidates materialized a `Vec<f32>` per candidate
and per query; at an oversample x ef candidate count that dominated the pass and
defeated SIMD.

The kernel now reads the centred query straight from the prepared payload bytes:
zero allocation per candidate and per query, one pass.

It is a `SimdRuntime` field (`l2_bbq`), selected in `detect()` beside the
f16/bf16 fused kernels — one dispatch system, one place that decides a tier for a
host. The per-tier kernels sit with their siblings in
`distance/simd/{avx512,avx2,neon,wasm_simd128}.rs`, the scalar reference in
`distance/simd/bbq.rs`, and the rerank codec calls `runtime().l2_bbq`.

The wasm simd128 arm is wired into `detect()` and not merely compiled: on
`wasm32` + `simd128` the dispatcher installs `wasm-simd128` instead of falling
through to the scalar arm. The arm's guard mirrors the tier module's own
compile-time gate, since wasm has no runtime feature probe.

Evidence: parity against an f64 oracle over dims 0..768
(`distance/simd/bbq.rs`), 18 tests pass; the whole detection table is covered by
`dispatch_selects_a_tier_that_matches_the_reference`, which calls through
`runtime().l2_bbq` rather than a tier directly. The 512-bit tier has no hardware
in this fleet and runs under Intel SDE; NEON under qemu; the wasm tier is
executed under Node's WASI runner, where a probe asserts the dispatched name is
`wasm-simd128` and that 14 dims match the oracle through the dispatched pointer.

The AVX10/256 arm the first revision carried is gone: `AVX512VL` implies
`AVX512F`, so no host could reach it while AVX10 detection is absent from
`std::arch`. Its test went with it.
@EnRaiha

EnRaiha commented Sep 24, 2026

Copy link
Copy Markdown
Contributor Author

One more fix before re-review, pushed as an amended commit (7f3998407, rebased shape unchanged): an audit of the wasm tier found it had the same defect you flagged on the avx10-256 arm — a tier that compiles and is tested but that production could not select.

detect() had no wasm32 arm, so a wasm32 + simd128 build fell through to the scalar fallback (l2_bbq: bbq::l2_bbq). The simd128 kernel was reachable only from its own test, which calls the tier module directly rather than the dispatch table.

Fixed in distance/simd/runtime.rs:

  • the arm now installs wasm-simd128 when targeting wasm32 with simd128, placed after neon and before the scalar fallback
  • its guard is the same target_feature = "simd128" condition as the tier module, since wasm has no runtime feature probe to consult

This is 25 lines in one file. The arm is verified by execution, not by compilation alone: a probe built for wasm32-wasip1 under Node 22's WASI runner reports dispatch name: wasm-simd128 and matches the f64 oracle across 14 dims (1, 7, 8, 9, 63, 64, 65, 127, 128, 129, 255, 256, 512, 768) when calling through rt.l2_bbq — the dispatched pointer, not the tier directly.

On the rebased tree:

Check Result
cargo test -p nodedb-vector --profile ci bbq 18 passed, 0 failed
cargo test -p nodedb-vector --profile ci --lib 433 passed, 0 failed
cargo clippy -p nodedb-vector --lib 0 warnings
cargo check -p nodedb-vector --all-targets 0 warnings, 0 errors
--target aarch64-unknown-linux-gnu clean
--target wasm32-wasip1 + +simd128 clean
--target wasm32-wasip1 without +simd128 clean — the arm and its import gate out, scalar fallback stands

Two things to flag plainly:

  1. This is a force-push, so the inline threads you opened now read as outdated — the file they were anchored to, rerank/codecs/bbq_kernel.rs, no longer exists and the commit they pointed at is gone. Both are answered in this thread.
  2. The bench figures in the description (7.3× / 10.3×) are carried over from the pre-rework build and are marked as not re-measured. The rework changed dispatch wiring and module layout, not the kernel bodies or the bench harness.

The description has been rewritten to match the current tree. Ready for another look.

@farhan-syah farhan-syah left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes. There is one blocker, plus a set of smaller fixes.

What I checked: I merged this head onto current main and ran:

  • cargo fmt --all --check: clean
  • cargo clippy -p nodedb-vector --all-targets -- -D warnings: clean
  • cargo nextest run -p nodedb-vector: 501/501 passed

The repo requires cargo nextest run, not cargo test. Use nextest for the validation you list.

What landed well:

  • One dispatch table: l2_bbq is a field on SimdRuntime, not a second feature detection.
  • Every tier uses the same MSB-first lane mapping and finishes its tail with the shared scalar helper.
  • The per-candidate allocation is gone.
  • The wasm arm's gate matches the module's own gate.

Blocker: the public safe tier functions (avx2, avx512, neon, wasm_simd128 l2_bbq) read out of bounds when given short slices. Details are inline.

Also inline: misplaced doc and #[inline] attributes, two tests that cannot fail as written, a false CI claim in the module doc, issue references in comments, and a missing value test for distance_prepared.


use super::bbq::{l2_scalar_from_bytes, recon_scale};
/// Safe entry for `SimdRuntime`; the feature guard lives in `SimdRuntime::detect`.
pub fn l2_bbq(centered: &[u8], packed: &[u8], residual_norm: f32, dim: usize) -> f32 {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocker: this is a public safe function, but the SIMD loop reads through loadu/get_unchecked on the assumption that centered.len() >= dim * 4 and packed.len() >= dim.div_ceil(8). Nothing checks either length. nodedb_vector::distance::simd is public and so is SimdRuntime::l2_bbq, so safe code such as l2_bbq(&[], &[], 1.0, 16) reads out of bounds, which is undefined behaviour. The existing kernels in this file guard their entry point (assert_eq!(a.len(), b.len(), "avx2 l2: length mismatch")). Check both slice lengths against dim here the same way, before the unsafe block, and in every tier.


use super::bbq::{l2_scalar_from_bytes, recon_scale};
/// Safe entry for `SimdRuntime`; the feature guard lives in `SimdRuntime::detect`.
pub fn l2_bbq(centered: &[u8], packed: &[u8], residual_norm: f32, dim: usize) -> f32 {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocker: this is a public safe function, but the SIMD loop reads through loadu/get_unchecked on the assumption that centered.len() >= dim * 4 and packed.len() >= dim.div_ceil(8). Nothing checks either length. nodedb_vector::distance::simd is public and so is SimdRuntime::l2_bbq, so safe code such as l2_bbq(&[], &[], 1.0, 16) reads out of bounds, which is undefined behaviour. The existing kernels in avx2.rs guard their entry point (assert_eq!(a.len(), b.len(), "avx2 l2: length mismatch")). Check both slice lengths against dim here the same way, before the unsafe block, and in every tier.


use super::bbq::{l2_scalar_from_bytes, recon_scale};
/// Safe entry for `SimdRuntime`; the feature guard lives in `SimdRuntime::detect`.
pub fn l2_bbq(centered: &[u8], packed: &[u8], residual_norm: f32, dim: usize) -> f32 {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocker: this is a public safe function, but the SIMD loop reads through loadu/get_unchecked on the assumption that centered.len() >= dim * 4 and packed.len() >= dim.div_ceil(8). Nothing checks either length. nodedb_vector::distance::simd is public and so is SimdRuntime::l2_bbq, so safe code such as l2_bbq(&[], &[], 1.0, 16) reads out of bounds, which is undefined behaviour. The existing kernels in avx2.rs guard their entry point (assert_eq!(a.len(), b.len(), "avx2 l2: length mismatch")). Check both slice lengths against dim here the same way, before the unsafe block, and in every tier.

/// 4-lane tier (`simd128`): `v128_bitselect` between `+scale` and `-scale`.
#[cfg(all(target_arch = "wasm32", target_feature = "simd128"))]
use super::bbq::{l2_scalar_from_bytes, recon_scale};
pub fn l2_bbq(centered: &[u8], packed: &[u8], residual_norm: f32, dim: usize) -> f32 {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same out-of-bounds read as the x86 tiers: v128_load reads centered without a length check. Check centered.len() and packed.len() against dim at entry. Also:

  • The doc comment and #[cfg] on line 236 attach to the use on line 237, not to this function. The #[cfg] is redundant because the module already has #![cfg(...)].
  • Move the function above mod tests.
  • The per-lane scalar loop that builds mask runs on every step. Precompute the masks in a table, as the AVX2 tier does.

_mm_cvtss_f32(sums2)
}

use super::bbq::{l2_scalar_from_bytes, recon_scale};

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Move this use to the top of the file with the other imports. The #[cfg(all(target_arch = "x86_64", ...))] on lines 153/158/161 repeats the module's own #![cfg(target_arch = "x86_64")]. Remove it. (The same applies to the mid-file use in avx512.rs and neon.rs.)

//!
//! The scalar kernel here is the reference; every SIMD tier in the sibling
//! modules must agree within `PARITY_REL` (see the tests below). The 512-bit
//! tier has no native hardware in this fleet: it is exercised under Intel SDE

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

.github/ has no SDE job, so "exercised under Intel SDE and the crate's CI" is false. On an AVX2-only runner, avx512_tier_matches_the_reference returns early and reports PASS. State what is true: the tier is checked manually under SDE, and CI skips it. "in this fleet" is also local context that does not belong in a crate doc.

const PARITY_REL: f64 = 1e-4;
const PARITY_ABS: f64 = 1e-6;

/// Dimensions under test: the 8..512 range the issue names, plus the

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remove the issue references from the code comments ("the range the issue names" here, and on lines 111 and 167). A comment must stand on its own after the issue closes. "32- and 8-lane tiers" is also wrong: AVX-512 f32 has 16 lanes, AVX2 8, and NEON/wasm 4.


/// The issue's acceptance range: every dim in `DIMS` against the oracle.
#[test]
fn every_available_kernel_matches_the_reference() {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This test calls only the scalar l2_bbq, so it duplicates the_scalar_kernel_always_matches_the_reference. Its name claims coverage of every kernel. Either run every tier compiled for the target, or remove it.

// The issue asks for parity against the scalar kernel directly.
let scalar = l2_bbq(&centered, &packed, residual_norm, dim);
assert!(
within_parity(got, scalar as f64) || within_parity(got, expected),

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This assert can never fail. within_parity(got, expected) already passed in the assert above, so the || within_parity(got, expected) branch is always true. Drop the || so the tier is actually checked against the scalar kernel.

.sum::<f32>()
.sqrt();
Ok(dist)
// Fused and allocation-free: the kernel reads the centered query

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No test pins the value distance_prepared returns: the existing bbq rerank tests cover only the error paths and the round-trip. The old decode-and-dequantize path is removed here, so add a test on a trained codec. It must check that distance_prepared equals the unfused L2 (centred query vs. ±residual_norm/√dim reconstruction) within the parity tolerance.

@EnRaiha

EnRaiha commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

Fixed, rebuilt into one commit. All ten points are addressed. The revised commit is ready but I cannot push it — see the last section.

Blocker — out-of-bounds reads on short slices. Reproduced before fixing: an external safe caller, nodedb_vector::distance::simd::avx2::l2_bbq(&[], &[], 1.0, 16), dies with signal: 11, SIGSEGV (exit 101), while the same call on well-formed buffers returns 35.284557. The four // SAFETY: comments asserted a caller guarantee that no caller provided — dispatch guarantees the host's features, not the shape of the data.

Fixed with one shared guard, bbq::assert_payload_shapes, called by the scalar kernel and all four tiers before any pointer arithmetic. It checks centered.len() >= dim * 4 and packed.len() >= dim.div_ceil(8), with saturating_mul so a plain dim * 4 cannot wrap to a small value and slip past the check. Each tier has a should_panic case on that guard, and the crate's own simd_length_safety suite gains bbq_rejects_short_slices, exercising the dispatched kernel from outside the crate the way its existing length-parity cases do.

Both new tests were checked by mutation rather than by green alone: deleting the guard from avx2::l2_bbq makes the suite case die with SIGSEGV, and shifting the seam to &payload[0..] makes the new value test fail (1.2574875 vs 0.2060269).

Also fixed, per your inline notes:

  • the mid-file use moved to the top in avx2.rs, avx512.rs, neon.rs, wasm_simd128.rs; the three redundant #[cfg] at avx2.rs:153/158/161 removed
  • wasm: doc and gate moved onto the function (the use no longer carries them), function moved above mod tests, per-step mask loop replaced with a const fn-built table as in the AVX2 tier
  • bbq.rs: the #[inline] and "Scalar accumulation…" doc moved to l2_scalar_from_bytes; the false CI claim replaced with what is true — checked manually under SDE, CI runs no SDE job, and on a host without AVX-512 the test returns early and passes; "in this fleet" dropped
  • issue references removed from the three comments; lane counts corrected to 16 (AVX-512 f32) / 8 (AVX2) / 4 (NEON, wasm)
  • every_available_kernel_matches_the_reference removed — you were right that it called only the scalar kernel and duplicated the_scalar_kernel_always_matches_the_reference while its name claimed every kernel
  • the || dropped, so the tier is genuinely checked against the scalar kernel now
  • distance_prepared_matches_the_unfused_l2 added: it pins the returned value against an f64 reference built from the documented wire layout (32-byte header, residual_norm at 8..12, then the sign bytes), not from the kernel

Runner. Correct — all evidence is now under cargo nextest. On the revised commit: cargo nextest run -p nodedb-vector --all-features --cargo-profile ci --profile ci → 466 passed, 0 failed; cargo fmt --all --check clean; cargo clippy -p nodedb-vector --profile ci --all-targets -- -D warnings clean; nodedb-preflight.sh pass; aarch64-unknown-linux-gnu and wasm32-wasip1 both with and without +simd128 check clean.

Your point about preflight landed well: it fails on the current head (three prose issue references) and passes on the revised commit.

Description fix. .github/ has no SDE job, so the claim that the 512-bit tier runs "under Intel SDE and the crate's CI" was wrong. The description now states it as manual and uses nextest for the validation list.

Blocker on my side, not in the code. I cannot push the revision: this account has permissions.push: false on NodeDB-Lab/nodedb (HTTP 403), and the branch lives upstream rather than in a fork, so there is no fork I can force-push to. The revised commit is 42abd92f7 — parent 1ff35512b, the same parent as the current head, so it is a clean rebuild rather than a rebase. A verified git bundle and a format-patch are in a handoff archive for whoever has write access:

git fetch /path/to/nodedb-pr351-fix.bundle 'refs/heads/feat/280-bbq-fused:refs/heads/pr351-fix'
git push --force-with-lease=feat/280-bbq-fused:7f39984075ebb1c6273fc29489aeeaf031757ce6 \
  origin refs/heads/pr351-fix:refs/heads/feat/280-bbq-fused

The bundle verifies, and the patch was checked to reproduce the same tree (7ac790de7afe…) from a clean checkout of 1ff35512b.

Question — the wasm tier (your call, no change needed for approval). Kept and fixed as you asked, but its reachability is worth deciding. nodedb-lite ships nodedb-lite-wasm (wasm-pack build --target web, plus Node and headless-Chrome suites) and depends on nodedb-vector = "0.5", so the crate really is built for wasm32. However rustc --print cfg --target wasm32-unknown-unknown (1.96.1) enables no simd128, and nodedb-lite sets no +simd128 RUSTFLAGS — so the tier compiles out of today's shipped wasm build and the browser product runs the scalar fallback. That is the same shape as the avx10-256 arm you had me remove. I left it in because you flagged the gate as landing well; say the word and I will delete the arm instead.

@EnRaiha EnRaiha changed the title feat(vector): fused BBQ dequantize-distance kernels with tiered SIMD dispatch feat(vector): fuse the BBQ rerank distance kernels Sep 28, 2026
@farhan-syah

Copy link
Copy Markdown
Member

you can just create new branch on your own fork, and create new PR for that, closing this PR now.

@EnRaiha

EnRaiha commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

Thanks — done in the fork style.

New PR: #391, EnRaiha:feat/280-bbq-fused-v2 → main, single commit 42abd92f7. This branch lives upstream and I have no push access to it, so the work moved to the fork as you suggested.

#391 carries the whole fused-kernel change plus every point from the 2026-09-27 review:

  • the blocker — a shared length guard (bbq::assert_payload_shapes) on all four tier entry points and the scalar kernel, with a should_panic case per tier and a case in the crate's own simd_length_safety suite so the dispatched kernel is checked from outside the crate
  • the misplaced doc/#[inline], the mid-file use and redundant #[cfg], the wasm doc/gate attachment, the function moved above mod tests, and the per-step mask loop replaced with a const fn table
  • the false CI claim corrected, issue references removed, lane counts fixed to 16/8/4
  • the duplicate test removed and the dead || dropped, so the tier is genuinely checked against the scalar kernel
  • a value test for distance_prepared against the documented wire layout
  • cargo nextest for all evidence

Two things worth flagging:

  • the two new tests were verified by mutation, not by green alone: deleting the guard from avx2::l2_bbq makes the suite case die with SIGSEGV — the original blocker reproducing itself — and shifting the seam to &payload[0..] makes the value test fail.
  • on the wasm tier: kept and fixed as you asked, but its reachability is worth your call. nodedb-lite ships nodedb-lite-wasm and depends on this crate, so it is built for wasm32 — but wasm32-unknown-unknown enables no simd128 by default and nodedb-lite sets no +simd128 RUSTFLAGS, so the tier compiles out of today's shipped wasm build and the browser product runs the scalar fallback. Same shape as the avx10-256 arm you had me remove. Say the word and I will delete the arm in feat(vector): fuse the BBQ rerank distance kernels #391 instead.

Closing this one is fine by me.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-ci Opt this PR into the full test suite; re-add to force a re-run

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Fused BBQ dequantize-distance kernels with tiered SIMD dispatch

3 participants