Repository navigation
build(deps): upgrade alcatraz to v0.14.0 - #5
Conversation
Ten releases behind (v0.4.0 -> v0.14.0). No API changes: the scanner only
uses NewEngine, Options{Entities,Threshold,AllowList} and Analyze, all
unchanged, so no scanner code moves. Still dependency-free - the new NER
model support lives in the separate alcatraz/ner submodule.
Detection improves at the action's default threshold of 0.8:
- international and Brazilian phone patterns are recognised at all (v0.7.0)
- entity spans land on word boundaries, so a parenthesised US number is
captured whole instead of clipped (v0.12.0)
- matches near a labelling word score higher, so a labelled phone (0.85)
and ip (0.95) now clear 0.8 where they previously scored 0.50 and 0.60
(v0.14.0)
Existing users will therefore see more findings on unchanged code: better
recall, not a regression.
Also corrects the README threshold guidance, which claimed phone numbers
always score 0.5 - true now only for unlabelled ones.
PR Summary by QodoUpgrade alcatraz dependency to v0.14.0 and align README threshold guidance
AI Description
Diagram
High-Level Assessment
Files changed (3)
|
Code Review by Qodo
1. IP_ADDRESS now exceeds default
|
alcatraz v0.14.0 installs a context enhancer on every engine, so a match near a word naming its entity type scores up to +0.35. That is the right default - it is how a labelled phone or ip reaches the 0.8 CI threshold - but it silently changes what an unchanged repo reports, and a caller who reads 0.8 as "checksum validated" has no way back to pattern-only scores. Mirror the escape hatch upstream already ships on both its CLIs: a -context flag defaulting to true, wired to Engine.SetContextEnhancer(nil), surfaced as the context-scoring action input. disableContext() matches the helper in alcatraz's own cmd/alcatraz/scan.go so the two stay in step. Adds a regression test pinning the boost at threshold 0.8 for a labelled email and ip - both drop out with context off - plus a CI step covering the flag end to end.
🪨 Alcatraz PII scanNo PII detected in this PR's diff. ✅ |
The self-scan job caught the phone and ip values the previous commit added to the tests, action.yml and the CI step - correctly, and at exactly the boosted scores this PR is about (0.85 and 0.95). They are reserved-range synthetic values, so they belong in .pii-allowlist next to the card, email and ssn fixtures already there.
Upgrades
github.com/hoophq/alcatrazfromv0.4.0tov0.14.0— ten releases behind, most of them shipped in the last two weeks.Why it's safe
No API breaks. The scanner only touches
NewEngine,Options{Entities,Threshold,AllowList}andAnalyze, all unchanged, so no scanner code moves — the diff isgo.mod/go.sumplus one README correction.Verified locally:
go build ./...,go vet ./...,go test ./...— all passalcatraz/nersubmodule and adds no weight hereWhat it buys
Detection output diffed between the two versions on a 10-line PII sample. At the action's default
threshold: 0.8, v0.4.0 flagged 3 lines and v0.14.0 flags 5:+55 11 98765-4321PHONE_NUMBER0.50phone: (415) 555-2671ip 192.168.1.44license AB1234567MEDICAL_LICENSE0.40US_DRIVER_LICENSE0.65The headline is v0.14.0's context-aware scoring: a match near a labelling word scores higher, so genuine PII that used to sit under the
0.8CI default now surfaces.Behavior change to expect
Repos already using this action will likely see more findings on unchanged code. That's improved recall rather than a regression, but it's worth expecting rather than being surprised by — hence
minorand notpatch.The
0.8default is unchanged; it's still the right precision/recall balance and now catches strictly more real PII.Docs
The README claimed phone numbers score
0.5and needthreshold: 0.4to catch. With context scoring that's now true only for unlabelled ones, so the passage is corrected. Checked the other version-sensitive claims: "45 entity types across 12 countries" still holds in both the README andaction.yml— the new phone patterns extendedPHONE_NUMBERrather than adding entity types (confirmed viaSupportedEntities: 45 in both versions).New:
context-scoringopt-out (added after review)Review flagged that context scoring is installed on every engine with no way off, which leaves callers who read
threshold: 0.8as "checksum validated" with no route back to pattern-only scores. Fair, so the escape hatch upstream already ships on both its CLIs is now mirrored here:context-scoringtruecontext-scoring: falseSetContextEnhancer(nil)— scores on pattern strength aloneAt
threshold: 0.8:disableContext()deliberately mirrors the helper in alcatraz's owncmd/alcatraz/scan.goso the two implementations don't drift.Review also proposed adding
IP_ADDRESSto the defaultignore-entities. Declined — the precedent cited is alcatraz's hook CLI (a live prompt guard, where your own LAN addresses are constant noise). The right analogue is its scan CLI, whoseignoredefault isDATE_TIME,URL— already identical to this action's. IPs are also personal data under GDPR Art. 4(1), so suppressing them by default would be a silent detection regression. Both opt-outs remain one line for anyone who wants them quiet. Full reasoning in the review thread.Test coverage
TestContextScoringpins the boost atthreshold: 0.8for a labelled email and IP, asserting each drops out with context off-context=true/falseend to end through the built binary