Skip to content

build(deps): upgrade alcatraz to v0.14.0 - #5

Merged
luanlorenzo merged 3 commits into
mainfrom
luanlorenzo/update-alcatraz-version
Aug 11, 2026
Merged

luanlorenzo merged 3 commits into
mainfrom
luanlorenzo/update-alcatraz-version

Conversation

@luanlorenzo

@luanlorenzo luanlorenzo commented Jul 31, 2026 •

Copy link
Copy Markdown
Contributor

Upgrades github.com/hoophq/alcatraz from v0.4.0 to v0.14.0 — ten releases behind, most of them shipped in the last two weeks.

Why it's safe

No API breaks. The scanner only touches NewEngine, Options{Entities,Threshold,AllowList} and Analyze, all unchanged, so no scanner code moves — the diff is go.mod/go.sum plus one README correction.

Verified locally:

  • go build ./..., go vet ./..., go test ./... — all pass
  • both CI end-to-end checks pass: fixture log → 2 findings, allowlist → 0 findings
  • still dependency-free; the new NER/model work lives in the separate alcatraz/ner submodule and adds no weight here

What it buys

Detection output diffed between the two versions on a 10-line PII sample. At the action's default threshold: 0.8, v0.4.0 flagged 3 lines and v0.14.0 flags 5:

Input v0.4.0 v0.14.0 Upstream change
+55 11 98765-4321 missed entirely PHONE_NUMBER 0.50 intl/BR phone patterns (v0.7.0)
phone: (415) 555-2671 0.50, span clipped 0.85, full span word-boundary fix (v0.12.0) + context scoring (v0.14.0)
ip 192.168.1.44 0.60 0.95 context scoring (v0.14.0)
license AB1234567 MEDICAL_LICENSE 0.40 0.75, plus US_DRIVER_LICENSE 0.65 context scoring (v0.14.0)

The headline is v0.14.0's context-aware scoring: a match near a labelling word scores higher, so genuine PII that used to sit under the 0.8 CI default now surfaces.

Behavior change to expect

Repos already using this action will likely see more findings on unchanged code. That's improved recall rather than a regression, but it's worth expecting rather than being surprised by — hence minor and not patch.

The 0.8 default is unchanged; it's still the right precision/recall balance and now catches strictly more real PII.

Docs

The README claimed phone numbers score 0.5 and need threshold: 0.4 to catch. With context scoring that's now true only for unlabelled ones, so the passage is corrected. Checked the other version-sensitive claims: "45 entity types across 12 countries" still holds in both the README and action.yml — the new phone patterns extended PHONE_NUMBER rather than adding entity types (confirmed via SupportedEntities: 45 in both versions).

New: context-scoring opt-out (added after review)

Review flagged that context scoring is installed on every engine with no way off, which leaves callers who read threshold: 0.8 as "checksum validated" with no route back to pattern-only scores. Fair, so the escape hatch upstream already ships on both its CLIs is now mirrored here:

Input Default Effect
context-scoring true Matches upstream and current behavior
context-scoring: false SetContextEnhancer(nil) — scores on pattern strength alone

At threshold: 0.8:

context-scoring: true    phone 0.85 ✓   ip 0.95 ✓   card 1.00 ✓    → 3 findings
context-scoring: false   pattern-only               card 1.00 ✓    → 1 finding

disableContext() deliberately mirrors the helper in alcatraz's own cmd/alcatraz/scan.go so the two implementations don't drift.

Review also proposed adding IP_ADDRESS to the default ignore-entities. Declined — the precedent cited is alcatraz's hook CLI (a live prompt guard, where your own LAN addresses are constant noise). The right analogue is its scan CLI, whose ignore default is DATE_TIME,URL — already identical to this action's. IPs are also personal data under GDPR Art. 4(1), so suppressing them by default would be a silent detection regression. Both opt-outs remain one line for anyone who wants them quiet. Full reasoning in the review thread.

Test coverage

  • TestContextScoring pins the boost at threshold: 0.8 for a labelled email and IP, asserting each drops out with context off
  • new CI step exercises -context=true/false end to end through the built binary

Ten releases behind (v0.4.0 -> v0.14.0). No API changes: the scanner only
uses NewEngine, Options{Entities,Threshold,AllowList} and Analyze, all
unchanged, so no scanner code moves. Still dependency-free - the new NER
model support lives in the separate alcatraz/ner submodule.

Detection improves at the action's default threshold of 0.8:

- international and Brazilian phone patterns are recognised at all (v0.7.0)
- entity spans land on word boundaries, so a parenthesised US number is
  captured whole instead of clipped (v0.12.0)
- matches near a labelling word score higher, so a labelled phone (0.85)
  and ip (0.95) now clear 0.8 where they previously scored 0.50 and 0.60
  (v0.14.0)

Existing users will therefore see more findings on unchanged code: better
recall, not a regression.

Also corrects the README threshold guidance, which claimed phone numbers
always score 0.5 - true now only for unlabelled ones.
@qodo-code-review

Copy link
Copy Markdown

PR Summary by Qodo

Upgrade alcatraz dependency to v0.14.0 and align README threshold guidance

⚙️ Configuration changes 📝 Documentation 🕐 10-20 Minutes

Grey Divider

AI Description

• Bump github.com/hoophq/alcatraz from v0.4.0 to v0.14.0 for improved detection scoring.
• Keep scanner integration unchanged (same NewEngine, Options, Analyze API surface).
• Update README to reflect context-aware scoring at the default threshold: 0.8.
Diagram

graph TD
  A["alcatraz-action (this repo)"] --> B["Go module deps (go.mod/go.sum)"] --> C["alcatraz library v0.14.0"] --> D["PII findings output"]
  A --> E["README guidance"]
Loading
High-Level Assessment

The following are alternative approaches to this PR:

1. Stage the upgrade via intermediate minor versions
  • ➕ Smaller behavioral deltas per step, easier to bisect detection changes
  • ➕ Lower risk if an upstream regression appears
  • ➖ More PRs/releases and slower time-to-benefit
  • ➖ Extra CI churn for effectively the same end state
2. Expose/adjust a more conservative default threshold in the action
  • ➕ Reduces surprise for existing users seeing more findings after upgrade
  • ➕ Can preserve historical pass/fail rates without requiring user config changes
  • ➖ Counteracts the main benefit (better recall at 0.8)
  • ➖ Risk of hiding real PII that now correctly clears 0.8
3. Add a small compatibility note / changelog entry in release docs
  • ➕ Sets expectations about increased findings without changing behavior
  • ➕ Low-effort mitigation for downstream consumers
  • ➖ Still requires users to read release notes
  • ➖ Does not help if users pin old assumptions about scoring

Recommendation: Keep the current approach (single-step upgrade) because the integration surface is unchanged and the PR already documents the scoring behavior change. If downstream surprise is a concern, the best follow-up is a brief release-note callout (rather than changing the default threshold and weakening detection).

Files changed (3) +6 / -5

Documentation (1) +3 / -2
README.mdClarify threshold behavior with context-aware scoring +3/-2

Clarify threshold behavior with context-aware scoring

• Updates the threshold guidance to explain that nearby context words can lift scores above the default 0.8. Keeps the note about unlabelled emails/phone numbers typically scoring ~0.5 and therefore requiring a lower threshold to catch.

README.md

Other (2) +3 / -3
go.modBump alcatraz module requirement to v0.14.0 +1/-1

Bump alcatraz module requirement to v0.14.0

• Updates the required version of 'github.com/hoophq/alcatraz' from v0.4.0 to v0.14.0 without modifying any other module settings.

go.mod

go.sumRefresh go.sum checksums for alcatraz v0.14.0 +2/-2

Refresh go.sum checksums for alcatraz v0.14.0

• Replaces the v0.4.0 checksums with v0.14.0 checksums to match the updated module requirement.

go.sum

@qodo-code-review

qodo-code-review Bot commented Jul 31, 2026 •

Copy link
Copy Markdown

Code Review by Qodo

🐞 Bugs (0) 📘 Rule violations (0) 🔗 Cross-repo conflicts (1) 📜 Skill insights (0)

Grey Divider


Remediation recommended

1. IP_ADDRESS now exceeds default 🔗 Cross-repo conflict ≡ Correctness
Description
With alcatraz v0.14.0, context scoring can boost IP_ADDRESS matches (base score 0.6) above this
action’s default threshold: 0.8, but alcatraz-action’s default ignore-entities does not include
IP_ADDRESS. This default mismatch (vs alcatraz’s own CLI/hook defaults) can cause downstream repos
to start failing CI on newly-reported IP findings after this dependency bump.
Code

go.mod[5]

+require github.com/hoophq/alcatraz v0.14.0
Evidence
The action defaults threshold to 0.8 and fails the workflow on findings. In alcatraz v0.14.x, IP
patterns score 0.6 and declare context words; alcatraz also documents a +0.35 boost from context,
which makes labeled IPs exceed 0.8. Alcatraz’s own hook CLI defaults to ignoring IP_ADDRESS, but
alcatraz-action’s default ignore list does not, so this upgrade can change results for downstream
repos.

action.yml[21-41]
action.yml[174-181]
External repo: hoophq/alcatraz, recognizers/generic.go [77-87]
External repo: hoophq/alcatraz, README.md [193-204]
External repo: hoophq/alcatraz, analyzer/engine.go [34-45]
External repo: hoophq/alcatraz, cmd/alcatraz/hook.go [82-87]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
After upgrading to alcatraz v0.14.0, IP addresses can now clear the action’s default `threshold: 0.8` due to context-aware boosting, but the action’s default `ignore-entities` remains `DATE_TIME,URL`. This may introduce noisy findings and unexpected CI failures for downstream repos.

## Issue Context
In alcatraz v0.14.x, context words increase scores by `+0.35` and IP patterns start at `0.6`, so a labeled IP can become `0.95` (> `0.8`). Alcatraz’s own hook CLI defaults to ignoring `IP_ADDRESS`, suggesting it’s commonly treated as noise in certain workflows.

## Fix Focus Areas
- action.yml[21-41]
- action.yml[174-181]
- README.md[148-153]

Potential implementation outline:
- Change `inputs.ignore-entities.default` in `action.yml` to `DATE_TIME,URL,IP_ADDRESS`.
- Update README guidance around thresholds/expected findings to mention IPs explicitly (and how to re-enable detection by removing it from `ignore-entities`).
- Consider adding/adjusting regression fixtures to cover a labeled IP (`ip=192.168.1.44`) at `threshold=0.8` so the behavior is intentional and stable.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


2. No context-scoring toggle ✓ Resolved 🔗 Cross-repo conflict ☼ Reliability
Description
After upgrading to alcatraz v0.14.0, context-aware scoring is enabled by default in the alcatraz
Engine, but alcatraz-action always constructs the engine with defaults and exposes no action
input/flag to disable context boosting. This prevents downstream repos from preserving the previous
pattern-only scoring behavior when upgrading this action.
Code

go.mod[5]

+require github.com/hoophq/alcatraz v0.14.0
Evidence
The PR upgrades alcatraz to v0.14.0, and the action constructs the engine with
alcatraz.NewEngine() without any way to disable context boosting. In alcatraz v0.14.x, the engine
installs a context enhancer by default and explicitly documents disabling it via
SetContextEnhancer(nil) or -context=false, creating a behavioral contract gap for
alcatraz-action consumers.

go.mod[1-5]
cmd/pii-scan/scan.go[37-49]
action.yml[10-45]
External repo: hoophq/alcatraz, analyzer/engine.go [34-45]
External repo: hoophq/alcatraz, analyzer/engine.go [59-64]
External repo: hoophq/alcatraz, README.md [214-229]
External repo: hoophq/alcatraz, cmd/alcatraz/hook.go [82-107]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Upgrading `github.com/hoophq/alcatraz` to v0.14.0 changes scoring behavior because alcatraz now boosts scores based on nearby context words by default. `alcatraz-action` currently provides no way to disable this behavior, which makes it hard for downstream repos to keep stable detection semantics across action upgrades.

## Issue Context
In alcatraz v0.14.x, context-aware scoring is installed by default on Engine construction, and upstream docs/CLI explicitly document ways to disable it (`SetContextEnhancer(nil)` / `-context=false`). The action’s scanner creates an engine via `alcatraz.NewEngine()` and never disables context scoring, and `action.yml` does not offer an input to control it.

## Fix Focus Areas
- cmd/pii-scan/scan.go[37-50]
- cmd/pii-scan/main.go[30-49]
- action.yml[10-45]
- action.yml[92-108]

Potential implementation outline:
- Add an action input like `context-scoring` (default `true` to match upstream alcatraz behavior).
- Thread it to the scanner binary as a `-context` boolean flag.
- When `-context=false`, call `engine.SetContextEnhancer(nil)` after `alcatraz.NewEngine()`.
- Document the behavior in README (especially how it interacts with `threshold`).

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Context
✅ Cross-repo context
  Explored: repo: hoophq/alcatraz (sha: 4afc98c0)

Grey Divider

Tip of the day
💡 Did you know, you can group findings by type and pick your Finding display, from Minimal to Full

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Qodo Logo

Comment thread go.mod
Comment thread go.mod
alcatraz v0.14.0 installs a context enhancer on every engine, so a match
near a word naming its entity type scores up to +0.35. That is the right
default - it is how a labelled phone or ip reaches the 0.8 CI threshold -
but it silently changes what an unchanged repo reports, and a caller who
reads 0.8 as "checksum validated" has no way back to pattern-only scores.

Mirror the escape hatch upstream already ships on both its CLIs: a
-context flag defaulting to true, wired to Engine.SetContextEnhancer(nil),
surfaced as the context-scoring action input. disableContext() matches the
helper in alcatraz's own cmd/alcatraz/scan.go so the two stay in step.

Adds a regression test pinning the boost at threshold 0.8 for a labelled
email and ip - both drop out with context off - plus a CI step covering
the flag end to end.
@github-actions

github-actions Bot commented Jul 31, 2026 •

Copy link
Copy Markdown

🪨 Alcatraz PII scan

No PII detected in this PR's diff. ✅

The self-scan job caught the phone and ip values the previous commit added
to the tests, action.yml and the CI step - correctly, and at exactly the
boosted scores this PR is about (0.85 and 0.95). They are reserved-range
synthetic values, so they belong in .pii-allowlist next to the card, email
and ssn fixtures already there.
@luanlorenzo
luanlorenzo merged commit 2fe78bd into main Aug 11, 2026
3 checks passed
@luanlorenzo
luanlorenzo deleted the luanlorenzo/update-alcatraz-version branch August 11, 2026 20:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant