Repository navigation
Conversation
clean_field strips everything up to the last dot, but the per-field counter that makes the rex/eval variable names unique was keyed by the original field name. a.x and b.x both became xCondition/xMatch, so the second eval overwrote the first (e.g. 'a.x|re or not b.x|re' became the tautology 'xCondition="true" OR NOT xCondition="true"'). Key the counter by the cleaned name. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Contributor
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The fix is narrowly scoped, addresses the described collision mechanism directly, and is backed by targeted regression and unit tests that exercise the failure mode.
Review effort: Lite
Findings: None
What changed in this PR
This PR fixes a correctness bug in the Splunk backend’s deferred OR-ed regex handling where dotted field names that “clean” to the same suffix (e.g., a.x and b.x) could collide on generated rex/eval variable names, causing later eval stages to overwrite earlier ones and produce incorrect (sometimes always-true) boolean logic in the final | search clause.
Changes:
- Key
SplunkDeferredORRegularExpression.field_countsby the cleaned field name (suffix after the last.) to ensure unique*Match/*Conditionvariable generation for colliding dotted fields. - Add regression and unit tests covering both end-to-end query output and helper naming APIs for the dotted-field collision case.
| File | Description |
|---|---|
sigma/backends/splunk/splunk.py |
Fixes the counter keying to avoid name collisions for deferred OR regex variables derived from dotted fields. |
tests/test_backend_splunk.py |
Adds regression/unit tests verifying distinct xCondition/xCondition2 and xMatch/xMatch2 generation for a.x vs b.x. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Resolve conflicts with upstream main, keeping both the upstream changes and this PR's change. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
BLUF
rex/evalvariables named<field>Match/<field>Condition, where<field>is the field name without its dotted prefix (clean_field).2,3, … to keep those names unique was keyed by the original field name, soa.xandb.xboth gotxMatch/xCondition, and the secondevaloverwrote the first.a.x|re or not b.x|rebecame| search xCondition="true" OR NOT xCondition="true", which is always true, so every event matched.a.x→xCondition,b.x→xCondition2. Output is unchanged whenever no two regex fields share a cleaned name.Priority: medium
Details
Dotted field names are common: JSON and cloud logs (
requestParameters.x,props.x) and CIM data-model fields (Processes.process,Filesystem.file_path). InSplunkDeferredORRegularExpression:clean_fieldstrips everything up to the last..add_field/get_field_suffixcounted occurrences per original field name.Two fields that differ only in their prefix therefore both got count 1 and the same variable names.
Before:
After:
add_fieldandget_field_suffixnow useclean_field(field)as the key.get_all_condition_fieldsneeds no change:clean_fieldis idempotent, so it now iterates the cleaned keys directly. It is used by the pipeline-hoisting logic, which therefore also sees the correct set of variable names.Testing
a.x|re or not b.x|re(exact query)get_field_condition,get_field_matchandget_all_condition_fieldsfora.x/b.xpytest141 passed, with Python 3.12 and pySigma 1.2.0 frompoetry.lock, run the same way as CI (poetry install+pytest). No existing expectations changed.black==24.1.1(pre-commit pin) reports nothing on the changed code.🤖 Generated with Claude Code