Skip to content

[medium] Avoid condition-variable collisions for dotted field names - #81

Open
elhoim wants to merge 2 commits into
SigmaHQ:mainfrom
elhoim:fix/dotted-field-condvar-collision
Open

elhoim wants to merge 2 commits into
SigmaHQ:mainfrom
elhoim:fix/dotted-field-condvar-collision

Conversation

@elhoim

@elhoim elhoim commented Sep 24, 2026

Copy link
Copy Markdown
Contributor

BLUF

  • For OR-ed regexes, the backend creates rex/eval variables named <field>Match/<field>Condition, where <field> is the field name without its dotted prefix (clean_field).
  • The counter that adds 2, 3, … to keep those names unique was keyed by the original field name, so a.x and b.x both got xMatch/xCondition, and the second eval overwrote the first.
  • Example: a.x|re or not b.x|re became | search xCondition="true" OR NOT xCondition="true", which is always true, so every event matched.
  • Fix: key the counter by the cleaned name. a.x → xCondition, b.x → xCondition2. Output is unchanged whenever no two regex fields share a cleaned name.

Priority: medium

Details

Dotted field names are common: JSON and cloud logs (requestParameters.x, props.x) and CIM data-model fields (Processes.process, Filesystem.file_path). In SplunkDeferredORRegularExpression:

  • clean_field strips everything up to the last ..
  • add_field/get_field_suffix counted occurrences per original field name.

Two fields that differ only in their prefix therefore both got count 1 and the same variable names.

Before:

| rex field=a.x "(?<xMatch>A)"
| eval xCondition=if(isnotnull(xMatch), "true", "false")
| rex field=b.x "(?<xMatch>B)"
| eval xCondition=if(isnotnull(xMatch), "true", "false")
| search xCondition="true" OR NOT xCondition="true"

After:

| rex field=a.x "(?<xMatch>A)"
| eval xCondition=if(isnotnull(xMatch), "true", "false")
| rex field=b.x "(?<xMatch2>B)"
| eval xCondition2=if(isnotnull(xMatch2), "true", "false")
| search xCondition="true" OR NOT xCondition2="true"

add_field and get_field_suffix now use clean_field(field) as the key. get_all_condition_fields needs no change: clean_field is idempotent, so it now iterates the cleaned keys directly. It is used by the pipeline-hoisting logic, which therefore also sees the correct set of variable names.

Testing

  • New tests:
    • the end-to-end conversion of a.x|re or not b.x|re (exact query)
    • a unit test of get_field_condition, get_field_match and get_all_condition_fields for a.x/b.x
  • Without the fix, both new tests fail. With it, the full suite passes: pytest 141 passed, with Python 3.12 and pySigma 1.2.0 from poetry.lock, run the same way as CI (poetry install + pytest). No existing expectations changed.
  • black==24.1.1 (pre-commit pin) reports nothing on the changed code.

🤖 Generated with Claude Code

clean_field strips everything up to the last dot, but the per-field counter that makes the rex/eval variable names unique was keyed by the original field name. a.x and b.x both became xCondition/xMatch, so the second eval overwrote the first (e.g. 'a.x|re or not b.x|re' became the tautology 'xCondition="true" OR NOT xCondition="true"'). Key the counter by the cleaned name.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The fix is narrowly scoped, addresses the described collision mechanism directly, and is backed by targeted regression and unit tests that exercise the failure mode.

Review effort: Lite
Findings: None

What changed in this PR

This PR fixes a correctness bug in the Splunk backend’s deferred OR-ed regex handling where dotted field names that “clean” to the same suffix (e.g., a.x and b.x) could collide on generated rex/eval variable names, causing later eval stages to overwrite earlier ones and produce incorrect (sometimes always-true) boolean logic in the final | search clause.

Changes:

  • Key SplunkDeferredORRegularExpression.field_counts by the cleaned field name (suffix after the last .) to ensure unique *Match / *Condition variable generation for colliding dotted fields.
  • Add regression and unit tests covering both end-to-end query output and helper naming APIs for the dotted-field collision case.
File Description
sigma/​backends/​splunk/​splunk.py Fixes the counter keying to avoid name collisions for deferred OR regex variables derived from dotted fields.
tests/​test_backend_splunk.py Adds regression/unit tests verifying distinct xCondition/xCondition2 and xMatch/xMatch2 generation for a.x vs b.x.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Resolve conflicts with upstream main, keeping both the upstream changes and this PR's change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants