Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion SECURITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -189,7 +189,7 @@ See [SSRF Protection](docs/features/ssrf-protection.md) for full details, includ

### Cache Key Redaction in Logs (CWE-532)

Cache keys can embed caller-supplied tenant/user identifiers, so **the SDK's own loggers** (`cachekit.*`) never emit them verbatim ([CWE-532][cwe-532]). Every cachekit log path — decorator error handling (structured and backwards-compat), cache-operation logs, and SWR/TTL-refresh debug logs — replaces the key with a fixed-length blake2b digest (`<redacted:…>`), keeping log lines correlatable without leaking the key. Error paths are covered centrally at the shared error sink (`FeatureOrchestrator.handle_cache_error` / `log_cache_operation`), so new call sites are redacted by construction. Both structured cache-operation sinks (`FeatureOrchestrator.log_cache_operation`, `UltraOptimizedStructuredLogger.cache_operation`) also sanitise an exception passed as `error=` themselves — pass the exception object, never `str(e)`, which is emitted as-is. `BackendError` redacts the key in its formatted text (`str(e)` carries `key=<redacted:…>`), while the `.key` attribute keeps the raw caller-supplied key for programmatic use — never log `e.key`. Its free-form `message` is caller-supplied and third-party exception text (a redis `ResponseError` naming the key, a pymemcache illegal-input error echoing it) has unknown provenance — so **no cachekit log line renders `str(e)`**. Every logging call that mentions an exception goes through `redact_error_for_log`, which emits only the exception type plus, for `BackendError`, its `BackendErrorType` classification; the full exception stays on the object (`original_exception`, `.message`) for programmatic access. Operators lose the provider's message text in the log line and keep it on the exception. An architecture test (`tests/unit/test_log_redaction_architecture.py`) walks every logging call in the package — `logger.*()`, `get_logger().*()`, `getattr(logger, level)()` — and fails CI if a key-shaped value reaches one unredacted in the message, `%s` arguments, or `extra=`; if an exception — any name bound by `except ... as`, a conventional name (`e`, `exc`, `err`, `error`, `*_err`), or an attribute of one — reaches one outside `redact_error_for_log`; or if a call emits a traceback (`logger.exception`, `exc_info=`). The guarantee does not depend on the next contributor remembering it. It is flow-insensitive: build log lines inline, not via a pre-formatted variable, and bind exceptions with `except ... as` or a conventional name (an `Exception`-typed parameter called `failure` is invisible to it), or the guard cannot see them.
Cache keys can embed caller-supplied tenant/user identifiers, so **the SDK's own loggers** (`cachekit.*`) never emit them verbatim ([CWE-532][cwe-532]). Every cachekit log path — decorator error handling (structured and backwards-compat), cache-operation logs, and SWR/TTL-refresh debug logs — replaces the key with a fixed-length blake2b digest (`<redacted:…>`), keeping log lines correlatable without leaking the key. Error paths are covered centrally at the shared error sink (`FeatureOrchestrator.handle_cache_error` / `log_cache_operation`), so new call sites are redacted by construction. Both structured cache-operation sinks (`FeatureOrchestrator.log_cache_operation`, `StructuredLogger.cache_operation`) also sanitise an exception passed as `error=` themselves — pass the exception object, never `str(e)`, which is emitted as-is. `BackendError` redacts the key in its formatted text (`str(e)` carries `key=<redacted:…>`), while the `.key` attribute keeps the raw caller-supplied key for programmatic use — never log `e.key`. Its free-form `message` is caller-supplied and third-party exception text (a redis `ResponseError` naming the key, a pymemcache illegal-input error echoing it) has unknown provenance — so **no cachekit log line renders `str(e)`**. Every logging call that mentions an exception goes through `redact_error_for_log`, which emits only the exception type plus, for `BackendError`, its `BackendErrorType` classification; the full exception stays on the object (`original_exception`, `.message`) for programmatic access. Operators lose the provider's message text in the log line and keep it on the exception. An architecture test (`tests/unit/test_log_redaction_architecture.py`) walks every logging call in the package — `logger.*()`, `get_logger().*()`, `getattr(logger, level)()` — and fails CI if a key-shaped value reaches one unredacted in the message, `%s` arguments, or `extra=`; if an exception — any name bound by `except ... as`, a conventional name (`e`, `exc`, `err`, `error`, `*_err`), or an attribute of one — reaches one outside `redact_error_for_log`; or if a call emits a traceback (`logger.exception`, `exc_info=`). The guarantee does not depend on the next contributor remembering it. It is flow-insensitive: build log lines inline, not via a pre-formatted variable, and bind exceptions with `except ... as` or a conventional name (an `Exception`-typed parameter called `failure` is invisible to it), or the guard cannot see them.

**Scope — transport logs are not covered.** The CachekitIO backend addresses entries by key in the request path (`GET /v1/cache/{key}`), and `httpx` logs every request line — method, full URL, status — at `INFO` on its own `httpx` logger. An application that enables `INFO` globally (`logging.basicConfig(level=logging.INFO)`) will therefore see raw keys in *httpx's* output on every operation, exactly as it would see any REST resource path. cachekit does not mute a third-party logger on your behalf; if your keys carry identifiers, silence or raise the level of that logger in your logging config:

Expand Down
6 changes: 3 additions & 3 deletions src/cachekit/decorators/orchestrator.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@

from ..hash_utils import redact_error_for_log, redact_key_for_log
from ..monitoring.correlation_tracking import CorrelationTracker
from ..monitoring.pool_monitor import OptimizedPoolMonitor
from ..monitoring.pool_monitor import PoolMonitor

# Import EXISTING modules - no duplication
from ..reliability import (
Expand Down Expand Up @@ -127,14 +127,14 @@ def correlation_tracker(self) -> Optional[CorrelationTracker]:
return self._correlation_tracker

@property
def pool_monitor(self) -> Optional[OptimizedPoolMonitor]:
def pool_monitor(self) -> Optional[PoolMonitor]:
"""Get pool monitor."""
return self._pool_monitor

def set_pool_manager(self, pool_manager) -> None:
"""Initialize pool monitor with pool manager."""
if pool_manager and not self._pool_monitor:
self._pool_monitor = OptimizedPoolMonitor(pool_manager)
self._pool_monitor = PoolMonitor(pool_manager)

def should_allow_request(self) -> bool:
"""Check if request should be allowed based on circuit breaker state."""
Expand Down
2 changes: 1 addition & 1 deletion src/cachekit/decorators/wrapper.py
Original file line number Diff line number Diff line change
Expand Up @@ -1606,7 +1606,7 @@ async def async_wrapper(*args: Any, **kwargs: Any) -> Any:
raise TypeError(f"key function must return str, got {type(custom_key).__name__}")
cache_key = f"{namespace or 'default'}:{custom_key}"
elif fast_mode:
# Ultra-fast key generation for hot paths (10-50μs savings)
# Fast-path key generation (10-50μs savings)
from ..hash_utils import cache_key_hash

cache_namespace = namespace or "default"
Expand Down
2 changes: 1 addition & 1 deletion src/cachekit/hash_utils.py
Original file line number Diff line number Diff line change
Expand Up @@ -105,7 +105,7 @@ def redact_error_for_log(error: object) -> str:


def fast_hash(data: Union[str, bytes], digest_size: int = 8) -> str:
"""Ultra-fast hash using BLAKE3 - optimized for hot paths.
"""Fast hash using BLAKE3 for hot paths.

Args:
data: String or bytes to hash
Expand Down
20 changes: 8 additions & 12 deletions src/cachekit/logging.py
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
"""Ultra-optimized structured logging with minimal overhead.
"""Structured logging with minimal overhead.

This module provides lock-free, sampling-based structured logging
that reduces overhead from 570% to <5% while maintaining functionality.
Expand Down Expand Up @@ -151,8 +151,8 @@ def _write_batch(self, entries: list[LogEntry]):
pass


class UltraOptimizedStructuredLogger:
"""Ultra-optimized structured logger with <5% overhead.
class StructuredLogger:
"""Structured logger with <5% overhead.

Features:
- Lock-free ring buffer
Expand Down Expand Up @@ -456,7 +456,7 @@ def __del__(self):
class SimpleSpan:
"""Simple span implementation for tracing integration."""

def __init__(self, logger: UltraOptimizedStructuredLogger, name: str, **kwargs):
def __init__(self, logger: StructuredLogger, name: str, **kwargs):
self.logger = logger
self.name = name
self.kwargs = kwargs
Expand All @@ -479,18 +479,18 @@ def __exit__(self, exc_type, exc_val, exc_tb):


# Global logger instances cache
_logger_instances: dict[str, UltraOptimizedStructuredLogger] = {}
_logger_instances: dict[str, StructuredLogger] = {}
_logger_lock = threading.Lock()


def get_structured_logger(name: str) -> UltraOptimizedStructuredLogger:
def get_structured_logger(name: str) -> StructuredLogger:
"""Get or create a structured logger instance.

Args:
name: Logger name (usually __name__)

Returns:
Ultra-optimized structured logger instance
Structured logger instance
"""
# Fast path - check if already exists
if name in _logger_instances:
Expand All @@ -500,14 +500,10 @@ def get_structured_logger(name: str) -> UltraOptimizedStructuredLogger:
with _logger_lock:
# Double-check pattern
if name not in _logger_instances:
_logger_instances[name] = UltraOptimizedStructuredLogger(name)
_logger_instances[name] = StructuredLogger(name)
return _logger_instances[name]


# Alias
StructuredRedisLogger = UltraOptimizedStructuredLogger


class JsonFormatter(logging.Formatter):
"""JSON formatter for log records."""

Expand Down
10 changes: 5 additions & 5 deletions src/cachekit/monitoring/pool_monitor.py
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
"""Optimized connection pool monitoring with <3% overhead.
"""Connection pool monitoring with <3% overhead.

This module provides lazy, sampling-based connection pool monitoring
that reduces overhead from 1409% to <3% while maintaining functionality.
Expand Down Expand Up @@ -84,8 +84,8 @@ def is_stale(self) -> bool:
return time.time() - self.last_update > self.cache_duration


class OptimizedPoolMonitor:
"""Optimized pool monitor with <3% overhead.
class PoolMonitor:
"""Pool monitor with <3% overhead.

Features:
- Lazy stats calculation with 5-second caching
Expand All @@ -101,7 +101,7 @@ class OptimizedPoolMonitor:
>>> mock_pool_manager = Mock()
>>> mock_pool_manager.is_sync_initialized = False
>>> mock_pool_manager.pool = None
>>> monitor = OptimizedPoolMonitor(mock_pool_manager, sampling_rate=0.01)
>>> monitor = PoolMonitor(mock_pool_manager, sampling_rate=0.01)
>>> monitor.sampling_rate
0.01

Expand All @@ -127,7 +127,7 @@ class OptimizedPoolMonitor:
"""

def __init__(self, pool_manager, sampling_rate: float = 0.01):
"""Initialize optimized monitor.
"""Initialize the monitor.

Args:
pool_manager: The connection pool manager to monitor
Expand Down
2 changes: 1 addition & 1 deletion src/cachekit/object_cache.py
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
"""Thread-safe in-memory object cache with TTL, LRU eviction, byte bounds, and SWR.

Stores Python object references directly — no serialization. Used by @cache.local()
and by @cache(backend=None) (L1-only mode) to provide ultra-low-latency (~50ns)
and by @cache(backend=None) (L1-only mode) to provide low-latency (~50ns)
caching for objects that do not need to cross process boundaries or survive restarts.
"""

Expand Down
8 changes: 4 additions & 4 deletions tests/unit/test_error_path_key_redaction.py
Original file line number Diff line number Diff line change
Expand Up @@ -31,7 +31,7 @@
from cachekit.decorators.orchestrator import FeatureOrchestrator
from cachekit.hash_utils import _SENTINEL_KEYS, redact_cache_key
from cachekit.key_generator import CacheKeyGenerator
from cachekit.logging import UltraOptimizedStructuredLogger
from cachekit.logging import StructuredLogger
from cachekit.serializers.base import SerializationError

TENANT_KEY = "ns:tenant-42-alice-secret:func:app.get_user:args:deadbeef:v1"
Expand Down Expand Up @@ -285,7 +285,7 @@ async def test_async_get_failure_redacts_key_in_exception_text(self, caplog: pyt


class TestStructuredLoggerCacheOperationRedaction:
"""``UltraOptimizedStructuredLogger.cache_operation`` is a direct sink.
"""``StructuredLogger.cache_operation`` is a direct sink.

``cache_hit``/``cache_miss``/``cache_stored`` all funnel through it, so this
one method is the whole surface. It must apply the *same* pass-through policy
Expand All @@ -296,7 +296,7 @@ class TestStructuredLoggerCacheOperationRedaction:
"""

def _emit(self, caplog: pytest.LogCaptureFixture, cache_key: str) -> str:
logger = UltraOptimizedStructuredLogger("test.cache_operation")
logger = StructuredLogger("test.cache_operation")

with caplog.at_level(logging.INFO, logger="test.cache_operation"):
logger.cache_operation("get", cache_key, hit=True)
Expand Down Expand Up @@ -418,7 +418,7 @@ class TestErrorKwargSanitisedAtSink:

@pytest.mark.parametrize(("error", "rendered"), ERROR_KWARGS)
def test_logging_sink(self, error: object, rendered: str, caplog: pytest.LogCaptureFixture) -> None:
logger = UltraOptimizedStructuredLogger("test.error_kwarg")
logger = StructuredLogger("test.error_kwarg")

with caplog.at_level(logging.INFO, logger="test.error_kwarg"):
logger.cache_operation("set", TENANT_KEY, error=error)
Expand Down
10 changes: 5 additions & 5 deletions tests/unit/test_structured_logging.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,18 +11,18 @@
from cachekit.backends.errors import BackendError, BackendErrorType
from cachekit.logging import (
JsonFormatter,
StructuredRedisLogger,
StructuredLogger,
get_structured_logger,
)


class TestStructuredRedisLogger:
"""Test StructuredRedisLogger functionality."""
class TestStructuredLogger:
"""Test StructuredLogger functionality."""

@pytest.fixture
def logger(self):
"""Create a test logger instance."""
return StructuredRedisLogger("test_logger")
return StructuredLogger("test_logger")

def test_logger_initialization(self, logger):
"""Test logger initialization."""
Expand Down Expand Up @@ -279,7 +279,7 @@ class TestFactoryFunction:
def test_get_structured_logger(self):
"""Test get_structured_logger factory returns one cached instance per name."""
logger1 = get_structured_logger("test1")
assert isinstance(logger1, StructuredRedisLogger)
assert isinstance(logger1, StructuredLogger)

logger1_again = get_structured_logger("test1")
assert logger1_again is logger1
Loading