Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .secrets.baseline

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

55 changes: 39 additions & 16 deletions docs/backends/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -170,43 +170,66 @@ def explicit_backend():

### 2. Module-Level Default Backend (Middle Priority)

```python notest
```python
import tempfile

from cachekit import cache
from cachekit.config.decorator import set_default_backend
from cachekit.backends.file import FileBackend, FileBackendConfig

# Set once at application startup
file_backend = FileBackend(FileBackendConfig(cache_dir="/var/cache/myapp"))
file_backend = FileBackend(FileBackendConfig(cache_dir=tempfile.mkdtemp()))
set_default_backend(file_backend)

# All decorators now use file backend — no backend= needed
@cache.minimal(ttl=300)
def fast_lookup():
return data()
def fast_lookup(x: int) -> int:
return x * 2

@cache.production(ttl=600)
def critical_function():
return data()
def critical_function(x: int) -> int:
return x * 3

assert fast_lookup(2) == 4 and critical_function(2) == 6

set_default_backend(None) # clear the default
```

Call `set_default_backend(None)` to clear the default. Works with any backend (Redis, File, CachekitIO, custom).

**Import order does not matter, but configuration must happen before a decorated
function's first call.** A decorator applied
without `backend=` pins the default when it is first seen — at decoration if
already set, otherwise at first call — so the usual layout (business modules
imported at the top of the file, `set_default_backend()` in `main()`) works.
Later `set_default_backend()` calls do not re-point already-pinned functions.
Exception: `stale_ttl` and `@cache.io`'s default stale window validate SWR
capability at decoration, so set a CachekitIO default *before* importing modules
that use them.

### 3. Environment Variable Auto-Detection (Lowest Priority)

```bash
# Primary: CACHEKIT_REDIS_URL
CACHEKIT_REDIS_URL=redis://prod.example.com:6379
If no explicit backend and no module-level default, `DefaultBackendProvider`
picks a backend from exactly one environment selector, in this order:

# Fallback: REDIS_URL
REDIS_URL=redis://localhost:6379
```
| Priority | Environment variable | Backend |
|----------|-----------------------------|--------------------|
| 1 | `CACHEKIT_API_KEY` | `CachekitIOBackend` (SaaS) |
| 2 | `CACHEKIT_REDIS_URL` | `RedisBackend` |
| 3 | `CACHEKIT_MEMCACHED_SERVERS`| `MemcachedBackend` |
| 4 | `CACHEKIT_FILE_CACHE_DIR` | `FileBackend` |
| 5 | `REDIS_URL`, or nothing set | `RedisBackend` (localhost fallback) |

If no explicit backend and no module-level default, cachekit creates a RedisBackend from environment variables.
Setting more than one of the four `CACHEKIT_*` selectors is ambiguous and raises
`ConfigurationError` at first call. The decorator catches it, logs a WARNING on
the `cachekit.decorators.orchestrator` logger, and runs the function uncached.
`REDIS_URL` is a 12-factor fallback and never counts as a conflict.

**Resolution order**:
1. Check for explicit `backend` parameter in `@cache(backend=...)`
2. Check for module-level default via `set_default_backend()`
3. Create RedisBackend from environment variables (CACHEKIT_REDIS_URL > REDIS_URL)
1. Explicit `backend` parameter in `@cache(backend=...)`
2. Module-level default via `set_default_backend()` (checked at decoration, and
again at first call if still unset)
3. Environment auto-detection per the table above

## Performance Considerations

Expand Down
20 changes: 13 additions & 7 deletions docs/features/distributed-locking.md
Original file line number Diff line number Diff line change
Expand Up @@ -41,7 +41,7 @@ report = await get_report("2025-01-15")
> [!NOTE]
> Locking requires **both** of:
>
> 1. A backend implementing the `LockableBackend` protocol. `RedisBackend` and `CachekitIOBackend` (the SaaS backend behind `api.cachekit.io`) both do. `FileBackend` and pure-L1 (zero-config) caching don't — they silently skip lock acquisition; the function still works, just without stampede protection.
> 1. A backend implementing the `LockableBackend` protocol. `CachekitIOBackend` (the SaaS backend behind `api.cachekit.io`) does, and so does the Redis backend you get from env auto-detection or `RedisBackendProvider` (`PerRequestRedisBackend`). A `RedisBackend` you construct yourself and pass as `backend=` does **not** — it has no `acquire_lock`. Neither do `FileBackend` or pure-L1 (zero-config) caching. All of them silently skip lock acquisition; the function still works, just without stampede protection.
> 2. An **async** decorated function. Sync wrappers never take the lock path on any backend — see [Async-only](#async-only-sync-functions-are-never-lock-protected).

---
Expand Down Expand Up @@ -192,9 +192,11 @@ leaderboard = await get_leaderboard()
### With Redis Backend (Explicit)
```python notest
from cachekit import cache
from cachekit.backends.redis import RedisBackend
from cachekit.backends.redis.provider import RedisBackendProvider, tenant_context

backend = RedisBackend() # Implements LockableBackend
# PerRequestRedisBackend implements LockableBackend; a bare RedisBackend() does not.
tenant_context.set("default")
backend = RedisBackendProvider(redis_url="redis://localhost:6379").get_backend()

@cache(ttl=300, backend=backend)
async def generate_stats(date):
Expand Down Expand Up @@ -238,16 +240,20 @@ def cheap_lookup(x):
The `LockableBackend` protocol defines how backends provide distributed locking:

```python notest
async def acquire_lock(
def acquire_lock(
self,
key: str, # Bare cache key (same key as get/set); backend derives lock namespace
timeout: float, # How long to hold the lock (seconds)
blocking_timeout: Optional[float] = None, # Max wait to acquire (None = non-blocking)
) -> AsyncIterator[bool]:
# Yields True if lock acquired, False if timeout waiting
) -> AbstractAsyncContextManager[bool]:
# Entering the context yields True if the lock was acquired, False if the wait timed out
...
```

Implementations are `async` generators wrapped in `@asynccontextmanager`, so the
protocol declares the *decorated* shape — a backend author must apply the
decorator for `async with` to work.

The decorator wrapper calls it with `timeout=30.0` (lock self-expiry) and
`blocking_timeout=5.0` (max wait to acquire) — see
[Lock Timing](#lock-timing-30-second-expiry-5-second-wait).
Expand Down Expand Up @@ -332,7 +338,7 @@ A: Your function takes longer than the 5 s `blocking_timeout`, so waiters fall t
**Q: Locking doesn't seem to be working**
A: Two things to check:
1. The decorated function must be **async** — sync wrappers never lock ([Async-only](#async-only-sync-functions-are-never-lock-protected)).
2. The backend must implement `LockableBackend` (`RedisBackend`, `CachekitIOBackend`). Check with `from cachekit.backends.base import LockableBackend; isinstance(backend, LockableBackend)`.
2. The backend must implement `LockableBackend` (`CachekitIOBackend`, or `PerRequestRedisBackend` from `RedisBackendProvider` — a bare `RedisBackend()` does not). Check it the way the SDK does: `hasattr(backend, "acquire_lock")`. Avoid `isinstance(backend, LockableBackend)` — from CPython 3.12 a `runtime_checkable` protocol check resolves members with `inspect.getattr_static`, so it reports `False` for a backend that delegates `acquire_lock` through `__getattr__`.

**Q: How do I know if stampedes are happening?**
A: Check Prometheus: a spike in `rate(redis_cache_operations_total{status="miss"}[1m])` = stampede risk. See [Prometheus Metrics](prometheus-metrics.md).
Expand Down
20 changes: 13 additions & 7 deletions src/cachekit/backends/base.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,7 @@
from __future__ import annotations

from collections.abc import AsyncIterator, Callable
from contextlib import AbstractAsyncContextManager
from typing import Any, BinaryIO, Optional, Protocol, runtime_checkable

# Re-export BackendError for convenience (public API)
Expand Down Expand Up @@ -269,8 +270,11 @@ class LockableBackend(Protocol):
features like cache stampede prevention and critical sections.

Not all backends support this capability:
- Supported: RedisBackend, CachekitIOBackend (SaaS ``POST /v1/cache/{key}/lock``)
- Not supported: FileBackend, L1-only (in-memory)
- Supported: ``PerRequestRedisBackend`` — what ``RedisBackendProvider`` and
therefore the env-resolved Redis path hand out — and ``CachekitIOBackend``
(SaaS ``POST /v1/cache/{key}/lock``).
- Not supported: ``RedisBackend`` constructed directly and passed as
``backend=``, ``FileBackend``, L1-only (in-memory).

Contract — bare cache key:
``acquire_lock`` receives the **bare cache key**, identical to what
Expand All @@ -293,12 +297,12 @@ class LockableBackend(Protocol):
>>> # result = expensive_computation()
"""

async def acquire_lock(
def acquire_lock(
self,
key: str,
timeout: float,
blocking_timeout: Optional[float] = None,
) -> AsyncIterator[bool]:
) -> AbstractAsyncContextManager[bool]:
"""Acquire a distributed lock on key.

Args:
Expand All @@ -309,9 +313,11 @@ async def acquire_lock(
timeout: How long to hold the lock (seconds) before auto-release
blocking_timeout: Max time to wait for lock acquisition (None = non-blocking)

Yields:
True if lock was acquired
False if timeout occurred waiting for lock
Returns:
An async context manager yielding True if the lock was acquired,
False if the wait timed out. Implementations are ``async``
generators wrapped in ``@asynccontextmanager``, so this protocol
declares the *decorated* shape.

Raises:
BackendError: If backend operation fails
Expand Down
29 changes: 29 additions & 0 deletions src/cachekit/cache_handler.py
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,7 @@
BufferHandle,
BufferReadableBackend,
BufferWritableBackend,
LockableBackend,
TTLInspectableBackend,
)
from cachekit.backends.provider import (
Expand Down Expand Up @@ -192,6 +193,34 @@ def supports_ttl_inspection(backend: BaseBackend) -> TypeGuard[TTLInspectableBac
return hasattr(backend, "get_ttl") and hasattr(backend, "refresh_ttl")


def supports_locking(backend: object) -> TypeGuard[LockableBackend]:
"""Type guard: backend provides distributed locking (stampede prevention).

Takes ``object``, not ``BaseBackend`` like its siblings, because the
decorator probes the lazily-resolved ``_backend`` cell, which is ``None``
until first call and may already be narrowed by another capability guard.

Checked on the INSTANCE, deliberately — unlike ``supports_swr`` below, which
is class-level to keep ``__getattr__`` proxies and mocks off the freshness
read path. The asymmetry is the failure direction: a false negative silently
drops stampede protection on a hot key, while a false positive raises out of
the ``async with`` — and does NOT fail open. The wrapper's handler degrades
to lockless execution only for a ``BackendError``; a ``TypeError`` from
calling a non-callable propagates to the caller and breaks the decorated
function. Hence ``callable``, not ``hasattr``: a backend carrying
``acquire_lock = None`` is not lockable, and must take the lockless path
rather than crash the call it was meant to protect.

Equally deliberate: not ``isinstance(backend, LockableBackend)``. Since
CPython 3.12 a ``runtime_checkable`` Protocol check resolves members with
``inspect.getattr_static``, which does not consult ``__getattr__`` — so a
delegating backend proxy locks on 3.10/3.11 and silently stops locking on
3.12+. Plain ``getattr`` consults ``__getattr__`` and is stable across
every supported interpreter.
"""
return callable(getattr(backend, "acquire_lock", None))


# Backend type names already warned about, so refresh_ttl_on_get degradation warns at most
# once per backend type per process (avoids per-hit log spam). Tests clear this set.
_TTL_REFRESH_UNSUPPORTED_WARNED: set[str] = set()
Expand Down
6 changes: 5 additions & 1 deletion src/cachekit/decorators/intent.py
Original file line number Diff line number Diff line change
Expand Up @@ -137,7 +137,11 @@ def decorator(f: F) -> F:
backend = manual_overrides.pop("backend", None)

# Tier 2 resolution: if no explicit backend and not L1-only mode,
# check module-level default set via set_default_backend()
# check module-level default set via set_default_backend(). Kept here
# (not only lazily) because decoration-time validation — the interop
# backend guard and stale_ttl/SWR capability (LAB-557) — needs the
# backend when it is already known. If the default is set LATER, the
# wrapper re-consults it at first call (_resolve_lazy_backend, LAB-4457).
if backend is None and not _explicit_l1_only:
from ..config.decorator import get_default_backend

Expand Down
38 changes: 27 additions & 11 deletions src/cachekit/decorators/wrapper.py
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,9 @@
get_logger,
handle_decrypt_failure,
redact_cache_key,
supports_locking,
supports_swr,
supports_ttl_inspection,
warn_ttl_refresh_unsupported,
)
from ..interop import (
Expand All @@ -46,8 +48,23 @@
from .tenant_context import TenantContextExtractor

if TYPE_CHECKING:
from ..backends.base import BaseBackend
from ..serializers.base import SerializerProtocol


def _resolve_lazy_backend() -> BaseBackend:
"""Backend for a decorator that was applied without ``backend=``.

Consulted at FIRST CALL, not at decoration, so ``set_default_backend()``
takes effect regardless of whether it ran before or after the module holding
the decorated function was imported (LAB-4457).
"""
from ..config.decorator import get_default_backend

default = get_default_backend()
return default if default is not None else get_backend_provider().get_backend()


F = TypeVar("F", bound=Callable[..., Any])

_logger = logging.getLogger(__name__)
Expand Down Expand Up @@ -907,9 +924,8 @@ async def _l2_swr_revalidate_async(cache_key: str, call_args: tuple[Any, ...], c
the caller already got the stale value; the entry hard-expires at evict_at and
the next request takes the ordinary synchronous miss path (spec degradation)."""
try:
_acquire_lock = getattr(_backend, "acquire_lock", None)
if _acquire_lock is not None:
async with _acquire_lock(cache_key, timeout=_l2_swr_lease_seconds, blocking_timeout=None) as got_lease:
if supports_locking(_backend):
async with _backend.acquire_lock(cache_key, timeout=_l2_swr_lease_seconds, blocking_timeout=None) as got_lease:
if not got_lease:
return # another client is revalidating — stale already served
await _l2_swr_recompute_store_async(cache_key, call_args, call_kwargs)
Expand Down Expand Up @@ -1263,7 +1279,7 @@ def sync_wrapper(*args: Any, **kwargs: Any) -> Any: # noqa: PLR0912

nonlocal _backend
if _backend is None:
_backend = get_backend_provider().get_backend()
_backend = _resolve_lazy_backend()

# Setup cache handler strategy on first use
handler = StandardCacheHandler(
Expand Down Expand Up @@ -1672,7 +1688,7 @@ async def async_wrapper(*args: Any, **kwargs: Any) -> Any:
if interop is not None:
if _backend is None:
try:
_backend = get_backend_provider().get_backend()
_backend = _resolve_lazy_backend()
except Exception as e:
# If Redis connection fails, execute function without caching - RETURN EARLY
# This prevents the decorator from breaking the application
Expand Down Expand Up @@ -1746,7 +1762,7 @@ async def async_wrapper(*args: Any, **kwargs: Any) -> Any:
# Initialize backend only when needed (lazy init for performance)
if _backend is None:
try:
_backend = get_backend_provider().get_backend()
_backend = _resolve_lazy_backend()
except Exception as e:
# If Redis connection fails, execute function without caching - RETURN EARLY
# This prevents the decorator from breaking the application
Expand Down Expand Up @@ -1805,7 +1821,7 @@ async def async_wrapper(*args: Any, **kwargs: Any) -> Any:
_l1_backfill_from_l2(cache_key, cached_data, _l2_is_stale, _l2_fresh_for)

# Handle TTL refresh if configured and threshold met
if refresh_ttl_on_get and ttl and hasattr(_backend, "get_ttl") and hasattr(_backend, "refresh_ttl"):
if refresh_ttl_on_get and ttl and supports_ttl_inspection(_backend):
try:
remaining_ttl = await _backend.get_ttl(cache_key)
if remaining_ttl and remaining_ttl < (ttl * ttl_refresh_threshold):
Expand Down Expand Up @@ -1859,7 +1875,7 @@ async def async_wrapper(*args: Any, **kwargs: Any) -> Any:
blocking_timeout = 5.0 # Wait up to 5 seconds to acquire lock

# Check if backend supports distributed locking
if hasattr(_backend, "acquire_lock"):
if supports_locking(_backend):
try:
# Use backend's async lock protocol
async with _backend.acquire_lock(
Expand Down Expand Up @@ -2003,7 +2019,7 @@ async def async_wrapper(*args: Any, **kwargs: Any) -> Any:
# Fall through to execute without locking

# Execute without locking (either backend doesn't support it or lock failed)
if not hasattr(_backend, "acquire_lock"):
if not supports_locking(_backend):
logger().debug(
f"Backend doesn't support locking for {redact_cache_key(cache_key)}, executing without thundering herd protection"
)
Expand Down Expand Up @@ -2078,7 +2094,7 @@ def invalidate_cache(*args: Any, **kwargs: Any) -> None:
# we should NOT try to get a backend from the provider
if not _l1_only_mode and _backend is None:
try:
_backend = get_backend_provider().get_backend()
_backend = _resolve_lazy_backend()
except Exception as e:
# If backend creation fails, can't invalidate L2
_logger.debug("Failed to get backend for invalidation: %s", redact_error_for_log(e))
Expand Down Expand Up @@ -2138,7 +2154,7 @@ async def ainvalidate_cache(*args: Any, **kwargs: Any) -> None:
# we should NOT try to get a backend from the provider
if not _l1_only_mode and _backend is None:
try:
_backend = get_backend_provider().get_backend()
_backend = _resolve_lazy_backend()
except Exception as e:
# If backend creation fails, can't invalidate L2
_logger.debug("Failed to get backend for async invalidation: %s", redact_error_for_log(e))
Expand Down
Loading
Loading