Skip to content
Merged
Show file tree
Hide file tree
Changes from 5 commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .secrets.baseline

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

55 changes: 39 additions & 16 deletions docs/backends/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -170,43 +170,66 @@ def explicit_backend():

### 2. Module-Level Default Backend (Middle Priority)

```python notest
```python
import tempfile

from cachekit import cache
from cachekit.config.decorator import set_default_backend
from cachekit.backends.file import FileBackend, FileBackendConfig

# Set once at application startup
file_backend = FileBackend(FileBackendConfig(cache_dir="/var/cache/myapp"))
file_backend = FileBackend(FileBackendConfig(cache_dir=tempfile.mkdtemp()))
set_default_backend(file_backend)

# All decorators now use file backend — no backend= needed
@cache.minimal(ttl=300)
def fast_lookup():
return data()
def fast_lookup(x: int) -> int:
return x * 2

@cache.production(ttl=600)
def critical_function():
return data()
def critical_function(x: int) -> int:
return x * 3

assert fast_lookup(2) == 4 and critical_function(2) == 6

set_default_backend(None) # clear the default
```

Call `set_default_backend(None)` to clear the default. Works with any backend (Redis, File, CachekitIO, custom).

**Import order does not matter, but configuration must happen before a decorated
function's first call.** A decorator applied
without `backend=` pins the default when it is first seen — at decoration if
already set, otherwise at first call — so the usual layout (business modules
imported at the top of the file, `set_default_backend()` in `main()`) works.
Later `set_default_backend()` calls do not re-point already-pinned functions.
Exception: `stale_ttl` and `@cache.io`'s default stale window validate SWR
capability at decoration, so set a CachekitIO default *before* importing modules
that use them.

### 3. Environment Variable Auto-Detection (Lowest Priority)

```bash
# Primary: CACHEKIT_REDIS_URL
CACHEKIT_REDIS_URL=redis://prod.example.com:6379
If no explicit backend and no module-level default, `DefaultBackendProvider`
picks a backend from exactly one environment selector, in this order:

# Fallback: REDIS_URL
REDIS_URL=redis://localhost:6379
```
| Priority | Environment variable | Backend |
|----------|-----------------------------|--------------------|
| 1 | `CACHEKIT_API_KEY` | `CachekitIOBackend` (SaaS) |
| 2 | `CACHEKIT_REDIS_URL` | `RedisBackend` |
| 3 | `CACHEKIT_MEMCACHED_SERVERS`| `MemcachedBackend` |
| 4 | `CACHEKIT_FILE_CACHE_DIR` | `FileBackend` |
| 5 | `REDIS_URL`, or nothing set | `RedisBackend` (localhost fallback) |

If no explicit backend and no module-level default, cachekit creates a RedisBackend from environment variables.
Setting more than one of the four `CACHEKIT_*` selectors is ambiguous and raises
`ConfigurationError` at first call. The decorator catches it, logs a WARNING on
the `cachekit.decorators.orchestrator` logger, and runs the function uncached.
`REDIS_URL` is a 12-factor fallback and never counts as a conflict.

**Resolution order**:
1. Check for explicit `backend` parameter in `@cache(backend=...)`
2. Check for module-level default via `set_default_backend()`
3. Create RedisBackend from environment variables (CACHEKIT_REDIS_URL > REDIS_URL)
1. Explicit `backend` parameter in `@cache(backend=...)`
2. Module-level default via `set_default_backend()` (checked at decoration, and
again at first call if still unset)
3. Environment auto-detection per the table above

## Performance Considerations

Expand Down
20 changes: 13 additions & 7 deletions docs/features/distributed-locking.md
Original file line number Diff line number Diff line change
Expand Up @@ -41,7 +41,7 @@ report = await get_report("2025-01-15")
> [!NOTE]
> Locking requires **both** of:
>
> 1. A backend implementing the `LockableBackend` protocol. `RedisBackend` and `CachekitIOBackend` (the SaaS backend behind `api.cachekit.io`) both do. `FileBackend` and pure-L1 (zero-config) caching don't — they silently skip lock acquisition; the function still works, just without stampede protection.
> 1. A backend implementing the `LockableBackend` protocol. `CachekitIOBackend` (the SaaS backend behind `api.cachekit.io`) does, and so does the Redis backend you get from env auto-detection or `RedisBackendProvider` (`PerRequestRedisBackend`). A `RedisBackend` you construct yourself and pass as `backend=` does **not** — it has no `acquire_lock`. Neither do `FileBackend` or pure-L1 (zero-config) caching. All of them silently skip lock acquisition; the function still works, just without stampede protection.
> 2. An **async** decorated function. Sync wrappers never take the lock path on any backend — see [Async-only](#async-only-sync-functions-are-never-lock-protected).

---
Expand Down Expand Up @@ -192,9 +192,11 @@ leaderboard = await get_leaderboard()
### With Redis Backend (Explicit)
```python notest
from cachekit import cache
from cachekit.backends.redis import RedisBackend
from cachekit.backends.redis.provider import RedisBackendProvider, tenant_context

backend = RedisBackend() # Implements LockableBackend
# PerRequestRedisBackend implements LockableBackend; a bare RedisBackend() does not.
tenant_context.set("default")
backend = RedisBackendProvider(redis_url="redis://localhost:6379").get_backend()

@cache(ttl=300, backend=backend)
async def generate_stats(date):
Expand Down Expand Up @@ -238,16 +240,20 @@ def cheap_lookup(x):
The `LockableBackend` protocol defines how backends provide distributed locking:

```python notest
async def acquire_lock(
def acquire_lock(
self,
key: str, # Bare cache key (same key as get/set); backend derives lock namespace
timeout: float, # How long to hold the lock (seconds)
blocking_timeout: Optional[float] = None, # Max wait to acquire (None = non-blocking)
) -> AsyncIterator[bool]:
# Yields True if lock acquired, False if timeout waiting
) -> AbstractAsyncContextManager[bool]:
# Entering the context yields True if the lock was acquired, False if the wait timed out
...
```

Implementations are `async` generators wrapped in `@asynccontextmanager`, so the
protocol declares the *decorated* shape — a backend author must apply the
decorator for `async with` to work.

The decorator wrapper calls it with `timeout=30.0` (lock self-expiry) and
`blocking_timeout=5.0` (max wait to acquire) — see
[Lock Timing](#lock-timing-30-second-expiry-5-second-wait).
Expand Down Expand Up @@ -332,7 +338,7 @@ A: Your function takes longer than the 5 s `blocking_timeout`, so waiters fall t
**Q: Locking doesn't seem to be working**
A: Two things to check:
1. The decorated function must be **async** — sync wrappers never lock ([Async-only](#async-only-sync-functions-are-never-lock-protected)).
2. The backend must implement `LockableBackend` (`RedisBackend`, `CachekitIOBackend`). Check with `from cachekit.backends.base import LockableBackend; isinstance(backend, LockableBackend)`.
2. The backend must implement `LockableBackend` (`CachekitIOBackend`, or `PerRequestRedisBackend` from `RedisBackendProvider` — a bare `RedisBackend()` does not). Check it the way the SDK does: `hasattr(backend, "acquire_lock")`. Avoid `isinstance(backend, LockableBackend)` — from CPython 3.12 a `runtime_checkable` protocol check resolves members with `inspect.getattr_static`, so it reports `False` for a backend that delegates `acquire_lock` through `__getattr__`.

**Q: How do I know if stampedes are happening?**
A: Check Prometheus: a spike in `rate(redis_cache_operations_total{status="miss"}[1m])` = stampede risk. See [Prometheus Metrics](prometheus-metrics.md).
Expand Down
20 changes: 13 additions & 7 deletions src/cachekit/backends/base.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,7 @@
from __future__ import annotations

from collections.abc import AsyncIterator, Callable
from contextlib import AbstractAsyncContextManager
from typing import Any, BinaryIO, Optional, Protocol, runtime_checkable

# Re-export BackendError for convenience (public API)
Expand Down Expand Up @@ -269,8 +270,11 @@ class LockableBackend(Protocol):
features like cache stampede prevention and critical sections.

Not all backends support this capability:
- Supported: RedisBackend, CachekitIOBackend (SaaS ``POST /v1/cache/{key}/lock``)
- Not supported: FileBackend, L1-only (in-memory)
- Supported: ``PerRequestRedisBackend`` — what ``RedisBackendProvider`` and
therefore the env-resolved Redis path hand out — and ``CachekitIOBackend``
(SaaS ``POST /v1/cache/{key}/lock``).
- Not supported: ``RedisBackend`` constructed directly and passed as
``backend=``, ``FileBackend``, L1-only (in-memory).

Contract — bare cache key:
``acquire_lock`` receives the **bare cache key**, identical to what
Expand All @@ -293,12 +297,12 @@ class LockableBackend(Protocol):
>>> # result = expensive_computation()
"""

async def acquire_lock(
def acquire_lock(
self,
key: str,
timeout: float,
blocking_timeout: Optional[float] = None,
) -> AsyncIterator[bool]:
) -> AbstractAsyncContextManager[bool]:
"""Acquire a distributed lock on key.

Args:
Expand All @@ -309,9 +313,11 @@ async def acquire_lock(
timeout: How long to hold the lock (seconds) before auto-release
blocking_timeout: Max time to wait for lock acquisition (None = non-blocking)

Yields:
True if lock was acquired
False if timeout occurred waiting for lock
Returns:
An async context manager yielding True if the lock was acquired,
False if the wait timed out. Implementations are ``async``
generators wrapped in ``@asynccontextmanager``, so this protocol
declares the *decorated* shape.

Raises:
BackendError: If backend operation fails
Expand Down
29 changes: 29 additions & 0 deletions src/cachekit/cache_handler.py
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,7 @@
BufferHandle,
BufferReadableBackend,
BufferWritableBackend,
LockableBackend,
TTLInspectableBackend,
)
from cachekit.backends.provider import (
Expand Down Expand Up @@ -192,6 +193,34 @@ def supports_ttl_inspection(backend: BaseBackend) -> TypeGuard[TTLInspectableBac
return hasattr(backend, "get_ttl") and hasattr(backend, "refresh_ttl")


def supports_locking(backend: object) -> TypeGuard[LockableBackend]:
"""Type guard: backend provides distributed locking (stampede prevention).

Takes ``object``, not ``BaseBackend`` like its siblings, because the
decorator probes the lazily-resolved ``_backend`` cell, which is ``None``
until first call and may already be narrowed by another capability guard.

Checked on the INSTANCE, deliberately — unlike ``supports_swr`` below, which
is class-level to keep ``__getattr__`` proxies and mocks off the freshness
read path. The asymmetry is the failure direction: a false negative silently
drops stampede protection on a hot key, while a false positive raises out of
the ``async with`` — and does NOT fail open. The wrapper's handler degrades
to lockless execution only for a ``BackendError``; a ``TypeError`` from
calling a non-callable propagates to the caller and breaks the decorated
function. Hence ``callable``, not ``hasattr``: a backend carrying
``acquire_lock = None`` is not lockable, and must take the lockless path
rather than crash the call it was meant to protect.

Equally deliberate: not ``isinstance(backend, LockableBackend)``. Since
CPython 3.12 a ``runtime_checkable`` Protocol check resolves members with
``inspect.getattr_static``, which does not consult ``__getattr__`` — so a
delegating backend proxy locks on 3.10/3.11 and silently stops locking on
3.12+. Plain ``getattr`` consults ``__getattr__`` and is stable across
every supported interpreter.
"""
return callable(getattr(backend, "acquire_lock", None))


# Backend type names already warned about, so refresh_ttl_on_get degradation warns at most
# once per backend type per process (avoids per-hit log spam). Tests clear this set.
_TTL_REFRESH_UNSUPPORTED_WARNED: set[str] = set()
Expand Down
6 changes: 5 additions & 1 deletion src/cachekit/decorators/intent.py
Original file line number Diff line number Diff line change
Expand Up @@ -137,7 +137,11 @@ def decorator(f: F) -> F:
backend = manual_overrides.pop("backend", None)

# Tier 2 resolution: if no explicit backend and not L1-only mode,
# check module-level default set via set_default_backend()
# check module-level default set via set_default_backend(). Kept here
# (not only lazily) because decoration-time validation — the interop
# backend guard and stale_ttl/SWR capability (LAB-557) — needs the
# backend when it is already known. If the default is set LATER, the
# wrapper re-consults it at first call (_resolve_lazy_backend, LAB-4457).
if backend is None and not _explicit_l1_only:
from ..config.decorator import get_default_backend

Expand Down
38 changes: 27 additions & 11 deletions src/cachekit/decorators/wrapper.py
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,9 @@
get_logger,
handle_decrypt_failure,
redact_cache_key,
supports_locking,
supports_swr,
supports_ttl_inspection,
warn_ttl_refresh_unsupported,
)
from ..interop import (
Expand All @@ -46,8 +48,23 @@
from .tenant_context import TenantContextExtractor

if TYPE_CHECKING:
from ..backends.base import BaseBackend
from ..serializers.base import SerializerProtocol


def _resolve_lazy_backend() -> BaseBackend:
"""Backend for a decorator that was applied without ``backend=``.

Consulted at FIRST CALL, not at decoration, so ``set_default_backend()``
takes effect regardless of whether it ran before or after the module holding
the decorated function was imported (LAB-4457).
"""
from ..config.decorator import get_default_backend

default = get_default_backend()
return default if default is not None else get_backend_provider().get_backend()


F = TypeVar("F", bound=Callable[..., Any])

_logger = logging.getLogger(__name__)
Expand Down Expand Up @@ -907,9 +924,8 @@ async def _l2_swr_revalidate_async(cache_key: str, call_args: tuple[Any, ...], c
the caller already got the stale value; the entry hard-expires at evict_at and
the next request takes the ordinary synchronous miss path (spec degradation)."""
try:
_acquire_lock = getattr(_backend, "acquire_lock", None)
if _acquire_lock is not None:
async with _acquire_lock(cache_key, timeout=_l2_swr_lease_seconds, blocking_timeout=None) as got_lease:
if supports_locking(_backend):
async with _backend.acquire_lock(cache_key, timeout=_l2_swr_lease_seconds, blocking_timeout=None) as got_lease:
if not got_lease:
return # another client is revalidating — stale already served
await _l2_swr_recompute_store_async(cache_key, call_args, call_kwargs)
Expand Down Expand Up @@ -1263,7 +1279,7 @@ def sync_wrapper(*args: Any, **kwargs: Any) -> Any: # noqa: PLR0912

nonlocal _backend
if _backend is None:
_backend = get_backend_provider().get_backend()
_backend = _resolve_lazy_backend()

# Setup cache handler strategy on first use
handler = StandardCacheHandler(
Expand Down Expand Up @@ -1672,7 +1688,7 @@ async def async_wrapper(*args: Any, **kwargs: Any) -> Any:
if interop is not None:
if _backend is None:
try:
_backend = get_backend_provider().get_backend()
_backend = _resolve_lazy_backend()
except Exception as e:
# If Redis connection fails, execute function without caching - RETURN EARLY
# This prevents the decorator from breaking the application
Expand Down Expand Up @@ -1746,7 +1762,7 @@ async def async_wrapper(*args: Any, **kwargs: Any) -> Any:
# Initialize backend only when needed (lazy init for performance)
if _backend is None:
try:
_backend = get_backend_provider().get_backend()
_backend = _resolve_lazy_backend()
except Exception as e:
# If Redis connection fails, execute function without caching - RETURN EARLY
# This prevents the decorator from breaking the application
Expand Down Expand Up @@ -1805,7 +1821,7 @@ async def async_wrapper(*args: Any, **kwargs: Any) -> Any:
_l1_backfill_from_l2(cache_key, cached_data, _l2_is_stale, _l2_fresh_for)

# Handle TTL refresh if configured and threshold met
if refresh_ttl_on_get and ttl and hasattr(_backend, "get_ttl") and hasattr(_backend, "refresh_ttl"):
if refresh_ttl_on_get and ttl and supports_ttl_inspection(_backend):
try:
remaining_ttl = await _backend.get_ttl(cache_key)
if remaining_ttl and remaining_ttl < (ttl * ttl_refresh_threshold):
Expand Down Expand Up @@ -1859,7 +1875,7 @@ async def async_wrapper(*args: Any, **kwargs: Any) -> Any:
blocking_timeout = 5.0 # Wait up to 5 seconds to acquire lock

# Check if backend supports distributed locking
if hasattr(_backend, "acquire_lock"):
if supports_locking(_backend):
try:
# Use backend's async lock protocol
async with _backend.acquire_lock(
Expand Down Expand Up @@ -2003,7 +2019,7 @@ async def async_wrapper(*args: Any, **kwargs: Any) -> Any:
# Fall through to execute without locking

# Execute without locking (either backend doesn't support it or lock failed)
if not hasattr(_backend, "acquire_lock"):
if not supports_locking(_backend):
logger().debug(
f"Backend doesn't support locking for {redact_cache_key(cache_key)}, executing without thundering herd protection"
)
Expand Down Expand Up @@ -2078,7 +2094,7 @@ def invalidate_cache(*args: Any, **kwargs: Any) -> None:
# we should NOT try to get a backend from the provider
if not _l1_only_mode and _backend is None:
try:
_backend = get_backend_provider().get_backend()
_backend = _resolve_lazy_backend()
except Exception as e:
# If backend creation fails, can't invalidate L2
_logger.debug("Failed to get backend for invalidation: %s", redact_error_for_log(e))
Expand Down Expand Up @@ -2138,7 +2154,7 @@ async def ainvalidate_cache(*args: Any, **kwargs: Any) -> None:
# we should NOT try to get a backend from the provider
if not _l1_only_mode and _backend is None:
try:
_backend = get_backend_provider().get_backend()
_backend = _resolve_lazy_backend()
except Exception as e:
# If backend creation fails, can't invalidate L2
_logger.debug("Failed to get backend for async invalidation: %s", redact_error_for_log(e))
Expand Down
Loading
Loading