Skip to content

Staff Engineer Guide ​

Table of Contents ​

The Core Insight ​

A two-level cache trades staleness for latency. Instead of claiming coherence, CoCache states three properties and builds each one from a specific mechanism:

PropertyMechanism
Bounded staleness: every inconsistency self-heals in finite timeFinite default ttl (3600 s); separate short missingTtl (60 s); L2 cleared on every event-channel (re)subscription
No lost invalidation: concurrent loads and out-of-order events never resurrect old valuesInvalidationStamps guard every write-back; an L1 miss is never inferred as "not found"
Cheap hit pathBounded Caffeine L2 (expiry checked on read, cached CacheClock); one atomic Lua round trip per L1 read; per-key SingleFlight

No distributed locks and no consensus are involved. Each instance coalesces its own loads, and L1 absorbs duplicates across instances.

Consistency Model ​

mermaid
graph TB
    subgraph writers ["Write side (any instance)"]
        W["update DB → evict(key)"]
    end
    subgraph readers ["Read side (every instance)"]
        L2["L2 copy"]
    end
    W -->|"1. bump stamp, delete L2, delete L1"| L1[("L1")]
    W -->|"2. publish key@@clientId"| PS(["Pub/Sub"])
    PS -->|"onEvicted: bump stamp, delete L2"| L2
    PS -.->|"reconnect: onReset clears L2"| L2
    L1 -->|"stamp-guarded fill"| L2

    style W fill:#2d333b,stroke:#6d5dfc,color:#e6edf3
    style L2 fill:#2d333b,stroke:#6d5dfc,color:#e6edf3
    style L1 fill:#2d333b,stroke:#6d5dfc,color:#e6edf3
    style PS fill:#2d333b,stroke:#6d5dfc,color:#e6edf3
    style writers fill:#161b22,stroke:#8b949e,color:#e6edf3
    style readers fill:#161b22,stroke:#8b949e,color:#e6edf3

Staleness bounds per failure mode:

FailureStale for at mostWhy
Event delayedPub/Sub latencyonEvicted evicts L2; a fill that raced with the event is undone
Event lost while connectedttlThe connection is TCP, so loss implies a disconnect, which triggers a reset. Otherwise the TTL bounds it.
Disconnect / reconnectReconnect timeonReset clears L2 after resubscription
Slow load vs concurrent update (cache-aside race)ttl in L1Inherent to cache-aside; undone in L2 when the event arrives, persists in L1 until expiry
Row created after a negative lookupmissingTtlNegative entries use their own TTL

Writers must update the source, then evict. Prefer evict over set unless the writer has the committed value at hand.

Read Path ​

mermaid
sequenceDiagram
autonumber
    participant App
    participant C as DefaultCoherentCache
    participant SF as SingleFlight
    participant L1 as Redis
    participant Src as CacheSource

    App->>C: getCache(key)
    C->>C: L2 hit? → return
    C->>C: keyFilter.notExist? → MissingValue
    C->>SF: execute(cacheKey)
    SF->>SF: stamp = current(cacheKey)
    SF->>L1: EVALSHA read-script → {TTL, value}
    alt hit
        SF->>C: fill L2 iff stamp unchanged (re-check after)
    else miss
        SF->>Src: loadCacheValue(key)
        Src-->>SF: value | null → MissingValue(missingTtl)
        SF->>L1: write iff stamp unchanged
        SF->>C: write L2; if stamp changed → undo L1+L2, publish
    end
    SF-->>App: value (followers share it)

Design Tradeoffs ​

Per-key SingleFlight vs distributed locks vs striped locks ​

SingleFlight (chosen)Striped locks (4.x)Distributed lock
ScopePer instance, exact keyPer instance, hashed stripeCluster
Unrelated keys block each otherNeverOn stripe collisionNo
Failure propagationLeader's original exception to all waitersEach waiter retries the loadLock TTL management
CostOne ConcurrentHashMap entry while in flightFixed lock arrayNetwork round trip

Cross-instance duplicate loads are accepted: they are rare, idempotent, and absorbed by L1.

Stamps vs per-key generation map vs locking writers ​

Stamps are a fixed AtomicLongArray(4096). They need no allocation per key and no cleanup, and they cover every invalidation source, local ones included. A stripe collision only costs an extra skipped write-back.

Reset-on-subscribe vs reliable messaging ​

Redis Pub/Sub is at-most-once. Rather than adding a durable broker, CoCache treats every (re)subscription as "messages may have been lost" and drops L2. The cost is a brief hit-rate dip after a reconnect.

Explicit negative type vs in-band sentinel ​

4.x detected negative entries by the shape of the value ("_nil_", {"_nil_"} ...), so real data could be misread as "not found". 5.0 uses a sealed MissingValue; the sentinel exists only as the Redis wire encoding.

JDK proxy vs AOP ​

Caches are declared at the interface level (UserCache : Cache<String, User>), so a JDK proxy is enough. One CacheInvocationHandler dispatches calls and rethrows original exceptions.

Extension Points ​

ExtensionSPI (cocache-api)DefaultsSpring override
L2ClientSideCache<V>CaffeineClientSideCachebean {cacheName}.ClientSideCache
L1DistributedCache<V>RedisDistributedCachebean {cacheName}.DistributedCache
ChannelCacheEvictedEventBusRedisCacheEvictedEventBus@Bean CacheEvictedEventBus (must call onReset on (re)subscription)
SourceCacheSource<K, V>noOp()bean {cacheName}.CacheSource or unique typed bean
Key conversionKeyConverter<K>ToStringKeyConverter / ExpKeyConverterbean {cacheName}.KeyConverter
Existence filterKeyFilterKeyFilter.NO_OPvia CoherentCacheConfiguration

Verify any new implementation with the matching TCK spec in cocache-test.

Performance Characteristics ​

PathLatencyNotes
L2 hit~100 ns – 1 µsCaffeine lookup + expiry check
L1 hit~0.5 – 2 msOne Lua round trip
L0 loadsource latencyCoalesced per key
Write / evict~1 RTT + publishPublish is fire-and-forget

Memory: L2 is bounded by maximumSize (default 10 000 entries per cache). Expired entries are evicted when read, or by size-based eviction. In-flight loads cost one map entry each.

Fan-out: with N instances and W writes/s per cache, Pub/Sub delivers N·W messages/s on that cache's channel. Successful loads publish nothing.

Operational Considerations ​

Redis Dependency ​

  • Read failures degrade to source loads; write failures log a warning (cocache.redis.strict-failure=false).
  • During an outage, L2 keeps serving what it holds until ttl. After recovery, resubscription clears L2.

Monitoring ​

SignalWhereWatch for
L2 size/actuator/cocacheClient/{name}Sustained maximumSize → raise it or shorten ttl
Cache composition and ttlPolicy/actuator/cocache/{name}Unexpected defaults
Channel[...] subscribed - reset subscriber (INFO)logsFrequent resets = unstable Redis connection
Discard the loaded value ... (WARN)logsHigh rate = write-heavy keys contending with loads

TTL Strategy ​

  • ttl: the longest staleness you accept for a missed invalidation. Keep it finite.
  • ttlAmplitude: about 2–10 % of ttl.
  • missingTtl: short (seconds to a minute); it bounds the invisibility of new rows.

Released under the Apache License 2.0.