← All notes

Sep 2026 · 5 min read

Cutting p99 latency from 900ms to 350ms with a cache-aside Redis layer

A mapping workflow looked up the same reference codes on every request. Measuring first, then adding a cache-aside layer, cut database hits by 38%.

RedisPerformanceCaching

A request/response mapping workflow in an integration service was slow at the tail. Averages looked fine, but p99 latency sat around 900ms, and that is what users actually feel when a batch of requests goes through.

Measure before changing anything

Tracing a slow request showed the same pattern every time: the workflow resolved a handful of reference codes, and each one was a separate database round trip. The codes change rarely, but they were read on every request.

  • Many small, repeated reads of rarely changing data
  • Each read paid a full database round trip
  • Tail latency grew with database load, not with the work itself

Cache-aside

Cache-aside keeps the database as the source of truth. The application checks Redis first, falls back to the database on a miss, and writes the result back with a TTL.

java
public String resolveCode(String key) {
    String cached = redis.get("code:" + key);
    if (cached != null) {
        return cached;                          // hit: no database round trip
    }
    String value = repository.findCode(key);    // miss: read the source of truth
    redis.setex("code:" + key, 3600, value);    // write back with a 1-hour TTL
    return value;
}

Choosing the TTL and invalidation

Because the codes change rarely, a TTL measured in hours is safe, and an explicit delete on update keeps the cache correct when a code does change. That keeps the design simple: no cache warming, no distributed locks.

Result

  • Database hits on the workflow dropped by 38%
  • p99 latency fell from 900ms to 350ms
  • The database has more headroom for writes