Sep 2026 · 5 min read
Cutting p99 latency from 900ms to 350ms with a cache-aside Redis layer
A mapping workflow looked up the same reference codes on every request. Measuring first, then adding a cache-aside layer, cut database hits by 38%.
A request/response mapping workflow in an integration service was slow at the tail. Averages looked fine, but p99 latency sat around 900ms, and that is what users actually feel when a batch of requests goes through.
Measure before changing anything
Tracing a slow request showed the same pattern every time: the workflow resolved a handful of reference codes, and each one was a separate database round trip. The codes change rarely, but they were read on every request.
- Many small, repeated reads of rarely changing data
- Each read paid a full database round trip
- Tail latency grew with database load, not with the work itself
Cache-aside
Cache-aside keeps the database as the source of truth. The application checks Redis first, falls back to the database on a miss, and writes the result back with a TTL.
public String resolveCode(String key) {
String cached = redis.get("code:" + key);
if (cached != null) {
return cached; // hit: no database round trip
}
String value = repository.findCode(key); // miss: read the source of truth
redis.setex("code:" + key, 3600, value); // write back with a 1-hour TTL
return value;
}Choosing the TTL and invalidation
Because the codes change rarely, a TTL measured in hours is safe, and an explicit delete on update keeps the cache correct when a code does change. That keeps the design simple: no cache warming, no distributed locks.
Result
- Database hits on the workflow dropped by 38%
- p99 latency fell from 900ms to 350ms
- The database has more headroom for writes