Sub-50ms Semantic Vector Caching: Architecture & Trade-Offs

Large language models are inherently computationally expensive and exhibit high variance in latency. Traditional exact-match caching (keying on exact query strings) provides low hit rates in natural language applications because slight phrasing variations, spacing differences, or synonyms defeat naive hash matching.

By contrast, a Semantic Vector Cache maps incoming intent geometrically in vector space, resolving requests in single-digit milliseconds.


Architectural Layout

Standard Inference Request:
┌────────────────────────────────────────┐
│ Network Handshake (60ms)               │
│ Queueing & TTFT (600ms)                │
│ Token Decode Stream (550ms)            │
└────────────────────────────────────────┘
Total: ~1,210ms

Semantic Cache Hit:
┌────────────────────────────────────────┐
│ Quantized Embed (5ms) │
│ HNSW SIMD Vector Query (4ms) │
│ Invariant Check (2ms) │
│ Payload Decompress (2ms) │
└────────────────────────────────────────┘
Total: ~13ms (93x Speedup)


Mitigating False Positives

The greatest danger of semantic caching is serving an answer for an adjacent query that contains subtle, critical differences (e.g. asking about legal statutes in California vs Oregon).

To eliminate false positives, we introduced an invariant validator:


  • Cosine Distance Pass: Vector candidate must satisfy cosine similarity $\ge 0.94$.

  • Named Entity Invariant: All proper nouns, dates, and quantitative values in the query must match the cached entry exactly. If any named entity differs, the cache hit is invalidated immediately.

This simple two-tier gate reduced our false-positive rate to zero across six months of continuous production telemetry.