Read a document once.
Serve it forever.

Engram Smart CAG is a Cache Augmented Generation solution that gives your fleet durable, indexed document memory — your document base cache, made permanent. Packaged as a pip-installable vLLM plug-in, Engram Smart CAG enables serving engines to access a document library faster and cheaper than existing RAG solutions, with the same answer quality.

Measured head-to-head — same model, same retrieval

See the numbers →

Text re-processed per question

Engram Smart CAG15 tokens
Standard RAG1,352 tokens

Time to first token

Engram Smart CAG28 ms
Standard RAG276 ms
91×
less text re-processed per question
10×
faster to the first token
Same quality
as a standard RAG architecture
Drop-in
stock vLLM connector, your VPC

Drops into the serving stack you already run

vLLM plug-in Open-weight models (Qwen3, Llama) OpenAI-compatible serving S3-tiered KV store Your VPC, your GPUs

pip install — a stock vLLM connector. No fork of your serving stack, no new inference engine.

$

Capacity back, every query

One forward pass per document — no training — then every answer re-processes dozens of tokens instead of thousands. Measured 3× the throughput per GPU at fleet scale vs same-model RAG — a fleet freed from re-reading serves more customers on the same GPUs.

Same answers, higher efficiency

Responses grounded in the source documents, with citations. In a head-to-head at 1,000-document scale: answer quality statistically identical to RAG on the same model and retrieval, evaluated by a second, independent judge. Quality you can put in front of your customers.

Flat latency at scale

A document's memory is already resident — first token in ~30 ms whether a tenant's library holds ten documents or ten thousand. Memories page across the fleet and multiplex on one shared base model.

How it works

From pip install to grounded answers.

See how it works →
1

Install

A stock vLLM connector — pip install into the fleet you already run.

2

Read once

One forward pass per document saves its KV memory. No training.

3

Retrieve

Bring your retrieval or use ours — it picks which memories to load, by document id.

4

Answer

The model answers from resident memory — dozens of tokens prefilled, not thousands.

Read once. Serve forever.

We're onboarding design partners — inference providers and platform teams running vLLM. We'll walk your team through the integration and the measured numbers.

No spam. We'll reach out to schedule.