Engram Smart CAG is a Cache Augmented Generation solution that gives your fleet durable, indexed document memory — your document base cache, made permanent. Packaged as a pip-installable vLLM plug-in, Engram Smart CAG enables serving engines to access a document library faster and cheaper than existing RAG solutions, with the same answer quality.
Measured head-to-head — same model, same retrieval
See the numbers →Text re-processed per question
Time to first token
Drops into the serving stack you already run
pip install — a stock vLLM connector. No fork of your serving stack, no new inference engine.
One forward pass per document — no training — then every answer re-processes dozens of tokens instead of thousands. Measured 3× the throughput per GPU at fleet scale vs same-model RAG — a fleet freed from re-reading serves more customers on the same GPUs.
Responses grounded in the source documents, with citations. In a head-to-head at 1,000-document scale: answer quality statistically identical to RAG on the same model and retrieval, evaluated by a second, independent judge. Quality you can put in front of your customers.
A document's memory is already resident — first token in ~30 ms whether a tenant's library holds ten documents or ten thousand. Memories page across the fleet and multiplex on one shared base model.
How it works
A stock vLLM connector — pip install into the fleet you already run.
One forward pass per document saves its KV memory. No training.
Bring your retrieval or use ours — it picks which memories to load, by document id.
The model answers from resident memory — dozens of tokens prefilled, not thousands.
We're onboarding design partners — inference providers and platform teams running vLLM. We'll walk your team through the integration and the measured numbers.
No spam. We'll reach out to schedule.