Economics
Pick a plan with a clear document and question allowance and pay one price a month. Onboarding is free, we never charge for input or context tokens, and you are never billed past your allowance.
See the plans ↓The re-reading tax
Set your monthly queries to see what re-sending the same context costs at list prices, next to the Engram CaaS plan that covers the same volume. Engram never charges for context tokens on any plan.
Every query resends the same documents. The model has no memory between calls, so you buy the same read again and again.
Context per query is fixed at , what a tuned retrieval stack sends per query, as measured.
The Engram CaaS row is the monthly price of the smallest plan whose question allowance covers your queries; past the largest plan it needs a custom plan. Competitor figures are input-token list prices as published, charged on the re-sent context only. Measured 30 September 2026 on the production serving stack, 2x H100, the same model both sides: Engram cartridges against a tuned retrieval stack that chunks, reranks and budgets its context, multi turn conversations, zero errors at every load, judged accuracy at parity.
Flat vs per-token
Your Engram plan is one price a month. A per-token bill grows with every query. Set your library size and monthly questions to see the plan that fits and what the same work costs per-token.
The Engram line is the flat monthly price of the smallest plan whose document and question allowance covers your sliders. It stops where the largest plan's allowance ends, because no published plan covers the shaded range. The per-token line is a per-token model at list price, charged on every query. Below the crossover a per-token model is cheaper today, and we say so.
Performance
Answers start streaming in well under a second and stay that way as more people and agents send queries at once. Retrieval based AI keeps up only while the server is idle: under real load its re-reading queues up, its first response takes several times longer and its full answer more than twice as long, while Engram keeps serving from memory. Measured on a single dedicated serving tier.
Measured on the production serving stack as concurrent queries climb from one to sixty four at once: one document base, multi turn conversations, against a tuned retrieval pipeline on identical hardware. The chart shows when the response begins streaming on screen, measured at the client, so it includes retrieval on both sides. The hover readout also shows the model's own time to first token and when the answer completes.
Plans
Every plan includes a set number of documents and questions a month.
How billing works
One price a month, an allowance you can see, and no surprises on the invoice.
Reach your document or question allowance and the workspace pauses that meter with a clear message. You are never billed for going over, so the number on the plan is the number on the invoice.
Move up a plan in the billing portal and the larger allowance applies right away. Questions reset at the start of each month; freeing documents frees space immediately.
Start free with no card, then subscribe, change or cancel yourself, with no order form and no sales call. Pay monthly, or annually at two months free.