REST API
Answers, chat and feedback
Ask a document base a question and get an answer with the documents it drew on. Non-streaming or token by token, from your own backend or from an agent.
Asking is one call. You send a question, you get an answer plus the documents it drew on, and the answer comes from memory that is already resident rather than from a retrieval pipeline you have to run.
Ask a question
curl -X POST https://api.engramdynamics.org/v1/corpora/c_7a1f.../chat \
-H "Authorization: Bearer <your key>" \
-H "Content-Type: application/json" \
-d '{
"question": "What is the refund window for EU customers?",
"k": 3
}'
{
"answer": "EU customers have 14 days from delivery to request a refund ...",
"used_docs": ["t_9a1..._handbook-ch3"],
"sources": [{"id": "t_9a1..._handbook-ch3", "title": "handbook/ch3.md"}],
"used_documents": [
{"id": "d_92f...", "filename": "handbook/ch3.md", "title": "handbook/ch3.md", "sections": []}
],
"condensed_q": null,
"documents_warming": 0,
"documents_ready": 2,
"documents_onboarding": 0,
"documents_failed": 0,
"last_synced_at": "2026-09-13T10:11:40Z"
}
Needs the query scope. The request body:
| Field | Default | What it does |
|---|---|---|
question | required | Up to 4,000 characters. |
k | 3 | How many documents to draw on, 1 to 20. |
history | empty | Prior turns of this conversation, oldest first, up to 20 turns of 4,000 characters. The model remembers the chat and the documents. |
doc_ids | null | Pin the answer to specific documents, using ids
you saw in used_docs. Up to 20. |
max_tokens | null | Answer budget. Clamped to the server ceiling, never raised past it. |
Python, with follow-up turns:
history = []
def ask(question):
body = {"question": question, "k": 3, "history": history}
res = requests.post(f"{BASE}/corpora/{base_id}/chat", headers=HEADERS, json=body, timeout=120)
res.raise_for_status()
out = res.json()
history.extend([{"role": "user", "content": question},
{"role": "assistant", "content": out["answer"]}])
return out
answer = ask("What is the refund window for EU customers?")
print(answer["answer"])
for source in answer["sources"]:
print(" source:", source["title"])
Stream the answer
Same request body, server-sent events, so the answer renders as it is written:
curl -N -X POST https://api.engramdynamics.org/v1/corpora/c_7a1f.../chat/stream \
-H "Authorization: Bearer <your key>" \
-H "Content-Type: application/json" \
-d '{"question": "What is the refund window for EU customers?"}'
Four frame types, in this order:
| Frame | When | Carries |
|---|---|---|
head | Once, first | used_docs,
sources and the readiness counts, so a UI can render citations before the first
word arrives. |
delta | Many | {"delta": "text"} as the model
writes. |
done | Once, last | Measured metrics for the turn. |
error | Instead of the rest | A failure that happened after the 200 had already gone out. Handle it in band. |
const res = await fetch(`${BASE}/corpora/${baseId}/chat/stream`, {
method: "POST",
headers: { ...headers, "Content-Type": "application/json" },
body: JSON.stringify({ question: "What is the refund window for EU customers?" }),
});
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buffer = "";
while (true) {
const { value, done } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const frames = buffer.split("\n\n");
buffer = frames.pop();
for (const frame of frames) {
const payload = JSON.parse(frame.replace(/^data: /, ""));
if (payload.head) console.log("sources", payload.sources);
else if (payload.delta) process.stdout.write(payload.delta);
else if (payload.error) throw new Error(payload.error);
}
}
What the answer cites
Three fields describe the evidence, and they answer different questions.
used_docsis the serving currency: the memory ids the answer was composed from. Pass any of them back asdoc_idson the next turn to pin the conversation to them.sourcesis the same list with titles, which is what you show a person.used_documentsis the file-level view: the document ids and filenames, each with the sections that answered when a long file was split. Cite the file the customer uploaded from here.
Two more fields are worth reading. condensed_q shows the standalone question a
follow-up turn was rewritten to before retrieval, which is the first thing to look at when a
follow-up answers oddly. documents_warming counts documents that were retrieved but
whose memory was not resident this turn; they are being warmed and the next turn includes them, so an
agent that wants the full evidence simply asks again.
The readiness counts ride on every answer, so an agent that just added files can see that three documents are still onboarding rather than concluding the answer is wrong.
When serving is busy
One shape covers every serving-side failure, so a client has one thing to handle:
{
"error": {
"code": "serving_unavailable",
"message": "Serving is busy right now. Your documents are safe and nothing was lost; ask again shortly.",
"status": 503,
"details": {"retry_after": 20, "state": "warming"},
"request_id": "b41d..."
}
}
Retry-After is on the response as a header too, and that is what a backing-off client
should read. state is for a human: offline, booting,
provisioning, warming, degraded or unknown.
Ingestion is never affected by this. Uploads, hashes, diffs, source syncs and parsing keep accepting work through a serving outage and the queue drains when serving is back. "Serving is busy" never means "we lost your documents".
Rate an answer
Answer quality is worth measuring, and the rating control in the app writes through this route. Use it from your own client and the same dashboards pick it up.
curl -X POST https://api.engramdynamics.org/corpora/c_7a1f.../feedback \
-H "Authorization: Bearer <your key>" \
-H "Content-Type: application/json" \
-d '{
"rating": "accurate",
"question": "What is the refund window for EU customers?",
"answer_excerpt": "EU customers have 14 days from delivery ..."
}'
rating is accurate, partial or inaccurate.
question and answer_excerpt are capped at 2,000 characters each: send a
snippet, not the whole transcript.
Feedback is one of a few routes that stay on the unversioned path only, so the
URL has no /v1 in it. It belongs to the console rather than to the published contract,
and it is listed that way in the endpoint index.
Ask over MCP
The MCP query route answers the same question and attributes its sources with the pinnable doc ids
agents pass back as doc_ids. It takes the same workspace API key:
curl -X POST https://api.engramdynamics.org/v1/mcp/c_7a1f.../query \
-H "Authorization: Bearer <your key>" \
-H "Content-Type: application/json" \
-d '{"question": "What is the refund window for EU customers?"}'
Two introspection routes sit next to it: GET /v1/mcp/{id}/meta for the base's name,
status and catalog description, and GET /v1/mcp/{id}/documents for the inventory. Both
return ids in the same pinnable currency as used_docs. Most agents never call these
directly: Engram's hosted MCP server calls them on the agent's behalf, which is on
MCP quick start.
Next
Every route in one table: Endpoint index.