REST API

Answers, chat and feedback

Ask a document base a question and get an answer with the documents it drew on. Non-streaming or token by token, from your own backend or from an agent.

Asking is one call. You send a question, you get an answer plus the documents it drew on, and the answer comes from memory that is already resident rather than from a retrieval pipeline you have to run.

Ask a question

curl -X POST https://api.engramdynamics.org/v1/corpora/c_7a1f.../chat \
  -H "Authorization: Bearer <your key>" \
  -H "Content-Type: application/json" \
  -d '{
        "question": "What is the refund window for EU customers?",
        "k": 3
      }' 
{
  "answer": "EU customers have 14 days from delivery to request a refund ...",
  "used_docs": ["t_9a1..._handbook-ch3"],
  "sources": [{"id": "t_9a1..._handbook-ch3", "title": "handbook/ch3.md"}],
  "used_documents": [
    {"id": "d_92f...", "filename": "handbook/ch3.md", "title": "handbook/ch3.md", "sections": []}
  ],
  "condensed_q": null,
  "documents_warming": 0,
  "documents_ready": 2,
  "documents_onboarding": 0,
  "documents_failed": 0,
  "last_synced_at": "2026-09-13T10:11:40Z"
}

Needs the query scope. The request body:

FieldDefaultWhat it does
questionrequiredUp to 4,000 characters.
k3How many documents to draw on, 1 to 20.
historyemptyPrior turns of this conversation, oldest first, up to 20 turns of 4,000 characters. The model remembers the chat and the documents.
doc_idsnullPin the answer to specific documents, using ids you saw in used_docs. Up to 20.
max_tokensnullAnswer budget. Clamped to the server ceiling, never raised past it.

Python, with follow-up turns:

history = []

def ask(question):
    body = {"question": question, "k": 3, "history": history}
    res = requests.post(f"{BASE}/corpora/{base_id}/chat", headers=HEADERS, json=body, timeout=120)
    res.raise_for_status()
    out = res.json()
    history.extend([{"role": "user", "content": question},
                    {"role": "assistant", "content": out["answer"]}])
    return out

answer = ask("What is the refund window for EU customers?")
print(answer["answer"])
for source in answer["sources"]:
    print(" source:", source["title"])

Stream the answer

Same request body, server-sent events, so the answer renders as it is written:

curl -N -X POST https://api.engramdynamics.org/v1/corpora/c_7a1f.../chat/stream \
  -H "Authorization: Bearer <your key>" \
  -H "Content-Type: application/json" \
  -d '{"question": "What is the refund window for EU customers?"}' 

Four frame types, in this order:

FrameWhenCarries
headOnce, firstused_docs, sources and the readiness counts, so a UI can render citations before the first word arrives.
deltaMany{"delta": "text"} as the model writes.
doneOnce, lastMeasured metrics for the turn.
errorInstead of the restA failure that happened after the 200 had already gone out. Handle it in band.
const res = await fetch(`${BASE}/corpora/${baseId}/chat/stream`, {
  method: "POST",
  headers: { ...headers, "Content-Type": "application/json" },
  body: JSON.stringify({ question: "What is the refund window for EU customers?" }),
});

const reader = res.body.getReader();
const decoder = new TextDecoder();
let buffer = "";

while (true) {
  const { value, done } = await reader.read();
  if (done) break;
  buffer += decoder.decode(value, { stream: true });
  const frames = buffer.split("\n\n");
  buffer = frames.pop();
  for (const frame of frames) {
    const payload = JSON.parse(frame.replace(/^data: /, ""));
    if (payload.head) console.log("sources", payload.sources);
    else if (payload.delta) process.stdout.write(payload.delta);
    else if (payload.error) throw new Error(payload.error);
  }
}

What the answer cites

Three fields describe the evidence, and they answer different questions.

Two more fields are worth reading. condensed_q shows the standalone question a follow-up turn was rewritten to before retrieval, which is the first thing to look at when a follow-up answers oddly. documents_warming counts documents that were retrieved but whose memory was not resident this turn; they are being warmed and the next turn includes them, so an agent that wants the full evidence simply asks again.

The readiness counts ride on every answer, so an agent that just added files can see that three documents are still onboarding rather than concluding the answer is wrong.

When serving is busy

One shape covers every serving-side failure, so a client has one thing to handle:

{
  "error": {
    "code": "serving_unavailable",
    "message": "Serving is busy right now. Your documents are safe and nothing was lost; ask again shortly.",
    "status": 503,
    "details": {"retry_after": 20, "state": "warming"},
    "request_id": "b41d..."
  }
}

Retry-After is on the response as a header too, and that is what a backing-off client should read. state is for a human: offline, booting, provisioning, warming, degraded or unknown.

Ingestion is never affected by this. Uploads, hashes, diffs, source syncs and parsing keep accepting work through a serving outage and the queue drains when serving is back. "Serving is busy" never means "we lost your documents".

Rate an answer

Answer quality is worth measuring, and the rating control in the app writes through this route. Use it from your own client and the same dashboards pick it up.

curl -X POST https://api.engramdynamics.org/corpora/c_7a1f.../feedback \
  -H "Authorization: Bearer <your key>" \
  -H "Content-Type: application/json" \
  -d '{
        "rating": "accurate",
        "question": "What is the refund window for EU customers?",
        "answer_excerpt": "EU customers have 14 days from delivery ..."
      }' 

rating is accurate, partial or inaccurate. question and answer_excerpt are capped at 2,000 characters each: send a snippet, not the whole transcript.

Feedback is one of a few routes that stay on the unversioned path only, so the URL has no /v1 in it. It belongs to the console rather than to the published contract, and it is listed that way in the endpoint index.

Ask over MCP

The MCP query route answers the same question and attributes its sources with the pinnable doc ids agents pass back as doc_ids. It takes the same workspace API key:

curl -X POST https://api.engramdynamics.org/v1/mcp/c_7a1f.../query \
  -H "Authorization: Bearer <your key>" \
  -H "Content-Type: application/json" \
  -d '{"question": "What is the refund window for EU customers?"}' 

Two introspection routes sit next to it: GET /v1/mcp/{id}/meta for the base's name, status and catalog description, and GET /v1/mcp/{id}/documents for the inventory. Both return ids in the same pinnable currency as used_docs. Most agents never call these directly: Engram's hosted MCP server calls them on the agent's behalf, which is on MCP quick start.

Next

Every route in one table: Endpoint index.