REST API
Documents and ingest
Load documents the way that fits your pipeline: one call for a single document, a four-step protocol for a folder that has to stay in sync, or a source that pulls so the bytes never touch your machine.
There are three ways to get documents in, and the right one depends on how many files you have and where they live. All three end in the same place: a document row, extracted text, and a build queued.
Three ways in
| You have | Use |
|---|---|
| One document, content already in hand | Upsert, a single call |
| A folder, a large file, or a sync that has to be exact | The ingest protocol, four steps |
| A bucket, a Drive folder or a SharePoint library | Sources, so Engram pulls and you never move the bytes |
The CLI implements the protocol below for you, so
engram push ./docs is the shortest path if you are not building an integration.
The ingest protocol
Four calls, in order. The shape exists to keep the bytes off the API and the round trips off your pipeline: you ask what has changed, upload only that, and register the result in one transaction.
1. Ask what is missing
Send the manifest, hashes only. This is metadata, so it costs one indexed query and moves no bytes.
curl -X POST https://api.engramdynamics.org/v1/corpora/c_7a1f.../documents/diff \
-H "Authorization: Bearer <your key>" \
-H "Content-Type: application/json" \
-d '{
"files": [
{"path": "handbook/ch1.md", "sha256": "9f86d0818...", "size": 4210},
{"path": "handbook/ch2.md", "sha256": "2c26b46b6...", "size": 9130}
]
}'
{
"new": ["handbook/ch2.md"],
"changed": [],
"unchanged": ["handbook/ch1.md"],
"missing": ["handbook/appendix-old.md"]
}
sha256 is of the raw bytes, which is the only hash you can compute: you have the
file, not our parser. The four buckets are disjoint. missing is what we hold and your
manifest did not mention, which is what drives a mirror-style delete.
2. Get somewhere to put the bytes
curl -X POST https://api.engramdynamics.org/v1/corpora/c_7a1f.../documents/upload-urls \
-H "Authorization: Bearer <your key>" \
-H "Content-Type: application/json" \
-d '{
"files": [
{"path": "handbook/ch2.md", "sha256": "2c26b46b6...", "size": 9130,
"content_type": "text/markdown"}
]
}'
{
"uploads": [
{
"path": "handbook/ch2.md",
"method": "PUT",
"url": "https://s3.amazonaws.com/...signature...",
"headers": {"Content-Type": "text/markdown"},
"expires_at": "2026-09-13T11:06:02Z"
}
]
}
The upload goes straight from your machine to storage, so a large file never passes through the
API. Send exactly the headers that came back: they were signed into the URL, and
different ones are rejected.
3. Put the bytes
curl -X PUT "https://s3.amazonaws.com/...signature..." \
-H "Content-Type: text/markdown" \
--data-binary @handbook/ch2.md
This step registers nothing. Storage holds the object, and no document row exists until you commit.
4. Commit the manifest
curl -X POST https://api.engramdynamics.org/v1/corpora/c_7a1f.../documents/commit \
-H "Authorization: Bearer <your key>" \
-H "Idempotency-Key: push-2026-09-13-run-7" \
-H "Content-Type: application/json" \
-d '{
"files": [
{"path": "handbook/ch1.md", "sha256": "9f86d0818...", "size": 4210},
{"path": "handbook/ch2.md", "sha256": "2c26b46b6...", "size": 9130}
],
"onboard": true,
"delete_missing": false
}'
{
"sync_run_id": "r_5d2c...",
"documents": [
{"path": "handbook/ch1.md", "id": "d_11a...", "lifecycle": "ready"},
{"path": "handbook/ch2.md", "id": "d_92f...", "lifecycle": "pending"}
],
"removed": []
}
Commit is the transaction. It checks the quotas on the whole manifest before a single row exists, registers the rows, removes what the manifest no longer lists when you ask for it, and queues one parse job per document. Nothing is parsed inside the call, so a fifty-thousand-file commit answers as fast as a two-file one.
Two flags decide how it behaves. onboard defaults to true and starts the build;
set it false to land a large initial load and build it later. delete_missing defaults to
false, because deleting documents is the one thing a sync must never do by accident. A file the diff
called unchanged is counted and skipped: no parse, no build, no cost.
Send an Idempotency-Key on the commit. A pipeline that reruns after a
timeout then replays the first answer instead of creating a second run. See
Conventions and errors.
When to use upsert instead
Upsert writes one document, content included, in a single call. Use it when the content is already in hand: an agent writing a note, a pipeline emitting markdown, a webhook handler mirroring a record. Using the four-step protocol for one small file is three round trips to save a copy that does not need saving.
curl -X POST https://api.engramdynamics.org/v1/corpora/c_7a1f.../documents/upsert \
-H "Authorization: Bearer <your key>" \
-H "Content-Type: application/json" \
-d '{
"path": "notes/2026-09-13-standup.md",
"text": "# Standup\n\nShipped the ingest docs.",
"onboard": true
}'
requests.post(
f"{BASE}/corpora/{base_id}/documents/upsert",
headers=HEADERS,
json={"path": "notes/2026-09-13-standup.md", "text": note_markdown},
timeout=30,
).raise_for_status()
Send exactly one of text or content_base64. Text is the common case and
skips base64 entirely; base64 is there for a PDF or a DOCX. Upsert is idempotent on
path, so the same path twice updates the same document rather than making a second one,
which is what a pipeline that reruns needs. It goes through the same queue as a bulk commit, so you
get the same lifecycle, the same run history and the same completion signal.
Use the protocol instead when you are syncing a folder (the diff is what makes an unchanged sync
free), when a file is larger than the upsert path should carry, or when you need
delete_missing to mirror deletions.
Watch the run
Every path into a base records a sync run, so uploads, pushes and source syncs share one history.
curl https://api.engramdynamics.org/v1/corpora/c_7a1f.../sync-runs/r_5d2c... \
-H "Authorization: Bearer <your key>"
{
"id": "r_5d2c...",
"corpus_id": "c_7a1f...",
"source": "upload",
"source_id": null,
"state": "succeeded",
"counters": {"added": 1, "updated": 0, "unchanged": 1, "removed": 0, "failed": 0},
"documents_total": 1,
"documents_done": 1,
"error": null,
"started_at": "2026-09-13T10:06:04Z",
"finished_at": "2026-09-13T10:06:39Z"
}
state is running, succeeded, failed,
limited or canceled. A run fails only when nothing it tried worked: one
unreadable file in a hundred is a counter, not a failed run. GET
/v1/corpora/{id}/sync-runs lists the history newest first.
Rather than polling, subscribe an endpoint and get told: Webhooks
fire sync_run.completed with these exact counters.
List and delete documents
curl "https://api.engramdynamics.org/v1/corpora/c_7a1f.../documents?limit=100" \
-H "Authorization: Bearer <your key>"
curl -X DELETE https://api.engramdynamics.org/v1/corpora/c_7a1f.../documents/d_92f... \
-H "Authorization: Bearer <your key>"
The listing returns files. Each document carries lifecycle
(pending, parsing, parsed, queued,
building, ready, failed or stale),
lifecycle_error when that needs words, content_hash and
cart_id so you can tell "same document, renamed" from "document changed", and
description once the post-onboarding pass has written one.
Deleting one document removes it everywhere: the raw file, the extracted text, the row, the retrieval index, and the built memory plus every warm copy of it when no other document in the workspace shares that cartridge. A build in flight does not block the delete.
Limits that are API constants
These are fixed ceilings in the API, quoted here with the setting each comes from. What your plan allows is separate and lives on Plans and allowances.
| Ceiling | Value | Setting |
|---|---|---|
| Files per manifest call (diff, upload-urls, commit) | 10,000 | ManifestReq.files max length |
| Per file on the presigned upload path and on upsert | 500 MB | MAX_API_UPLOAD_MB |
| Upload URL lifetime | 1 hour | UPLOAD_URL_TTL_S |
Per file on the multipart route POST /v1/corpora/{id}/documents |
25 MB | MAX_UPLOAD_MB |
| Per request on that same multipart route | 200 MB | MAX_REQUEST_MB |
| Documents in one base | 50,000 | MAX_CORPUS_DOCUMENTS |
The base ceiling is about keeping one base fast rather than about what you bought: routing quality and answer latency are properties of a single base, so split a very large set across several and query both.
Next
Watch the build and know when the base is answering: Onboarding and jobs. What we can read, and how a file becomes a document, is on What counts as a document.