REST API

Documents and ingest

Load documents the way that fits your pipeline: one call for a single document, a four-step protocol for a folder that has to stay in sync, or a source that pulls so the bytes never touch your machine.

There are three ways to get documents in, and the right one depends on how many files you have and where they live. All three end in the same place: a document row, extracted text, and a build queued.

Three ways in

You haveUse
One document, content already in handUpsert, a single call
A folder, a large file, or a sync that has to be exactThe ingest protocol, four steps
A bucket, a Drive folder or a SharePoint librarySources, so Engram pulls and you never move the bytes

The CLI implements the protocol below for you, so engram push ./docs is the shortest path if you are not building an integration.

The ingest protocol

Four calls, in order. The shape exists to keep the bytes off the API and the round trips off your pipeline: you ask what has changed, upload only that, and register the result in one transaction.

1. Ask what is missing

Send the manifest, hashes only. This is metadata, so it costs one indexed query and moves no bytes.

curl -X POST https://api.engramdynamics.org/v1/corpora/c_7a1f.../documents/diff \
  -H "Authorization: Bearer <your key>" \
  -H "Content-Type: application/json" \
  -d '{
        "files": [
          {"path": "handbook/ch1.md", "sha256": "9f86d0818...", "size": 4210},
          {"path": "handbook/ch2.md", "sha256": "2c26b46b6...", "size": 9130}
        ]
      }' 
{
  "new": ["handbook/ch2.md"],
  "changed": [],
  "unchanged": ["handbook/ch1.md"],
  "missing": ["handbook/appendix-old.md"]
}

sha256 is of the raw bytes, which is the only hash you can compute: you have the file, not our parser. The four buckets are disjoint. missing is what we hold and your manifest did not mention, which is what drives a mirror-style delete.

2. Get somewhere to put the bytes

curl -X POST https://api.engramdynamics.org/v1/corpora/c_7a1f.../documents/upload-urls \
  -H "Authorization: Bearer <your key>" \
  -H "Content-Type: application/json" \
  -d '{
        "files": [
          {"path": "handbook/ch2.md", "sha256": "2c26b46b6...", "size": 9130,
           "content_type": "text/markdown"}
        ]
      }' 
{
  "uploads": [
    {
      "path": "handbook/ch2.md",
      "method": "PUT",
      "url": "https://s3.amazonaws.com/...signature...",
      "headers": {"Content-Type": "text/markdown"},
      "expires_at": "2026-09-13T11:06:02Z"
    }
  ]
}

The upload goes straight from your machine to storage, so a large file never passes through the API. Send exactly the headers that came back: they were signed into the URL, and different ones are rejected.

3. Put the bytes

curl -X PUT "https://s3.amazonaws.com/...signature..." \
  -H "Content-Type: text/markdown" \
  --data-binary @handbook/ch2.md

This step registers nothing. Storage holds the object, and no document row exists until you commit.

4. Commit the manifest

curl -X POST https://api.engramdynamics.org/v1/corpora/c_7a1f.../documents/commit \
  -H "Authorization: Bearer <your key>" \
  -H "Idempotency-Key: push-2026-09-13-run-7" \
  -H "Content-Type: application/json" \
  -d '{
        "files": [
          {"path": "handbook/ch1.md", "sha256": "9f86d0818...", "size": 4210},
          {"path": "handbook/ch2.md", "sha256": "2c26b46b6...", "size": 9130}
        ],
        "onboard": true,
        "delete_missing": false
      }' 
{
  "sync_run_id": "r_5d2c...",
  "documents": [
    {"path": "handbook/ch1.md", "id": "d_11a...", "lifecycle": "ready"},
    {"path": "handbook/ch2.md", "id": "d_92f...", "lifecycle": "pending"}
  ],
  "removed": []
}

Commit is the transaction. It checks the quotas on the whole manifest before a single row exists, registers the rows, removes what the manifest no longer lists when you ask for it, and queues one parse job per document. Nothing is parsed inside the call, so a fifty-thousand-file commit answers as fast as a two-file one.

Two flags decide how it behaves. onboard defaults to true and starts the build; set it false to land a large initial load and build it later. delete_missing defaults to false, because deleting documents is the one thing a sync must never do by accident. A file the diff called unchanged is counted and skipped: no parse, no build, no cost.

Send an Idempotency-Key on the commit. A pipeline that reruns after a timeout then replays the first answer instead of creating a second run. See Conventions and errors.

When to use upsert instead

Upsert writes one document, content included, in a single call. Use it when the content is already in hand: an agent writing a note, a pipeline emitting markdown, a webhook handler mirroring a record. Using the four-step protocol for one small file is three round trips to save a copy that does not need saving.

curl -X POST https://api.engramdynamics.org/v1/corpora/c_7a1f.../documents/upsert \
  -H "Authorization: Bearer <your key>" \
  -H "Content-Type: application/json" \
  -d '{
        "path": "notes/2026-09-13-standup.md",
        "text": "# Standup\n\nShipped the ingest docs.",
        "onboard": true
      }' 
requests.post(
    f"{BASE}/corpora/{base_id}/documents/upsert",
    headers=HEADERS,
    json={"path": "notes/2026-09-13-standup.md", "text": note_markdown},
    timeout=30,
).raise_for_status()

Send exactly one of text or content_base64. Text is the common case and skips base64 entirely; base64 is there for a PDF or a DOCX. Upsert is idempotent on path, so the same path twice updates the same document rather than making a second one, which is what a pipeline that reruns needs. It goes through the same queue as a bulk commit, so you get the same lifecycle, the same run history and the same completion signal.

Use the protocol instead when you are syncing a folder (the diff is what makes an unchanged sync free), when a file is larger than the upsert path should carry, or when you need delete_missing to mirror deletions.

Watch the run

Every path into a base records a sync run, so uploads, pushes and source syncs share one history.

curl https://api.engramdynamics.org/v1/corpora/c_7a1f.../sync-runs/r_5d2c... \
  -H "Authorization: Bearer <your key>" 
{
  "id": "r_5d2c...",
  "corpus_id": "c_7a1f...",
  "source": "upload",
  "source_id": null,
  "state": "succeeded",
  "counters": {"added": 1, "updated": 0, "unchanged": 1, "removed": 0, "failed": 0},
  "documents_total": 1,
  "documents_done": 1,
  "error": null,
  "started_at": "2026-09-13T10:06:04Z",
  "finished_at": "2026-09-13T10:06:39Z"
}

state is running, succeeded, failed, limited or canceled. A run fails only when nothing it tried worked: one unreadable file in a hundred is a counter, not a failed run. GET /v1/corpora/{id}/sync-runs lists the history newest first.

Rather than polling, subscribe an endpoint and get told: Webhooks fire sync_run.completed with these exact counters.

List and delete documents

curl "https://api.engramdynamics.org/v1/corpora/c_7a1f.../documents?limit=100" \
  -H "Authorization: Bearer <your key>"

curl -X DELETE https://api.engramdynamics.org/v1/corpora/c_7a1f.../documents/d_92f... \
  -H "Authorization: Bearer <your key>" 

The listing returns files. Each document carries lifecycle (pending, parsing, parsed, queued, building, ready, failed or stale), lifecycle_error when that needs words, content_hash and cart_id so you can tell "same document, renamed" from "document changed", and description once the post-onboarding pass has written one.

Deleting one document removes it everywhere: the raw file, the extracted text, the row, the retrieval index, and the built memory plus every warm copy of it when no other document in the workspace shares that cartridge. A build in flight does not block the delete.

Limits that are API constants

These are fixed ceilings in the API, quoted here with the setting each comes from. What your plan allows is separate and lives on Plans and allowances.

CeilingValueSetting
Files per manifest call (diff, upload-urls, commit)10,000 ManifestReq.files max length
Per file on the presigned upload path and on upsert500 MB MAX_API_UPLOAD_MB
Upload URL lifetime1 hourUPLOAD_URL_TTL_S
Per file on the multipart route POST /v1/corpora/{id}/documents 25 MBMAX_UPLOAD_MB
Per request on that same multipart route200 MB MAX_REQUEST_MB
Documents in one base50,000MAX_CORPUS_DOCUMENTS

The base ceiling is about keeping one base fast rather than about what you bought: routing quality and answer latency are properties of a single base, so split a very large set across several and query both.

Next

Watch the build and know when the base is answering: Onboarding and jobs. What we can read, and how a file becomes a document, is on What counts as a document.