Sources

Sources

Point a document base at where your documents already live and Engram pulls them, keeps them current, and records every run. No bytes through your laptop, no keys to paste.

A source is a standing connection between a document base and somewhere your documents already live. You register it once and Engram pulls: nothing is uploaded from a laptop, no bytes go through your CI, and a folder that changes is picked up on its own.

The three kinds

There is a fourth way in that is not a source: ask for an hour of scoped S3 credentials and push into our bucket with any S3 client you already use.

curl -X POST https://api.engramdynamics.org/v1/corpora/c_7a1f.../upload-credentials \
  -H "Authorization: Bearer <your key>" 
{
  "bucket": "engram-platform-documents",
  "prefix": "corpora/c_7a1f.../docs/",
  "region": "us-east-1",
  "access_key_id": "ASIA...",
  "secret_access_key": "...",
  "session_token": "...",
  "expires_at": "2026-09-13T11:14:02Z",
  "sync_command": "aws s3 sync ./documents s3://engram-platform-documents/corpora/c_7a1f.../docs/ --region us-east-1",
  "session_policy": "{ ... the exact policy these credentials carry ... }"
}

The credentials can write and list one document base's prefix and nothing else, so aws s3 sync, DataSync and bucket replication all work without Engram sitting in the data path. An hour is the ceiling rather than a preference: long enough for a large sync, short enough that a leaked set stops working the same afternoon.

Where sources live in the app

Every document base's Documents tab opens with a Sources section, and it holds one card per place the documents came from, so you can see what each folder or bucket is contributing without opening anything. The readiness counts for the whole base sit below it, with Sync history at the bottom. Links to the old Sources tab open the Documents tab.

Every registered Google Drive folder, SharePoint library and S3 bucket gets a card, and one more card, Website uploads, covers everything that did not come from a registered source: files uploaded on the website, files sent with engram push, and one-off imports through the wizard's Drive or SharePoint picker. Each card names the source, says what kind it is, counts the documents in the base that came from it, and offers Download manifest.

CardSet it upWhat the card carries
Google Drive or SharePointIn the app, with Keep a folder in sync, or on the setup wizard's Review step after an import. The CLI works too.The sync settings, plus Sync now and Remove
Amazon S3With the CLI or the REST API, which is where the IAM template you apply comes fromRead-only, with the line "Managed with the CLI, where you can check access, sync it now or remove it."
Website uploadsNothing to register. Anything that was not pulled from a source lands hereAdd documents, which opens the wizard's Documents step

Removing a source stops the syncing, not the documents. What it already delivered stays in the base and moves under Website uploads, so disconnecting a folder never costs you answers.

Adding a bucket starts in the same section, which carries the line "Pulling from another cloud source? Set it up with the CLI: engram sources add s3".

Download a manifest

Download manifest on a card saves that source's documents as a CSV, one row per document. It is where file-by-file detail lives: open it in a spreadsheet, filter to the failures, and you have the list to fix. The tab itself carries the counts, so the page answers "how is this base doing" and the manifest answers "what happened to this file".

ColumnWhat it holds
filenameThe document's name in the base.
size_bytesHow big the file is.
statusThe lifecycle state: ready, building, failed, stale and the rest, set out on What counts as a document.
reasonWhy a document failed. Blank when there is nothing wrong.
last_changeWhen the state last moved.
content_hashIdentifies the extracted text, so a renamed file reads as the same document.
sourceThe kind of source the document came from, or upload for anything loaded on the website, pushed from the CLI or imported once.

Register, validate, sync, schedule

Every source follows the same four steps, and the API mirrors them one to one.

StepCallWhat it settles
RegisterPOST /v1/corpora/{id}/sourcesWhich bucket or folder, which prefix, which mode, how often.
ValidatePOST /v1/corpora/{id}/sources/{source_id}/validate Can we actually read it, and how many of the objects there are documents we can ingest.
SyncPOST /v1/corpora/{id}/sources/{source_id}/syncPull now and return a run to watch.
Scheduleschedule_minutes on the sourceRe-walk on a timer, unattended.
UpdatePUT /v1/corpora/{id}/sources/{source_id}Change the mode, the schedule and the credentials in place, without re-registering.

Registration deliberately does not require a working grant. It records the source, tells you plainly what is not right yet, and hands back the setup you still have to apply. Validate, or simply the first sync, flips it to ready.

Additive or mirror

mode decides what a sync is allowed to do:

Additive is the default because deleting documents is the one thing a sync must never do by accident. Prefixes are stripped, so a base fed from s3://acme/kb/ holds handbook/ch1.md rather than kb/handbook/ch1.md: your folder structure, not your bucket layout.

In either mode, a sync stays inside its own lane. A document belongs to the source that delivered it, identified by the document base, the source and that stripped path, and a sync only matches, updates, re-attributes or removes documents of its own source. Three things follow, and they are what let one base pull from several places safely:

Watch a sync

A sync writes the same run record every other ingestion path writes, so one history covers uploads, pushes and every source.

curl "https://api.engramdynamics.org/v1/corpora/c_7a1f.../sync-runs?limit=10" \
  -H "Authorization: Bearer <your key>" 
{
  "items": [
    {
      "id": "r_5d2c...",
      "source": "s3",
      "source_id": "s_3e8b...",
      "state": "succeeded",
      "counters": {"added": 12, "updated": 3, "unchanged": 480, "skipped": 37, "removed": 0, "failed": 1},
      "skipped_reasons": {"unsupported_type": 30, "empty": 3, "too_large": 2, "archived": 2},
      "documents_total": 533,
      "documents_done": 533,
      "started_at": "2026-09-13T10:06:04Z",
      "finished_at": "2026-09-13T10:09:51Z"
    }
  ],
  "next_cursor": null
}

skipped is the objects the run looked at and did not ingest, and skipped_reasons sits beside the counters naming which reasons fired and how many objects each one took: only the reasons that fired appear, and it is {} when nothing was skipped. documents_total is every object the run considered rather than only the ones it downloaded, so a run that finds nothing changed still reports what it walked. The five reasons are on What counts as a document.

Better than polling: subscribe an endpoint and sync_run.completed arrives with these counters and reasons the moment a run lands. See Webhooks.

Two more useful behaviours. A run counts as failed only when nothing it tried worked, so one unreadable file in five hundred is a counter rather than a failure. And a sync of a source that is already syncing loses immediately with a 409 saying so, rather than two walks of one folder fetching the same objects twice.

Which plan turns this on

Registering a source and triggering a sync by hand come with the Pro plan and every tier above. What each tier includes is on Plans and allowances.

Scheduled syncs of a source that already exists keep running whatever plan the workspace moves to. Stopping a customer's documents from tracking their bucket, silently, because a plan changed is a worse outcome than an unenforced promise.

Fixed limits

These are API constants rather than plan allowances, quoted with the setting each comes from.

LimitValueSetting
Sources on one document base10MAX_SOURCES_PER_CORPUS
Schedule interval15 minutes to 7 days schedule_minutes, minimum 15, maximum 10080
One object pulled from a source500 MB MAX_SOURCE_OBJECT_MB
Files considered in one Drive or SharePoint import run2,000 the connector walk cap

A tighter schedule than 15 minutes costs listing calls for nothing. If you want new files in minutes, wire the bucket notification described on Amazon S3 buckets instead of polling harder.

Next

Set one up: Amazon S3 buckets from the CLI, Google Drive or SharePoint. Turning a Drive or SharePoint folder into a standing source from the app is on Keeping a folder in sync. What we can read out of a file is on What counts as a document.