Sources
Sources
Point a document base at where your documents already live and Engram pulls them, keeps them current, and records every run. No bytes through your laptop, no keys to paste.
A source is a standing connection between a document base and somewhere your documents already live. You register it once and Engram pulls: nothing is uploaded from a laptop, no bytes go through your CI, and a folder that changes is picked up on its own.
The three kinds
There is a fourth way in that is not a source: ask for an hour of scoped S3 credentials and push into our bucket with any S3 client you already use.
curl -X POST https://api.engramdynamics.org/v1/corpora/c_7a1f.../upload-credentials \
-H "Authorization: Bearer <your key>"
{
"bucket": "engram-platform-documents",
"prefix": "corpora/c_7a1f.../docs/",
"region": "us-east-1",
"access_key_id": "ASIA...",
"secret_access_key": "...",
"session_token": "...",
"expires_at": "2026-09-13T11:14:02Z",
"sync_command": "aws s3 sync ./documents s3://engram-platform-documents/corpora/c_7a1f.../docs/ --region us-east-1",
"session_policy": "{ ... the exact policy these credentials carry ... }"
}
The credentials can write and list one document base's prefix and nothing else, so
aws s3 sync, DataSync and bucket replication all work without Engram sitting in the data
path. An hour is the ceiling rather than a preference: long enough for a large sync, short enough
that a leaked set stops working the same afternoon.
Where sources live in the app
Every document base's Documents tab opens with a Sources section, and it holds one card per place the documents came from, so you can see what each folder or bucket is contributing without opening anything. The readiness counts for the whole base sit below it, with Sync history at the bottom. Links to the old Sources tab open the Documents tab.
Every registered Google Drive folder, SharePoint library and S3 bucket gets a card, and one more card,
Website uploads, covers everything that did not come from a registered source: files
uploaded on the website, files sent with engram push, and one-off imports through the wizard's
Drive or SharePoint picker. Each card names the source, says what kind it is, counts the documents in the
base that came from it, and offers Download manifest.
| Card | Set it up | What the card carries |
|---|---|---|
| Google Drive or SharePoint | In the app, with Keep a folder in sync, or on the setup wizard's Review step after an import. The CLI works too. | The sync settings, plus Sync now and Remove |
| Amazon S3 | With the CLI or the REST API, which is where the IAM template you apply comes from | Read-only, with the line "Managed with the CLI, where you can check access, sync it now or remove it." |
| Website uploads | Nothing to register. Anything that was not pulled from a source lands here | Add documents, which opens the wizard's Documents step |
Removing a source stops the syncing, not the documents. What it already delivered stays in the base and moves under Website uploads, so disconnecting a folder never costs you answers.
Adding a bucket starts in the same section, which carries the line "Pulling from another cloud source?
Set it up with the CLI: engram sources add s3".
Download a manifest
Download manifest on a card saves that source's documents as a CSV, one row per document. It is where file-by-file detail lives: open it in a spreadsheet, filter to the failures, and you have the list to fix. The tab itself carries the counts, so the page answers "how is this base doing" and the manifest answers "what happened to this file".
| Column | What it holds |
|---|---|
filename | The document's name in the base. |
size_bytes | How big the file is. |
status | The lifecycle state: ready, building,
failed, stale and the rest, set out on
What counts as a document. |
reason | Why a document failed. Blank when there is nothing wrong. |
last_change | When the state last moved. |
content_hash | Identifies the extracted text, so a renamed file reads as the same document. |
source | The kind of source the document came from, or upload
for anything loaded on the website, pushed from the CLI or imported once. |
Register, validate, sync, schedule
Every source follows the same four steps, and the API mirrors them one to one.
| Step | Call | What it settles |
|---|---|---|
| Register | POST /v1/corpora/{id}/sources | Which bucket or folder, which prefix, which mode, how often. |
| Validate | POST /v1/corpora/{id}/sources/{source_id}/validate |
Can we actually read it, and how many of the objects there are documents we can ingest. |
| Sync | POST /v1/corpora/{id}/sources/{source_id}/sync | Pull now and return a run to watch. |
| Schedule | schedule_minutes on the source | Re-walk on a timer, unattended. |
| Update | PUT /v1/corpora/{id}/sources/{source_id} | Change the mode, the schedule and the credentials in place, without re-registering. |
Registration deliberately does not require a working grant. It records the source, tells you plainly what is not right yet, and hands back the setup you still have to apply. Validate, or simply the first sync, flips it to ready.
Additive or mirror
mode decides what a sync is allowed to do:
- additive (the default) adds new files and updates changed ones. Deleting something in your bucket leaves the document in the base.
- mirror also removes documents this source delivered whose source object is gone, so the base tracks the folder exactly.
Additive is the default because deleting documents is the one thing a sync must never do by
accident. Prefixes are stripped, so a base fed from s3://acme/kb/ holds
handbook/ch1.md rather than kb/handbook/ch1.md: your folder structure, not
your bucket layout.
In either mode, a sync stays inside its own lane. A document belongs to the source that delivered it, identified by the document base, the source and that stripped path, and a sync only matches, updates, re-attributes or removes documents of its own source. Three things follow, and they are what let one base pull from several places safely:
- A mirror removes only what its own source delivered. It never removes a website upload, an
engram push, or another source's documents. - Uploads and pushes belong to no source, so only another upload or push touches them.
- A filename belongs to whoever delivered it first. If another source, or an upload, already holds
a
reports/keep.txt, the sync leaves that document exactly as it is and counts the object aspath_taken, so the clash is reported rather than settled silently behind you. Give one of the two its own name, folder or base and both come through.
Watch a sync
A sync writes the same run record every other ingestion path writes, so one history covers uploads, pushes and every source.
curl "https://api.engramdynamics.org/v1/corpora/c_7a1f.../sync-runs?limit=10" \
-H "Authorization: Bearer <your key>"
{
"items": [
{
"id": "r_5d2c...",
"source": "s3",
"source_id": "s_3e8b...",
"state": "succeeded",
"counters": {"added": 12, "updated": 3, "unchanged": 480, "skipped": 37, "removed": 0, "failed": 1},
"skipped_reasons": {"unsupported_type": 30, "empty": 3, "too_large": 2, "archived": 2},
"documents_total": 533,
"documents_done": 533,
"started_at": "2026-09-13T10:06:04Z",
"finished_at": "2026-09-13T10:09:51Z"
}
],
"next_cursor": null
}
skipped is the objects the run looked at and did not ingest, and
skipped_reasons sits beside the counters naming which reasons fired and how many objects
each one took: only the reasons that fired appear, and it is {} when nothing was skipped.
documents_total is every object the run considered rather than only the ones it
downloaded, so a run that finds nothing changed still reports what it walked. The five reasons are on
What counts as a document.
Better than polling: subscribe an endpoint and sync_run.completed arrives with these
counters and reasons the moment a run lands. See Webhooks.
Two more useful behaviours. A run counts as failed only when nothing it tried worked, so one unreadable file in five hundred is a counter rather than a failure. And a sync of a source that is already syncing loses immediately with a 409 saying so, rather than two walks of one folder fetching the same objects twice.
Which plan turns this on
Registering a source and triggering a sync by hand come with the Pro plan and every tier above. What each tier includes is on Plans and allowances.
Scheduled syncs of a source that already exists keep running whatever plan the workspace moves to. Stopping a customer's documents from tracking their bucket, silently, because a plan changed is a worse outcome than an unenforced promise.
Fixed limits
These are API constants rather than plan allowances, quoted with the setting each comes from.
| Limit | Value | Setting |
|---|---|---|
| Sources on one document base | 10 | MAX_SOURCES_PER_CORPUS |
| Schedule interval | 15 minutes to 7 days | schedule_minutes, minimum 15, maximum 10080 |
| One object pulled from a source | 500 MB | MAX_SOURCE_OBJECT_MB |
| Files considered in one Drive or SharePoint import run | 2,000 | the connector walk cap |
A tighter schedule than 15 minutes costs listing calls for nothing. If you want new files in minutes, wire the bucket notification described on Amazon S3 buckets instead of polling harder.
Next
Set one up: Amazon S3 buckets from the CLI, Google Drive or SharePoint. Turning a Drive or SharePoint folder into a standing source from the app is on Keeping a folder in sync. What we can read out of a file is on What counts as a document.