Sources
Amazon S3 buckets
Feed a document base straight from your own bucket, set up from the CLI in a few commands. No access key changes hands, the grant is read-only, and you revoke it by deleting a role in your own account.
The S3 source is the one to use at scale. Engram reads one prefix of one bucket through a role you create in your own account, on the schedule you choose, so documents never pass through a laptop and nobody has to remember to upload them. Setup ends in your AWS account, so it lives where you already work with AWS: the Engram CLI, or the REST API if you are scripting it.
In the app, a registered bucket gets its own card in the Sources section at the top of the document base's Documents tab, counting the documents it has delivered, with the line "Managed with the CLI, where you can check access, sync it now or remove it." A bucket registered before its role works carries a Waiting for access badge until it can be read. Download manifest on the card saves those documents as a CSV, one row per file: see Download a manifest.
The sequence
This flow needs CLI 0.2.1 or later. Run engram --version to check, and
pip install -U engram-dynamics to upgrade.
- Print the template for your bucket. The CLI generates the external id and shows it with the role name and the exact command to run next.
- Apply it in the AWS account that owns the bucket, and copy the role ARN it outputs.
- Register the bucket with that external id and role ARN. Engram checks access as it saves.
- Sync now, and let the schedule keep the base current from there.
engram sources setup --bucket acme-knowledge-base --prefix handbooks/ --corpus Sales > engram_role.yaml
aws cloudformation deploy --template-file engram_role.yaml \
--stack-name engram-source-reader --capabilities CAPABILITY_NAMED_IAM
engram sources add s3 --corpus Sales --bucket acme-knowledge-base --prefix handbooks/ \
--external-id <external id> --role-arn <role ARN> --every 1h
engram sync --corpus Sales --wait
Building the role first means the ARN you register already exists, so the source is usually ready to sync the moment it is saved. Registering a source comes with the Pro plan and every tier above, as set out on Sources.
1. Print the template
engram sources setup --bucket acme-knowledge-base --prefix handbooks/ --corpus Sales > engram_role.yaml
engram sources setup --bucket acme-knowledge-base --prefix handbooks/ --corpus Sales --format terraform > engram_role.tf
With --bucket, no source has to exist yet. The template is the only thing written to standard
output, so the file is ready to apply. Your terminal shows the rest: the external id, the name of the role
the template creates, and the exact engram sources add s3 command to run once the role is in
place. Keep the external id, because the role trusts that exact value.
| Option | What it sets |
|---|---|
--bucket/-b | The bucket to build the role for. No source needed. |
--prefix | The folder to grant. Leave it off for the whole bucket. |
--format/-f | cloudformation (the default) or
terraform. Both grant exactly the same access. |
--corpus/-c | The document base, filled into the command it prints. |
--kms-key-arn | The customer-managed KMS key the bucket is encrypted with, if it has one. |
--inventory-bucket and --inventory-prefix | Where the daily S3 Inventory report lands. Both or neither. See S3 Inventory. |
--external-id | Build the role around an id of your own. Leave it off and one is generated for you. |
--json | Both templates plus external_id and
role_name, for a script that wants the parts separately. |
For a bucket that is already registered, engram sources setup --corpus Sales --source
<source id> prints its template again, with the same external id, role name and
ready-to-paste registration command the --bucket form prints. Repairing a source reads
exactly like setting one up.
2. Apply it in your AWS account
Apply the template in the account that owns the bucket, then read the role ARN from its output:
RoleArn on the CloudFormation stack, engram_role_arn in Terraform.
# CloudFormation
aws cloudformation deploy --template-file engram_role.yaml \
--stack-name engram-source-reader --capabilities CAPABILITY_NAMED_IAM
aws cloudformation describe-stacks --stack-name engram-source-reader \
--query "Stacks[0].Outputs[?OutputKey=='RoleArn'].OutputValue" --output text
# Terraform, from a folder with your AWS provider configured
terraform apply
terraform output -raw engram_role_arn
3. Register the bucket
engram sources add s3 --corpus Sales \
--bucket acme-knowledge-base \
--prefix handbooks/ \
--external-id engram-8f3c1d7b4a2e9051 \
--role-arn arn:aws:iam::111122223333:role/EngramSourceReader-1d7b4a2e9051 \
--every 1h
Engram checks access as it saves. When the role works, the source is ready and the receipt names the sync command to run next. When the role is not applied yet, or IAM has not caught up with it, the source is saved anyway and waits for access: the receipt says what to do next, and the app shows Waiting for access until validate confirms it.
| Option | What it sets |
|---|---|
--bucket/-b | The bucket name, or an s3:// URL.
Required. |
--role-arn | The role you applied, which Engram assumes. Required. |
--external-id | The external id the role trusts, as printed by
sources setup --bucket. Leave it off and Engram generates one. |
--prefix | The folder to read. Leave it off to read the whole bucket. |
--region | Where the bucket lives. Engram confirms the real region through the role and stores that, so the value here is a hint rather than the last word. See Region and egress. |
--kms-key-arn | The customer-managed KMS key the bucket is encrypted with, if it has one. |
--inventory-bucket and --inventory-prefix | Where the daily S3 Inventory report lands. Both or neither. |
--allow-cross-region | Accept a bucket outside Engram's region. See Region and egress. |
--mode | additive (the default) adds and updates.
mirror also removes documents this source delivered whose object is gone. |
--every | How often to check for changes, such as 15m,
1h or 1d. Leave it off to sync only when you ask. |
--json | The saved source as JSON, for scripts. |
A few mistakes are refused on the spot, because no amount of waiting fixes them: a malformed bucket name
or role ARN, the same bucket and prefix twice on one base, an external id that is too short or already in
use, or a bucket outside Engram's region without --allow-cross-region.
Prefer to register first? Leave out --external-id, then print that source's role with
engram sources setup --corpus Sales --source <source id>, apply it, and validate.
What the template creates
- A role that trusts Engram and nothing else. Only Engram's AWS account can assume it, and only while presenting the external id for this one source.
- A read-only policy on the prefix. The role can list the folder you shared and read the objects in it. There is no write action anywhere.
- KMS decrypt, only for a customer-managed key. Pass
--kms-key-arnand the policy gains decrypt on that one key. Otherwise it is left out. - Optional: a new-file notification, so new files arrive in minutes rather than on the next scheduled check.
- Optional: S3 Inventory, for a prefix holding millions of keys.
The trust policy
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": {"AWS": "arn:aws:iam::<the Engram AWS account id>:root"},
"Action": "sts:AssumeRole",
"Condition": {
"StringEquals": {"sts:ExternalId": "engram-8f3c1d7b4a2e9051"}
}
}
]
}
The condition is the whole security argument. Without it, any AWS customer who learned the role ARN could ask Engram's account to assume it. That is the confused deputy problem, and the external id is AWS's named answer to it.
The read-only policy
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "ListOnlyThisPrefix",
"Effect": "Allow",
"Action": "s3:ListBucket",
"Resource": "arn:aws:s3:::acme-knowledge-base",
"Condition": {"StringLike": {"s3:prefix": "handbooks/*"}}
},
{
"Sid": "ReadObjectsUnderPrefix",
"Effect": "Allow",
"Action": "s3:GetObject",
"Resource": "arn:aws:s3:::acme-knowledge-base/handbooks/*"
}
]
}
The prefix condition on ListBucket matters more than it looks: without it, "read this one
folder" quietly becomes "see every key name you own".
Encrypted buckets
For a bucket encrypted with a customer-managed KMS key, the policy adds one statement scoped to that key alone. If your key policy is restrictive, grant the role the same use on the key itself.
{
"Sid": "DecryptBucketObjects",
"Effect": "Allow",
"Action": ["kms:Decrypt", "kms:DescribeKey"],
"Resource": "arn:aws:kms:us-east-1:111122223333:key/2b7c..."
}
New files in minutes: the bucket notification
The template ends with an optional bucket notification that tells Engram the moment an object lands under your prefix, so new and changed files arrive within minutes. Every message is checked back against the source you registered before anything is read. The two work together: notifications for new files, the schedule for catching up on everything else.
- Terraform adds it as an
aws_s3_bucket_notificationresource. - CloudFormation can only set notifications on a bucket it creates, so the template
gives you an
aws s3api put-bucket-notification-configurationcommand in a comment instead.
Both replace the bucket's whole notification setup. In
Terraform, aws_s3_bucket_notification owns every notification on the bucket, so applying it
removes any the bucket already sends to a Lambda function, another queue or EventBridge. The AWS CLI
command works the same way. If the bucket already sends events elsewhere, merge those into the same
configuration first, or add Engram's queue to wherever that configuration is already managed.
Millions of keys: S3 Inventory
For a very large prefix, Engram can read the daily S3 Inventory report instead of listing the bucket on
every sync. Most document bases do not need it; the table below shows where it starts to pay. Pass
--inventory-bucket and --inventory-prefix to both sources setup and
sources add s3, and the template gains three pieces: read access for the role on the report
folder, the inventory configuration itself (an aws_s3_bucket_inventory resource in Terraform, a
put-bucket-inventory-configuration command in the CloudFormation comments), and a bucket policy
statement that lets S3 write the report into the destination bucket. Until the first report arrives, Engram
lists the bucket as usual and says so in inventory_notice.
Adding an inventory to a bucket that is already registered takes two steps: set the destination on the
source through the API, then render the updated template with sources setup --bucket, giving it the
source's own bucket, prefix and --external-id plus the two inventory flags.
curl -X PUT https://api.engramdynamics.org/v1/corpora/c_7a1f.../sources/s_3e8b... \
-H "Authorization: Bearer <your key>" \
-H "Content-Type: application/json" \
-d '{"inventory_bucket": "acme-inventory", "inventory_prefix": "engram/"}'
The figures
| What | Value |
|---|---|
| Listing a bucket | One call per 1,000 keys |
| Where S3 Inventory starts to pay | Past roughly 1 million keys |
| How fresh the inventory report is | Up to 1 day old |
| Time to the first inventory report | Up to 48 hours |
| An external id you supply | At least 16 characters, unique |
The schedule range and the largest object Engram pulls are on Sources.
Waiting for access? Validate
engram sources validate --corpus Sales --source s_3e8b... --json
{
"status": "ready",
"message": "Engram can read s3://acme-knowledge-base/handbooks/. 8 of the first objects listed are documents it can ingest.",
"objects_listed": 42,
"ingestible_sample": ["handbooks/ch1.md", "handbooks/ch2.md"],
"skipped_unsupported_type": 30,
"skipped_too_large": 2,
"skipped_archived": 2,
"inventory_configured": false
}
Validate assumes the role and lists one page, so you know in seconds whether access works.
--source is optional while the base has one source. The skip counts answer "my bucket has 900
files and Engram found 12" without anyone reading a log: unsupported types, objects over the size ceiling,
and objects in archived storage classes that need a restore first. If access still fails, the command says
why and the source keeps that reason, so fix the trust policy and run it again.
4. Sync and keep it current
engram sync --corpus Sales --wait
engram sync runs --corpus Sales --limit 5
A sync works out what changed from a listing and does the downloads on Engram's side, so a large bucket
starts as fast as an empty one. --wait follows the run to the end and exits non-zero if it
fails; without it the command returns as soon as the run is open. A source still waiting for access is
checked again first, so applying the role and running sync is enough.
From there, --every keeps the base current unattended. Every run shows in
Sync history on the Documents tab, and sync_run.completed
tells your own systems the moment one lands.
Every run says what it did: added, updated, unchanged, skipped, removed and failed, with the skipped
objects broken down by reason in skipped_reasons, so a bucket full of images or archived
objects explains itself without anyone reading a log. The counts are the same in engram sync
runs, in engram sync --wait, on the run over the API and in the run list on the
Documents tab, and the run's documents_total counts every object it
considered rather than only the ones it downloaded. The five reasons are on
What gets skipped.
Asking for a sync of a source that is already syncing is refused on the spot with "A sync of this source is already running". Two walks of one bucket would fetch the same objects twice and the second one adds nothing.
Each bucket keeps its own documents
Register as many buckets on one base as the source limit allows and none of them can tread on another. A document belongs to the source that delivered it, identified by the document base, the source and the file's path inside the prefix, and a sync only ever matches, updates, re-attributes or removes documents of its own source.
- A mirror only removes what its own source delivered. It never removes a website
upload, an
engram push, or another bucket's documents. - Uploads and pushes belong to no source. Only another upload or push touches them, whatever the buckets on the base are doing.
- A filename belongs to whoever delivered it first. Point one source at
s3://acme/finance/and another ats3://acme/legal/and areports/keep.txtunder each is a clash. The document that is already there keeps the name, the second run leaves it alone and counts the object aspath_taken, and you read that on the run rather than finding out from an answer. Give one of the two its own name, folder or base and both come through.
The same rule covers Drive and SharePoint folders: see Additive or mirror.
Change a registered bucket
engram sources update --corpus Sales --source s_3e8b... --every 6h
engram sources update --corpus Sales --source s_3e8b... --mode mirror
engram sources update --corpus Sales --source s_3e8b... \
--role-arn arn:aws:iam::111122223333:role/EngramSourceReader-9c04b1e27d3a \
--external-id engram-8f3c1d7b4a2e9051
A registered bucket is editable in place. The schedule, the mode, the role ARN and the external id
all change with one command, nothing already ingested is disturbed, and --source is
optional while the base has one source. Change the role ARN or the external id and Engram re-checks
access there and then and tells you the verdict, so a wrong ARN is a one-line repair rather than a
delete and a re-registration.
The bucket and the prefix stay as registered. Pointing at a different folder is a different source, so register that one and remove this one.
Over the API it is PUT /v1/corpora/{id}/sources/{source_id}, which takes
mode, schedule_minutes, role_arn and external_id
alongside the inventory fields.
Remove a bucket
engram sources list --corpus Sales
engram sources remove --corpus Sales --source s_3e8b...
Engram stops pulling, and the documents the bucket already delivered stay in the base, moved under the Website uploads card. To end access for good, delete the role in your own account as well.
Over the REST API
Every command above is one API call, which is the route for a pipeline or your own tooling. The same template-first order works: render the templates from values you choose, apply them, then register with the role ARN and the external id.
curl -G https://api.engramdynamics.org/v1/sources/s3/setup \
-H "Authorization: Bearer <your key>" \
--data-urlencode "external_id=engram-8f3c1d7b4a2e9051" \
--data-urlencode "bucket=acme-knowledge-base" \
--data-urlencode "prefix=handbooks/" | python -c "import json,sys; print(json.load(sys.stdin)['terraform'])"
The setup call is a pure renderer over the values you give it. Add kms_key_arn, or
inventory_bucket and inventory_prefix, to render those pieces too. Then
register:
curl -X POST https://api.engramdynamics.org/v1/corpora/c_7a1f.../sources \
-H "Authorization: Bearer <your key>" \
-H "Content-Type: application/json" \
-d '{
"kind": "s3",
"bucket": "acme-knowledge-base",
"prefix": "handbooks/",
"region": "us-east-1",
"external_id": "engram-8f3c1d7b4a2e9051",
"role_arn": "arn:aws:iam::111122223333:role/EngramSourceReader-1d7b4a2e9051",
"mode": "additive",
"schedule_minutes": 60
}'
Registering answers 201 whatever happens next, and status says where you are:
ready when the role already works, pending_validation when it does not yet, with the
reason in last_error and the templates in setup:
{
"id": "s_3e8b...",
"corpus_id": "c_7a1f...",
"kind": "s3",
"bucket": "acme-knowledge-base",
"prefix": "handbooks/",
"region": "us-east-1",
"role_arn": "arn:aws:iam::111122223333:role/EngramSourceReader-1d7b4a2e9051",
"external_id": "engram-8f3c1d7b4a2e9051",
"mode": "additive",
"schedule_minutes": 60,
"status": "pending_validation",
"last_error": "Engram could not assume that role yet.",
"setup": {
"external_id": "engram-8f3c1d7b4a2e9051",
"platform_account_id": "<the Engram AWS account id>",
"platform_region": "us-east-1",
"instructions": ["Apply either template below ...", "Copy the role ARN it outputs ...", "Run validate ..."],
"cloudformation": "AWSTemplateFormatVersion: '2010-09-09' ...",
"terraform": "data \"aws_iam_policy_document\" \"engram_trust\" ..."
}
}
Registering needs the ingest scope and a workspace admin. Leave out external_id
and Engram generates one. Either way it is readable back from the source on purpose: it is a secret shared
with you rather than kept from you, and a setup you cannot re-read is a setup you cannot repair.
PUT on the source changes the mode, the schedule, the role and the external id later.
Validate, sync and remove are one call each:
curl -X POST https://api.engramdynamics.org/v1/corpora/c_7a1f.../sources/s_3e8b.../validate \
-H "Authorization: Bearer <your key>"
curl -X POST https://api.engramdynamics.org/v1/corpora/c_7a1f.../sources/s_3e8b.../sync \
-H "Authorization: Bearer <your key>"
curl -X DELETE https://api.engramdynamics.org/v1/corpora/c_7a1f.../sources/s_3e8b... \
-H "Authorization: Bearer <your key>"
{
"id": "r_5d2c...",
"corpus_id": "c_7a1f...",
"source": "s3",
"source_id": "s_3e8b...",
"state": "running",
"counters": {"added": 0, "updated": 0, "unchanged": 0, "skipped": 0, "removed": 0, "failed": 0},
"skipped_reasons": {},
"documents_total": 15,
"documents_done": 0,
"started_at": "2026-09-13T10:06:04Z",
"finished_at": null
}
The sync call returns the run above straight away. Watch it with GET
/v1/corpora/{id}/sync-runs/{run_id}. If a source waiting for access still cannot read the bucket, the
409 carries the provider's own reason rather than a generic "not ready".
Region and egress
curl https://api.engramdynamics.org/v1/platform/info
{
"region": "us-east-1",
"region_guidance": "Engram runs in us-east-1. Keep your S3 bucket in us-east-1 and the data transfer to Engram is free and fast ...",
"max_source_object_mb": 500,
"supported_extensions": [".csv", ".doc", ".docx", ".htm", ".html", ".md", ".pdf", ".tsv", ".txt", ".xls", ".xlsx"]
}
This route is public, because "which region should my bucket be in" is a question people have before they
have an account. Same region means the transfer is free and fast. A bucket elsewhere works, and you opt into
it with --allow-cross-region (allow_cross_region over the API), but AWS bills you
cross-region transfer on every object Engram reads, and the first sync is when that bill is largest.
The region is resolved, not taken on trust. Engram asks S3 where the bucket really is, through the
role it assumes, and stores that answer on the source. A bucket outside Engram's region is refused
unless the registration opts in, so leaving --region off cannot quietly sign you up for a
transfer bill, and a source that says us-east-1 really is in us-east-1.