Sources
What counts as a document
A document is one file you loaded. Here is what we can read, what we skip and say so, and what happens to a file between arriving and answering questions.
Whatever route documents take in, they all end up the same way: the file is stored, its text is extracted, and that text is what gets read and answered from. This page is what that means in practice.
What counts as a document
A document is one file you loaded. A 200-page PDF is one document, a two-line note is one document, and both count as one against your allowances. The internal pieces a long file is served in are never counted, so a long document never quietly costs you twenty.
What each plan allows is on Plans and allowances.
File types we read
The accepted list is served live, so it can never disagree with what the API actually takes:
curl https://api.engramdynamics.org/v1/platform/info
| Extensions | How it is read |
|---|---|
.txt, .md | Directly. |
.pdf | The text layer. A scanned PDF with no text layer falls back to OCR at upload rather than onboarding as an empty document. |
.docx | Directly. .doc in the older RTF flavour is read
too; a genuine legacy binary .doc may need converting to .docx first,
and the error says so. |
.html, .htm | Text content, with the markup stripped. |
.xlsx, .xls, .csv, .tsv |
Cell content, sheet by sheet. A very large sheet is truncated rather than failed. |
An unsupported type is refused where you send it, before anything is stored, and the message lists what is accepted. That is deliberate: we do not want to be holding bytes we cannot turn into words.
Text only
Engram answers from the words in a document. Images, charts and layout are not read, so a slide that carries its meaning in a diagram contributes very little, and a scanned page contributes whatever the OCR could make out.
When extraction fails, the document is marked failed with a short reason in
lifecycle_error: an encrypted PDF, a corrupt file, an image-only page that OCR could not
read. A failed document is never onboarded as garbage, and never silently included in an answer. Three
places tell you: documents_failed on the base, the reason column of the
manifest you download from a source card, and the
document.failed webhook with the path and the reason.
How big a document is
Memory works in units of about 4,000 tokens, which is roughly six pages of prose. That is the
setting CAG_MAX_DOC_TOK, currently 4,096, and it is a property of the serving hardware
rather than of your plan.
A file longer than that is still read in full. Its whole text is indexed, and when you ask a question the passages that answer it are brought in alongside the resident memory, so page 150 of a long PDF is reachable. What the unit size means in practice is that short, well-titled documents retrieve more precisely than very long ones: if you control how the documents are produced, a folder of chapters beats one enormous export.
Size ceilings are separate from this and are listed on Documents and ingest: 500 MB per file on the upload and upsert paths, 500 MB per object pulled from a source, 50,000 documents in one base.
What gets skipped, and how you find out
A source sync reports its skips rather than hiding them. Every run carries a skipped
counter beside added, updated, unchanged, removed and failed, so "the bucket holds 900 files and Engram
found 12" is answered by the run itself. The counts read the same in engram sync runs, in
engram sync --wait, on the run over the API and in the run list on the
Documents tab.
Beside the counters, skipped_reasons says which reasons fired and how many objects each
one took. Only the reasons that actually fired appear, and it is {} when nothing was
skipped, so the object reads as a short list of what to fix:
| Reason | What it covers |
|---|---|
unsupported_type | An extension not on the list above, such as a
.jpg or a .zip. |
empty | A zero-byte object. An empty document is only ever a bad answer waiting to happen, so it is skipped rather than ingested. |
too_large | An object over the per-object ceiling. |
archived | An object in an archival storage class, which cannot be read without a restore in your own account. |
path_taken | Another source in this document base, or an upload, already holds that filename. The document that is already there keeps it, and the collision is counted on the run instead of one source quietly replacing another's words. |
path_taken is the one that asks something of you: give the file its own name, its own
folder or its own document base, and both come through.
Validate a source and you get the same picture before you commit to a run:
{
"status": "ready",
"message": "Engram can read s3://acme-knowledge-base/handbooks/. 8 of the first objects listed are documents it can ingest.",
"objects_listed": 42,
"skipped_unsupported_type": 30,
"skipped_too_large": 2,
"skipped_archived": 2
}
The lifecycle of one document
Every document carries a lifecycle field, which is the honest answer to "what is
happening to this file":
| State | Meaning |
|---|---|
pending | Registered, queued behind other work. |
parsing | Text extraction is running. |
parsed | Text extracted, waiting to be built. |
queued | Part of a build that has been dispatched. |
building | Memory is being compiled for it. |
ready | Answering questions. |
failed | Could not be read or built. lifecycle_error says
why. |
stale | The file changed and its memory has not been rebuilt yet. |
Identity is decided by content, not by name: a renamed file whose text is unchanged is recognised
as the same document and costs no rebuild, and re-saving a PDF whose words did not change costs
nothing either. content_hash and cart_id on a document are how you tell
"same document, renamed" from "document changed" without guessing from filenames.
In the app, the Documents tab counts how many documents are in each state, which is
the fast answer to "is this base ready". For one file, press Download manifest on its
source card and read the status, reason, last_change and
content_hash columns. See Download a manifest. The same
fields come back per document from the API and from
engram docs list.
Next
Load some: Documents and ingest, or point a base at a folder with Sources.