Sources

What counts as a document

A document is one file you loaded. Here is what we can read, what we skip and say so, and what happens to a file between arriving and answering questions.

Whatever route documents take in, they all end up the same way: the file is stored, its text is extracted, and that text is what gets read and answered from. This page is what that means in practice.

What counts as a document

A document is one file you loaded. A 200-page PDF is one document, a two-line note is one document, and both count as one against your allowances. The internal pieces a long file is served in are never counted, so a long document never quietly costs you twenty.

What each plan allows is on Plans and allowances.

File types we read

The accepted list is served live, so it can never disagree with what the API actually takes:

curl https://api.engramdynamics.org/v1/platform/info
ExtensionsHow it is read
.txt, .mdDirectly.
.pdfThe text layer. A scanned PDF with no text layer falls back to OCR at upload rather than onboarding as an empty document.
.docxDirectly. .doc in the older RTF flavour is read too; a genuine legacy binary .doc may need converting to .docx first, and the error says so.
.html, .htmText content, with the markup stripped.
.xlsx, .xls, .csv, .tsv Cell content, sheet by sheet. A very large sheet is truncated rather than failed.

An unsupported type is refused where you send it, before anything is stored, and the message lists what is accepted. That is deliberate: we do not want to be holding bytes we cannot turn into words.

Text only

Engram answers from the words in a document. Images, charts and layout are not read, so a slide that carries its meaning in a diagram contributes very little, and a scanned page contributes whatever the OCR could make out.

When extraction fails, the document is marked failed with a short reason in lifecycle_error: an encrypted PDF, a corrupt file, an image-only page that OCR could not read. A failed document is never onboarded as garbage, and never silently included in an answer. Three places tell you: documents_failed on the base, the reason column of the manifest you download from a source card, and the document.failed webhook with the path and the reason.

How big a document is

Memory works in units of about 4,000 tokens, which is roughly six pages of prose. That is the setting CAG_MAX_DOC_TOK, currently 4,096, and it is a property of the serving hardware rather than of your plan.

A file longer than that is still read in full. Its whole text is indexed, and when you ask a question the passages that answer it are brought in alongside the resident memory, so page 150 of a long PDF is reachable. What the unit size means in practice is that short, well-titled documents retrieve more precisely than very long ones: if you control how the documents are produced, a folder of chapters beats one enormous export.

Size ceilings are separate from this and are listed on Documents and ingest: 500 MB per file on the upload and upsert paths, 500 MB per object pulled from a source, 50,000 documents in one base.

What gets skipped, and how you find out

A source sync reports its skips rather than hiding them. Every run carries a skipped counter beside added, updated, unchanged, removed and failed, so "the bucket holds 900 files and Engram found 12" is answered by the run itself. The counts read the same in engram sync runs, in engram sync --wait, on the run over the API and in the run list on the Documents tab.

Beside the counters, skipped_reasons says which reasons fired and how many objects each one took. Only the reasons that actually fired appear, and it is {} when nothing was skipped, so the object reads as a short list of what to fix:

ReasonWhat it covers
unsupported_typeAn extension not on the list above, such as a .jpg or a .zip.
emptyA zero-byte object. An empty document is only ever a bad answer waiting to happen, so it is skipped rather than ingested.
too_largeAn object over the per-object ceiling.
archivedAn object in an archival storage class, which cannot be read without a restore in your own account.
path_takenAnother source in this document base, or an upload, already holds that filename. The document that is already there keeps it, and the collision is counted on the run instead of one source quietly replacing another's words.

path_taken is the one that asks something of you: give the file its own name, its own folder or its own document base, and both come through.

Validate a source and you get the same picture before you commit to a run:

{
  "status": "ready",
  "message": "Engram can read s3://acme-knowledge-base/handbooks/. 8 of the first objects listed are documents it can ingest.",
  "objects_listed": 42,
  "skipped_unsupported_type": 30,
  "skipped_too_large": 2,
  "skipped_archived": 2
}

The lifecycle of one document

Every document carries a lifecycle field, which is the honest answer to "what is happening to this file":

StateMeaning
pendingRegistered, queued behind other work.
parsingText extraction is running.
parsedText extracted, waiting to be built.
queuedPart of a build that has been dispatched.
buildingMemory is being compiled for it.
readyAnswering questions.
failedCould not be read or built. lifecycle_error says why.
staleThe file changed and its memory has not been rebuilt yet.

Identity is decided by content, not by name: a renamed file whose text is unchanged is recognised as the same document and costs no rebuild, and re-saving a PDF whose words did not change costs nothing either. content_hash and cart_id on a document are how you tell "same document, renamed" from "document changed" without guessing from filenames.

In the app, the Documents tab counts how many documents are in each state, which is the fast answer to "is this base ready". For one file, press Download manifest on its source card and read the status, reason, last_change and content_hash columns. See Download a manifest. The same fields come back per document from the API and from engram docs list.

Next

Load some: Documents and ingest, or point a base at a folder with Sources.