Webhooks
Delivery and retries
Six attempts over an hour and three quarters, a clear rule about which responses are retried, and every attempt recorded in your audit log.
Deliveries run on a worker, not on the path that produced the event, so nothing you do with your endpoint can slow down or break ingestion. What that worker does is worth knowing before you write a receiver.
What one delivery looks like
A single POST to your URL, with the JSON body and the four headers listed on
Overview and events. Two details shape how you should answer it:
- The timeout is 10 seconds (
WEBHOOK_TIMEOUT_S). Longer than that and we treat the attempt as failed and try again. - Redirects are not followed. A 3xx is recorded as the status your receiver returned, and nothing at the redirect target is fetched.
The URL is re-checked against the same public-HTTPS rule at delivery time, not only when you subscribed, because a DNS record can change under a saved endpoint.
The retry ladder
One attempt, then five retries. Six deliveries in all, the last of them 1 hour 42 minutes after the event:
| Attempt | Waits | Elapsed |
|---|---|---|
| 1 | immediate | 0 |
| 2 (retry 1) | 30 seconds | 30 seconds |
| 3 (retry 2) | 2 minutes | 2.5 minutes |
| 4 (retry 3) | 10 minutes | 12.5 minutes |
| 5 (retry 4) | 30 minutes | 42.5 minutes |
| 6 (retry 5) | 1 hour | 1 hour 42.5 minutes, then we stop |
Those are the code constants DELIVERY_RETRY_MAX and
DELIVERY_RETRY_INTERVALS, not a plan setting. DELIVERY_RETRY_MAX is 5
because it counts the RETRIES; the number of deliveries your receiver sees is one higher. The ladder
covers the realistic outage: a deploy, a restart, a brief provider blip. Past that, an endpoint that
has been down for nearly two hours is not coming back inside a retry budget, and hammering it helps
nobody.
Pausing an endpoint is the one thing that ends a ladder early. It stops new deliveries and the retries already scheduled, and nothing is saved up while it is off, so resuming starts with the next event. See Pause and resume.
Events are not queued for replay after the ladder runs out. If your receiver was
down for longer, reconcile from the API: GET /v1/corpora/{id}/sync-runs is the full
ingestion history, and GET /v1/corpora/{id} carries the readiness counts.
What gets retried and what does not
| Your response | We | Why |
|---|---|---|
| 2xx | Mark it delivered | Done. |
| 429 | Retry | You asked us to slow down, which is a timing problem, not a request problem. |
| 5xx | Retry | Your side had a bad moment and might not next time. |
| Any other 4xx | Do not retry | You are telling us the request itself is wrong, and resending an unchanged request to an unchanged endpoint cannot fix that. |
| Connection error or timeout | Retry | Nothing was heard, so it may yet work. |
| 3xx | Do not retry | Recorded as your status. Point the endpoint at its real destination instead. |
Build a receiver that holds
- Verify, then answer, then work. Check the signature, queue the payload, return 2xx. A receiver that does the work inline will eventually exceed 10 seconds and turn a successful delivery into a retried one.
- Be idempotent. A retry after a lost response means the same event can arrive
twice. Key your processing on
data.sync_run_id,data.job_idordata.document_id. - Return 4xx only when you mean it. A 400 stops the ladder, so use it for a payload you genuinely cannot accept and a 500 for "not right now".
- Do not order events. Two deliveries can overlap and retries reshuffle timing. Treat each as a fact about a moment, and read current state from the API when you need ordering.
@app.post("/engram")
async def engram_webhook(request: Request, x_engram_signature: str | None = Header(default=None)):
body = await request.body()
if not verify(SECRET, body, x_engram_signature):
raise HTTPException(401, "bad signature")
payload = json.loads(body)
key = (payload["data"].get("sync_run_id")
or payload["data"].get("job_id")
or payload["data"].get("document_id"))
if already_handled(key):
return {"ok": True} # a retry of something we finished; answer 2xx and stop
enqueue(payload["event"], payload["data"], dedupe_key=key)
return {"ok": True}
See what happened
Two places tell you, and neither needs you to check your own logs.
The endpoint itself carries the last outcome:
curl https://api.engramdynamics.org/v1/webhooks \
-H "Authorization: Bearer <your key>"
{
"items": [
{
"id": "w_6c4a...",
"url": "https://hooks.example.com/engram",
"events": [],
"active": true,
"last_delivery_at": "2026-09-13T10:11:43Z",
"last_status": "200",
"created_at": "2026-09-13T10:12:20Z"
}
],
"next_cursor": null
}
Every list route under /v1 returns that envelope, so read items rather
than indexing the response. The timestamps are UTC and end in Z, the same as
sent_at inside a payload.
last_status is the HTTP status as a string, error for a connection
failure, or refused when the URL stopped being a public HTTPS address.
Every attempt, successful or not, also lands in the audit log as a webhook.delivery
event with the outcome, the reason, and whether it was retryable, so "did you actually call us?" has
an answer. See Audit log.
Next
Feed a base from a bucket or a shared drive and let the events tell you when it lands: Sources.