Webhooks

Delivery and retries

Six attempts over an hour and three quarters, a clear rule about which responses are retried, and every attempt recorded in your audit log.

Deliveries run on a worker, not on the path that produced the event, so nothing you do with your endpoint can slow down or break ingestion. What that worker does is worth knowing before you write a receiver.

What one delivery looks like

A single POST to your URL, with the JSON body and the four headers listed on Overview and events. Two details shape how you should answer it:

The URL is re-checked against the same public-HTTPS rule at delivery time, not only when you subscribed, because a DNS record can change under a saved endpoint.

The retry ladder

One attempt, then five retries. Six deliveries in all, the last of them 1 hour 42 minutes after the event:

AttemptWaitsElapsed
1immediate0
2 (retry 1)30 seconds30 seconds
3 (retry 2)2 minutes2.5 minutes
4 (retry 3)10 minutes12.5 minutes
5 (retry 4)30 minutes42.5 minutes
6 (retry 5)1 hour1 hour 42.5 minutes, then we stop

Those are the code constants DELIVERY_RETRY_MAX and DELIVERY_RETRY_INTERVALS, not a plan setting. DELIVERY_RETRY_MAX is 5 because it counts the RETRIES; the number of deliveries your receiver sees is one higher. The ladder covers the realistic outage: a deploy, a restart, a brief provider blip. Past that, an endpoint that has been down for nearly two hours is not coming back inside a retry budget, and hammering it helps nobody.

Pausing an endpoint is the one thing that ends a ladder early. It stops new deliveries and the retries already scheduled, and nothing is saved up while it is off, so resuming starts with the next event. See Pause and resume.

Events are not queued for replay after the ladder runs out. If your receiver was down for longer, reconcile from the API: GET /v1/corpora/{id}/sync-runs is the full ingestion history, and GET /v1/corpora/{id} carries the readiness counts.

What gets retried and what does not

Your responseWeWhy
2xxMark it deliveredDone.
429RetryYou asked us to slow down, which is a timing problem, not a request problem.
5xxRetryYour side had a bad moment and might not next time.
Any other 4xxDo not retryYou are telling us the request itself is wrong, and resending an unchanged request to an unchanged endpoint cannot fix that.
Connection error or timeoutRetryNothing was heard, so it may yet work.
3xxDo not retryRecorded as your status. Point the endpoint at its real destination instead.

Build a receiver that holds

@app.post("/engram")
async def engram_webhook(request: Request, x_engram_signature: str | None = Header(default=None)):
    body = await request.body()
    if not verify(SECRET, body, x_engram_signature):
        raise HTTPException(401, "bad signature")

    payload = json.loads(body)
    key = (payload["data"].get("sync_run_id")
           or payload["data"].get("job_id")
           or payload["data"].get("document_id"))
    if already_handled(key):
        return {"ok": True}          # a retry of something we finished; answer 2xx and stop
    enqueue(payload["event"], payload["data"], dedupe_key=key)
    return {"ok": True}

See what happened

Two places tell you, and neither needs you to check your own logs.

The endpoint itself carries the last outcome:

curl https://api.engramdynamics.org/v1/webhooks \
  -H "Authorization: Bearer <your key>" 
{
  "items": [
    {
      "id": "w_6c4a...",
      "url": "https://hooks.example.com/engram",
      "events": [],
      "active": true,
      "last_delivery_at": "2026-09-13T10:11:43Z",
      "last_status": "200",
      "created_at": "2026-09-13T10:12:20Z"
    }
  ],
  "next_cursor": null
}

Every list route under /v1 returns that envelope, so read items rather than indexing the response. The timestamps are UTC and end in Z, the same as sent_at inside a payload.

last_status is the HTTP status as a string, error for a connection failure, or refused when the URL stopped being a public HTTPS address.

Every attempt, successful or not, also lands in the audit log as a webhook.delivery event with the outcome, the reason, and whether it was retryable, so "did you actually call us?" has an answer. See Audit log.

Next

Feed a base from a bucket or a shared drive and let the events tell you when it lands: Sources.