CortexDB Docs
API Reference

/v1/erasures

GDPR reference-counted erasure — preview, execute, and poll the job. The true WAL-deletion path.

Erasures are CortexDB's reference-counted GDPR deletion path — the sanctioned exception to the append-only WAL. Unlike forget (which redacts or drops derived layers), a completed erasure deletes events from the WAL. The flow is preview → execute → poll.

Preview

POST /v1/erasures/preview with the same selector and confirmation you will execute with, e.g. { scope, confirm_all: true, audit_note } → 200. A preview is bound to that exact operation: executing with a different selector or confirmation (for example previewing { scope, audit_note } and then executing with confirm_all: true) → 422 UNKNOWN_PREVIEW_ID. from_preview_id is optional; an execute without it also works. Preview ids live in the server's memory, so a restart invalidates them (same 422).

{
  "preview_id": "ervw_01HX...",
  "scope": "org:acme/user:alice",
  "estimated_affected": { "events": 5, "episodes": 0, "facts": 0, "beliefs": 0, "understanding": 0 },
  "refcount_breakdown": {
    "in_scope_events": 5,
    "events_to_delete": 5,
    "events_to_redact": 0,
    "events_under_legal_hold": 0
  },
  "estimated_duration_ms": 650
}
  • estimated_affected adds blobs when the erasure would delete uploaded blobs (see Blobs).
  • estimated_duration_ms is a planning figure for choosing a client timeout: 250 ms plus 80 ms per in-scope event. Real erasures are usually much faster.
  • Erasing a parent scope counts the events stored in its child scopes under refcount_breakdown.events_to_redact (and estimated_affected.events then reads 0). Those events are removed completely all the same, so read refcount_breakdown.in_scope_events for the total.

What a self-hosted preview contains

Self-hosted previews carry the fields above; there are no cross_scope_propagation, legal_holds or manifest_url blocks, and no separate manifest route. The preview can also carry a top-level code_repos_to_delete count (Code Intelligence Plane repositories bound to the scope that execute would remove); it is absent when zero, so a deployment without the code plane sees the body above unchanged.

Execute

POST /v1/erasures → 202:

{
  "erasure_id": "erasure_01HX...",
  "status": "completed",
  "manifest_url": "/v1/erasures/erasure_01HX...",
  "lifecycle_stream": "/v1/lifecycle/stream?erasure_id=erasure_01HX..."
}

A small scope often reports "status": "completed" already in this response; a larger one reports "running". Either way, poll manifest_url (GET /v1/erasures/{id}) for progress and the result. The lifecycle_stream URL is the general capture and indexing stream; it doesn't report erasure progress.

Execute requires confirm_all or an explicit selector

Without selector.memory_ids, execute → 422 EMPTY_SELECTOR_WITHOUT_CONFIRMATION, even with from_preview_id set — an erasure with no selector erases the entire scope, so you must pass confirm_all: true (or a selector.memory_ids list to erase specific records). The manifest_url in the response is /v1/erasures/{id} (no /manifest suffix).

{
  "scope": "org:acme/user:alice",
  "from_preview_id": "ervw_01HX...",
  "confirm_all": true,
  "audit_note": "DSR #1234"
}

An execute takes one server cutoff before it reads anything: captures the server stamped before it are erased, captures stamped after it are kept (and are erased by a later execute). Send the captures you need to keep after the erasure returns. Like forget, an execute runs detached from the HTTP request, so a client disconnect does not stop it.

Status

GET /v1/erasures/{id} → the job record:

{
  "erasure_id": "erasure_01HX...",
  "status": "completed",
  "phase": "audit",
  "summary": {
    "deleted_events": 3,
    "redacted_events": 0,
    "demoted_beliefs": 0,
    "deleted_facts": 0,
    "deleted_episodes": 0,
    "deleted_concepts": 0,
    "deleted_artifacts": 0,
    "artifacts_pending_reevaluation": 0,
    "purge": { "events": 3, "documents": 3, "edges": 0, "blobs": 0, "retained_shared_blobs": 0,
               "twins_redacted": 3, "bitemporal_records_erased": 0, "wal_rows_redacted": 0,
               "artifacts_cascaded": 0 }
  },
  "log_residual_cleared_at": null,
  "log_residual_pending": { "reason": "pass_pending" }
}

Reading the job record

The record is { erasure_id, status, phase, summary, log_residual_cleared_at, log_residual_pending }. Progress and results are in summary; phase runs enumerate → refcount → categorize → delete → demote → audit. summary.deleted_events is the count of events this job removed.

Jobs are kept in the server's memory: after a restart, or for an unknown id, GET /v1/erasures/{id} → 404 NOT_FOUND ("erasure job not found"). Re-submitting an erasure is safe.

Backend failures surface as 502 ERASURE_BACKEND_FAILED (retriable — erasure is idempotent), and an accepted job can still finish with status: "failed" and an error field — always poll rather than assuming acceptance means completion.

Blobs and leftover bytes

  • Blobs (capability erasure_blob_cascade_v1). An event made from an uploaded text blob records it in source_blob_ids. Erasing the event deletes the blob unless another live event still uses it, reported as summary.purge.blobs and summary.purge.retained_shared_blobs; afterwards GET /v1/blobs/{id} answers 404. Events written before v0.10.1 carry no link until an operator backfills it with POST /v1/admin/index-audit/source-blob-links; GET /v1/admin/index-audit/orphan-blobs then lists blobs nothing links (see Admin).

  • Leftover bytes. A delete in the storage engines leaves older copies of the bytes in their files until they are compacted. After an erasure a throttled background pass compacts them out. The status reports log_residual_cleared_at once the pass has covered the job (seconds for a small job; log_residual_pending is then null). While it is still null on a completed job, log_residual_pending.reason says why:

    reasonMeaning
    pass_pendingThe pass has not finished this job's work yet
    fulltext_segment_over_budgetA full-text segment larger than the per-pass budget still holds erased documents; on a store whose segments exceed the budget it may never clear
    not_coveredNo pass covers this job: CORTEX_ERASURE_RESIDUAL=0, or the legacy dual WAL (CORTEX_UNIFIED_WAL=0)

    A graceful stop drains the pass for up to CORTEX_ERASURE_RESIDUAL_SHUTDOWN_SECS (default 20) and resumes the rest at the next start. CORTEX_ERASURE_PERIODIC_COMPACTION_SECS sets a periodic compaction backstop. GET /v1/admin/metrics reports the pass under erasure_residual.

  • Erased facts stay erased after a crash. An erased fact id is recorded in a synced fence before it leaves memory, and every start drops fenced ids, so a hard kill cannot bring an erased fact back.

Speed and the erasure index

Since v0.10.1 an erasure's cost follows what it erases, not the size of the store: the server keeps an erasure index beside the event log (a one-id erasure on a 1M-event store takes well under a second). The first start after the upgrade backfills the index; until erasure_index.wal.ready is true on GET /v1/admin/metrics, erasures read the whole store as before. CORTEX_ERASURE_INDEX=0 keeps the index maintained but unread, and CORTEX_ERASURE_INDEX_BACKFILL=0 skips the backfill.

Concurrency and shutdown

Nothing is changed in any of these cases; retry after Retry-After. Erasures usually finish in well under a second, so these are rare.

StatusCodeWhen
409AUTHORIZATION_STATE_CHANGEDScopes were registered or removed inside the erasure's footprint on each of 5 attempts to authorize a preview, execute, status read or cancel (v0.10.2; Retry-After: 1, details.retry_after_seconds: 1)
409DESTRUCTIVE_OPERATION_IN_PROGRESSA capture that would register a new scope inside a running erasure's footprint waited CORTEX_DESTRUCTIVE_LEASE_WAIT_SECS (default 120) (Retry-After: 5)
409SCOPE_STATE_CHANGED/v1/scopes register, member edit or delete inside the footprint
409REDERIVE_IN_PROGRESSA derived-data repair job holds the scope (Retry-After: 30)
503SERVER_SHUTTING_DOWNA graceful stop has begun (Retry-After: 5); running erasures get up to CORTEX_SHUTDOWN_DESTRUCTIVE_WAIT_SECS (default 60) to finish

Captures into existing scopes and reads never wait for an erasure.

A scope registered outside the footprint (any capture, from any tenant, that creates a brand-new scope) no longer affects an erasure. Before v0.10.2 that race refused a preview or execute with 403 POLICY_DENIED, nothing erased, and turned a GET /v1/erasures/{id} poll or a cancel into a spurious 404. While new child scopes keep appearing inside the erased scope, retry after them.

Things to know

  • On a whole-scope erasure of an enrichment-on scope, summary.purge.events also counts derived content rows (facts, glosses, question keys); use summary.deleted_events and redacted_events for the event count.
  • Preview again if the scope changed. An execute with a from_preview_id erases what is in the scope when it runs, including events captured after the preview. Preview right before you execute when the preview itself is your record of what was erased.
  • Cancel applies to a job that is still running. A finished job can't be cancelled (POST /v1/erasures/{id}/cancel → 404 NOT_FOUND); read its result with GET /v1/erasures/{id}.

On this page