CortexDB Docs
API Reference

POST /v1/answer

Recall a pack and answer a natural-language question over it, with citations.

POST
/v1/answer
AuthorizationBearer <token>

PASETO v4 public token (or deployment gate key). Auth-disabled dev instances accept any caller.

In: header

Request Body

application/json

TypeScript Definitions

Use the request body type in TypeScript.

Response Body

application/json

application/json

curl -X POST "https://example.com/v1/answer" \  -H "Content-Type: application/json" \  -d '{    "scope": "string",    "question": "string"  }'
Empty

503 ANSWER_UNAVAILABLE without an answer lane

On a self-hosted server with no CORTEX_ANSWER_* lane configured, POST /v1/answer returns 503 ANSWER_UNAVAILABLE ("no answer LLM is configured on this deployment; set CORTEX_ANSWER_API_KEY …"). It is deployment-wide and scope-independent. See LLM & Answer.

The question field

The field is question, not query

/v1/answer requires question (not query). Sending { "query": … } → 422 "missing field question". It shares view, include, temporal, budgets, and citation_mode semantics with recall.

question_type is ignored

The server picks its retrieval route and question shape from the question itself. A question_type field is accepted and discarded so older clients keep working; the v0.10.2 spec marks it deprecated. Verified on v0.10.2: "multi-session", "temporal-reasoning", "single-session-user" and no question_type all give the same inferred_question_shape for the same question.

Retrieval without the LLM

skip_answer_llm returns an empty answer with citations

skip_answer_llm: true returns 200 with answer: "" and diagnostics.answer_llm_skipped_by_request: true — it runs recall + citation shaping but skips the LLM. Use it when you want the answer endpoint's citation shaping without paying for generation. cite_sources (top-level) is accepted but has no observable effect on a content-only instance.

Response diagnostics

The response is { pack_id, answer, citations[{ marker, layer, id, support_strength }], provenance, diagnostics, as_of }. as_of is a top-level field here (unlike recall, where it does not exist).

diagnostics fields (v0.9.9)

diagnostics carries recall_ms, llm_ms, total_ms, answer_model, answer_llm_skipped_by_request, tokens_input, tokens_output, and a time_ms phase-map (v0.9.13 adds a long list of artifact_* and wal_* counters, and inferred_question_shape / question_shape_applied when the server infers a question shape from the question). Token counts are flat tokens_input / tokens_output — there is no pack_used field and no answer_tokens: { prompt, completion }. The model used is at diagnostics.answer_model. The token keys are omitted (absent, not null) when the answer came from the multi-session executor lane and on any path that does not call the answer model. The SSE stream's diagnostics event carries both tokens_input and tokens_output since v0.10.1 (v0.9.13 sent only tokens_output), and omits a count it does not know instead of sending null.

Degraded runs and abstentions (v0.10.1)

  • diagnostics.degraded_sources appears when the answer ran without part of its evidence, as stable codes only (never error text): coordinator_recall, wal_supplement, derived_<lane> (e.g. derived_facts), artifact_search, index_divergence, derived_backfill, facet_plan_unavailable, facet_recall_failed, ms_executor_fallback, else unknown. The answer is still returned; read a non-empty list as "this answer saw less evidence than a healthy run would" and check the server log. On a healthy server it is absent.
  • Abstentions ("The memories contain no such information.") are detected as such. When the question names things the context never mentions, diagnostics.premise_unsupported_terms lists them (for example, a question about a Singapore office's wifi password over IT tickets lists Singapore, office, wifi, password, branch). General nouns such as "issue" or "problem" no longer count as unsupported premises, so "How was the VPN issue fixed?" is answered. When an extractive retry runs, diagnostics also carry extractive_retry_outcome (answered, still_abstained or premise_ungrounded).
  • A request carrying the evaluation-only facet_plan_override is refused with 422 EVAL_OVERRIDE_DISABLED unless the server runs with CORTEX_EVAL_FACET_PLAN_OVERRIDE=1. Never set a CORTEX_EVAL_ variable on a production server.
  • /v1/answer merges several recall passes; since v0.10.1 an event served whole by one pass and as a budget excerpt by another appears once, whole.

Model and streaming

answer_model default is env-driven; streaming is not token-SSE self-hosted

The claude-opus-4-6 default is a managed-cloud value; self-hosted, answer_model resolves to whatever the answer lane (CORTEX_ANSWER_*) is configured with, and /v1/answer is disabled until that lane is set. Streaming is selected by the body field "stream": true, which on v0.9.13 returns text/event-stream with the events token ({ text }), citations, provenance, diagnostics (a reduced set: answer_model, recall_ms, llm_ms, total_ms, tokens_output) and done ({ pack_id, as_of }). The query parameter ?stream=true is ignored and returns the normal JSON object. See Self-hosting defaults.

Notes

  • use_pack_id reuses a recall pack (same-scope only; a pack from a different scope → 404).
  • Temporal questions trigger internal "question shaping" — diagnostics.inferred_question_shape / question_shape_applied reflect it; no request field controls it.
  • order (relevance | recency) is forwarded to recall. When you omit it, recency-shaped questions ("recent…", "latest…", "currently…", "next best step…") auto-infer recency; an explicit value always wins. As with recall, the recency timeline tail is observable in the default holistic view (see POST /v1/recall). Kill-switches: CORTEX_RECENCY_INFERENCE=0 (disable the inference), CORTEX_RECENCY_ORDER=0 (disable the mode) — see Recall Tuning.

On this page