POST /v1/answer
Recall a pack and answer a natural-language question over it, with citations.
Authorization
bearer PASETO v4 public token (or deployment gate key). Auth-disabled dev instances accept any caller.
In: header
Request Body
application/json
TypeScript Definitions
Use the request body type in TypeScript.
Response Body
application/json
application/json
curl -X POST "https://example.com/v1/answer" \ -H "Content-Type: application/json" \ -d '{ "scope": "string", "question": "string" }'503 ANSWER_UNAVAILABLE without an answer lane
On a self-hosted server with no CORTEX_ANSWER_* lane configured, POST /v1/answer returns
503 ANSWER_UNAVAILABLE ("no answer LLM is configured on this deployment; set CORTEX_ANSWER_API_KEY …").
It is deployment-wide and scope-independent.
See LLM & Answer.
The question field
The field is question, not query
/v1/answer requires question (not query). Sending { "query": … } → 422 "missing field question". It shares view, include, temporal, budgets, and citation_mode semantics with
recall.
question_type is ignored
The server picks its retrieval route and question shape from the question itself. A
question_type field is accepted and discarded so older clients keep working; the v0.10.2 spec marks it
deprecated. Verified on v0.10.2: "multi-session", "temporal-reasoning", "single-session-user" and
no question_type all give the same inferred_question_shape for the same question.
Retrieval without the LLM
skip_answer_llm returns an empty answer with citations
skip_answer_llm: true returns 200 with answer: "" and
diagnostics.answer_llm_skipped_by_request: true — it runs recall + citation shaping but skips the
LLM. Use it when you want the answer endpoint's citation shaping without paying for generation.
cite_sources (top-level) is accepted but has no observable effect on a content-only instance.
Response diagnostics
The response is { pack_id, answer, citations[{ marker, layer, id, support_strength }], provenance, diagnostics, as_of }. as_of is a top-level field here (unlike recall, where it does not exist).
diagnostics fields (v0.9.9)
diagnostics carries recall_ms, llm_ms, total_ms, answer_model,
answer_llm_skipped_by_request, tokens_input, tokens_output, and a time_ms phase-map (v0.9.13
adds a long list of artifact_* and wal_* counters, and inferred_question_shape /
question_shape_applied when the server infers a question shape from the question). Token counts are flat tokens_input /
tokens_output — there is no pack_used field and no answer_tokens: { prompt, completion }. The
model used is at diagnostics.answer_model. The token keys are omitted (absent, not null) when
the answer came from the multi-session executor lane and on any path that does not call the answer model. The SSE stream's diagnostics event
carries both tokens_input and tokens_output since v0.10.1 (v0.9.13 sent only tokens_output), and
omits a count it does not know instead of sending null.
Degraded runs and abstentions (v0.10.1)
diagnostics.degraded_sourcesappears when the answer ran without part of its evidence, as stable codes only (never error text):coordinator_recall,wal_supplement,derived_<lane>(e.g.derived_facts),artifact_search,index_divergence,derived_backfill,facet_plan_unavailable,facet_recall_failed,ms_executor_fallback, elseunknown. The answer is still returned; read a non-empty list as "this answer saw less evidence than a healthy run would" and check the server log. On a healthy server it is absent.- Abstentions ("The memories contain no such information.") are detected as such. When the question
names things the context never mentions,
diagnostics.premise_unsupported_termslists them (for example, a question about a Singapore office's wifi password over IT tickets listsSingapore,office,wifi,password,branch). General nouns such as "issue" or "problem" no longer count as unsupported premises, so "How was the VPN issue fixed?" is answered. When an extractive retry runs, diagnostics also carryextractive_retry_outcome(answered,still_abstainedorpremise_ungrounded). - A request carrying the evaluation-only
facet_plan_overrideis refused with422 EVAL_OVERRIDE_DISABLEDunless the server runs withCORTEX_EVAL_FACET_PLAN_OVERRIDE=1. Never set aCORTEX_EVAL_variable on a production server. /v1/answermerges several recall passes; since v0.10.1 an event served whole by one pass and as a budget excerpt by another appears once, whole.
Model and streaming
answer_model default is env-driven; streaming is not token-SSE self-hosted
The claude-opus-4-6 default is a managed-cloud value; self-hosted, answer_model resolves to
whatever the answer lane (CORTEX_ANSWER_*) is configured with, and /v1/answer is disabled until
that lane is set. Streaming is selected by the body field "stream": true, which on v0.9.13
returns text/event-stream with the events token ({ text }), citations, provenance,
diagnostics (a reduced set: answer_model, recall_ms, llm_ms, total_ms, tokens_output) and
done ({ pack_id, as_of }). The query parameter ?stream=true is ignored and returns the normal
JSON object. See
Self-hosting defaults.
Notes
use_pack_idreuses a recall pack (same-scope only; a pack from a different scope →404).- Temporal questions trigger internal "question shaping" —
diagnostics.inferred_question_shape/question_shape_appliedreflect it; no request field controls it. order(relevance|recency) is forwarded to recall. When you omit it, recency-shaped questions ("recent…", "latest…", "currently…", "next best step…") auto-inferrecency; an explicit value always wins. As with recall, the recency timeline tail is observable in the defaultholisticview (see POST /v1/recall). Kill-switches:CORTEX_RECENCY_INFERENCE=0(disable the inference),CORTEX_RECENCY_ORDER=0(disable the mode) — see Recall Tuning.