CortexDB Docs
Features

Languages & Multilingual Memory

Store memories in any language and script, tag them, and recall across languages — what v0.10.1 does, what it doesn't, and how to gate it.

CortexDB never translates your memories. You store the original text, optionally with hints about its language, script and time zone, and the server keeps both. The original text is what recall returns and what answers cite. Since v0.10.1 the server also tags every event with a language, can use hints on recall, and keeps non-Latin names apart.

Gate these fields on capabilities, not on the version

Each feature below has a capability name in GET /v1/admin/version → capabilities[]. A server that doesn't list a capability refuses its request fields with 422 INVALID_BODY (v0.9.13 lists none of them). Check the list once at start-up and only send what it names.

CapabilityAdds
capture_language_v1language, scripts[], context.timezone on captures; language tags on events
refers_to_v1context.refers_to[] on captures; temporal.refers_during on recall and answer
recall_language_v1language, scripts[] on /v1/recall and /v1/recall/stream
query_variants_v1query_variants[] on /v1/recall and /v1/recall/stream
temporal_lenient_v1temporal.natural_mode, temporal.timezone
blob_text_decode_v1Charset-safe decoding of non-UTF-8 text blobs

Capture

POST /v1/experience?wait=indexed
{
  "scope": "org:acme/team:support",
  "modality": "text",
  "content": { "kind": "text", "text": "kal mera dentist appointment hai" },
  "language": "hi-Latn",
  "scripts": ["Latn"],
  "context": { "observed_at": "2026-09-23T20:30:00Z", "timezone": "Asia/Kolkata" }
}
  • language is a BCP-47 tag (hi-Latn, ta, ar) and scripts[] are ISO 15924 codes (at most 8). Values are canonicalized: HI-latn is stored as hi-Latn, latn as Latn.
  • context.timezone is an exact IANA name. Asia/Kolkata and its older alias Asia/Calcutta work; IST, +05:30 and Asia/Mumbai are refused.
  • context.refers_to[] records the dates the memory talks about (local dates or date-times with no offset, with an optional zone and precision), so recall can favour them later. See Experience.
  • A malformed value in any of these → 422 INVALID_ENVELOPE.

Every event is classified, even without hints. Your language wins; then a lang:<tag> label in context.labels; then detection from the text. We saw these on v0.10.1:

TextStored languagelanguage_source
"कल ऑफिस में मीटिंग है" (Devanagari)und-Devadetected
"yaar kal office nahi aana hai mujhe" (Hinglish)hi-Latndetected
any text with "language": "HI-latn"hi-Latnhint
any text with label lang:tatalabel
"the meeting is tomorrow"(no tag)(none)

An event that is English by default stores no tag, so it reads back exactly as it did before v0.10.1; an explicit "language": "en" is stored. Tags come back as language, scripts and language_source on GET /v1/events/{id}, on recall layers.events items and in exports.

The hints are not part of the write identity. Retrying with the same idempotency_key and body but different hints replays the first write (its hints win) and adds a conflicting_hint_on_replay warning to the response. An erasure clears context.timezone and context.refers_to; the language tag stays.

Recall

Language and script hints

/v1/recall and /v1/recall/stream take language and scripts[], validated like the capture hints (a malformed value → 422 INVALID_BODY with details.field). They choose the query stoplist for keyword search: English words are stripped with no hint or an en, hi or hi-Latn hint, and nothing is stripped for other languages or a scripts hint without Latn. The query text alone never changes the stoplist.

Query variants: bridge scripts yourself

POST /v1/recall
{
  "scope": "org:acme/team:support",
  "query": "Does Priya like tea?",
  "query_variants": [ { "text": "प्रिया चाय पसंद", "language": "hi", "script": "Deva" } ]
}
  • Up to 4 variants of { text, language?, script? }, each at most 4,096 bytes. More entries, an empty text or a longer one → 422 INVALID_BODY with details { field, max_items, max_text_bytes }. A variant that repeats the query or an earlier variant is dropped silently.
  • Each variant runs its own vector and keyword search, merged with the question's results (question first), and reranking still scores against the question only. With diagnostics: "full", each returned item shows its rank per variant and matched_via (for example ["primary", "variant:0"]).
  • Two variants add about 100 ms at p50, in one embedding call.

Add a variant in the memories' script

Recall doesn't translate the question. For a Latin-script question over memories written in Devanagari (or another script), send a variant in the memories' script, as above, for the best results.

Entity grounding keeps non-English packs

When a question names an entity that no retrieved memory mentions, recall normally empties the pack rather than return memories about someone else. That still happens for English questions over an all-English scope. Since v0.10.1 the pack is kept, with a warning, when the question or a variant has a non-English hint, a retrieved memory is in a script the question lacks, or the scope holds at least two non-English memories:

"warnings": ["entity_grounding_advisory: query entities [\"zorblax\"] not found in retrieved memories; pack kept (script_mismatch) - results may be about a different entity"]

With diagnostics: "summary" or "full", diagnostics.filters_applied also lists entity_grounding:advisory:<reason>.

Local time in the context

A non-English event captured with context.timezone gets a local-time bracket after its UTC one in context_block, so the model reads the day the user meant:

[2026-09-23 20:30 UTC] [2026/09/24 (Thu) 02:00 Asia/Kolkata] प्रिया को चाय बहुत पसंद है

English events render exactly as before.

Time phrases and time zones

  • temporal.timezone (IANA) makes calendar phrases such as today, yesterday and last month resolve on that zone's calendar; without it they resolve in UTC. IST → 422 INVALID_BODY with details.field: "temporal.timezone".
  • temporal.natural_mode: "lenient" drops a phrase the server can't parse instead of failing with 422 UNPARSEABLE_TEMPORAL, and adds a temporal_natural_ignored: … warning; no time filter is applied.
  • temporal.refers_during { from, to, zone? } is a boost, not a filter: memories whose refers_to (or capture time) falls in the window move to the front; nothing is added or removed. A malformed window → 422 INVALID_BODY.

These work on /v1/recall and /v1/answer; /v1/answer does not return the pack's warnings[].

Answers

/v1/answer has no language or scripts field. It accepts and silently ignores any field it doesn't know, so sending language there changes nothing. Put the hints and variants on a /v1/recall call if you need them, and the question in the user's language.

Names, facts and enrichment of non-Latin text

These apply when enrichment is on:

  • Entity ids keep non-Latin names apart. A name with any non-ASCII character gets a hashed id: José → ent_jose-0f66606aa13f1258, प्रिया → ent_u-0133576006acc0d5. v0.9.13 turned every such character into _, so different people in one script could share an id. Old ids still resolve on reads and forgets. See Facts.
  • Facts keep the original wording. Subjects and objects stay in the source script; a non-English fact gets an English gloss used only as a search key, never shown in the context or citations.
  • CORTEX_ENRICHMENT_SKIP_NON_EN=1 (off by default) skips LLM enrichment for text in a non-Latin script; those events stay searchable as raw events. Latin-script Hinglish is always enriched.
  • GET /v1/admin/metrics → enrichment_multilingual reports non_en_script_events_enriched, non_en_script_events_skipped, triples_rejected and the fact_surface counters.

Text files in other encodings

A text blob in a legacy encoding (Windows-1252, Shift-JIS, UTF-16) must say so on upload (Content-Type: text/plain; charset=windows-1252); otherwise a capture that uses it fails with 422 BLOB_TEXT_ENCODING_UNKNOWN rather than storing garbled text. See Experience.

From the SDKs

The Python SDK (cortexdbai 0.12.0) passes these through: experience() takes language=, scripts=, timezone= and refers_to=; recall() takes language=, scripts= and query_variants=; capabilities() returns the list to gate on. See Python SDK.

On this page