Languages & Multilingual Memory
Store memories in any language and script, tag them, and recall across languages — what v0.10.1 does, what it doesn't, and how to gate it.
CortexDB never translates your memories. You store the original text, optionally with hints about its language, script and time zone, and the server keeps both. The original text is what recall returns and what answers cite. Since v0.10.1 the server also tags every event with a language, can use hints on recall, and keeps non-Latin names apart.
Gate these fields on capabilities, not on the version
Each feature below has a capability name in GET /v1/admin/version → capabilities[]. A server that
doesn't list a capability refuses its request fields with 422 INVALID_BODY (v0.9.13 lists none of
them). Check the list once at start-up and only send what it names.
| Capability | Adds |
|---|---|
capture_language_v1 | language, scripts[], context.timezone on captures; language tags on events |
refers_to_v1 | context.refers_to[] on captures; temporal.refers_during on recall and answer |
recall_language_v1 | language, scripts[] on /v1/recall and /v1/recall/stream |
query_variants_v1 | query_variants[] on /v1/recall and /v1/recall/stream |
temporal_lenient_v1 | temporal.natural_mode, temporal.timezone |
blob_text_decode_v1 | Charset-safe decoding of non-UTF-8 text blobs |
Capture
POST /v1/experience?wait=indexed
{
"scope": "org:acme/team:support",
"modality": "text",
"content": { "kind": "text", "text": "kal mera dentist appointment hai" },
"language": "hi-Latn",
"scripts": ["Latn"],
"context": { "observed_at": "2026-09-23T20:30:00Z", "timezone": "Asia/Kolkata" }
}languageis a BCP-47 tag (hi-Latn,ta,ar) andscripts[]are ISO 15924 codes (at most 8). Values are canonicalized:HI-latnis stored ashi-Latn,latnasLatn.context.timezoneis an exact IANA name.Asia/Kolkataand its older aliasAsia/Calcuttawork;IST,+05:30andAsia/Mumbaiare refused.context.refers_to[]records the dates the memory talks about (local dates or date-times with no offset, with an optionalzoneandprecision), so recall can favour them later. See Experience.- A malformed value in any of these →
422 INVALID_ENVELOPE.
Every event is classified, even without hints. Your language wins; then a lang:<tag>
label in context.labels; then detection from the text. We saw these on v0.10.1:
| Text | Stored language | language_source |
|---|---|---|
| "कल ऑफिस में मीटिंग है" (Devanagari) | und-Deva | detected |
| "yaar kal office nahi aana hai mujhe" (Hinglish) | hi-Latn | detected |
any text with "language": "HI-latn" | hi-Latn | hint |
any text with label lang:ta | ta | label |
| "the meeting is tomorrow" | (no tag) | (none) |
An event that is English by default stores no tag, so it reads back exactly as it did before v0.10.1;
an explicit "language": "en" is stored. Tags come back as language, scripts and
language_source on GET /v1/events/{id}, on recall layers.events items and in exports.
The hints are not part of the write identity. Retrying with the same idempotency_key and body but
different hints replays the first write (its hints win) and adds a conflicting_hint_on_replay warning
to the response. An erasure clears context.timezone and context.refers_to; the language tag stays.
Recall
Language and script hints
/v1/recall and /v1/recall/stream take language and scripts[], validated like the capture hints
(a malformed value → 422 INVALID_BODY with details.field). They choose the
query stoplist for keyword search: English words are stripped with no hint or an en, hi or
hi-Latn hint, and nothing is stripped for other languages or a scripts hint without Latn. The
query text alone never changes the stoplist.
Query variants: bridge scripts yourself
POST /v1/recall
{
"scope": "org:acme/team:support",
"query": "Does Priya like tea?",
"query_variants": [ { "text": "प्रिया चाय पसंद", "language": "hi", "script": "Deva" } ]
}- Up to 4 variants of
{ text, language?, script? }, each at most 4,096 bytes. More entries, an emptytextor a longer one →422 INVALID_BODYwithdetails { field, max_items, max_text_bytes }. A variant that repeats the query or an earlier variant is dropped silently. - Each variant runs its own vector and keyword search, merged with the question's results
(question first), and reranking still scores against the question only. With
diagnostics: "full", each returned item shows its rank per variant andmatched_via(for example["primary", "variant:0"]). - Two variants add about 100 ms at p50, in one embedding call.
Add a variant in the memories' script
Recall doesn't translate the question. For a Latin-script question over memories written in Devanagari (or another script), send a variant in the memories' script, as above, for the best results.
Entity grounding keeps non-English packs
When a question names an entity that no retrieved memory mentions, recall normally empties the pack rather than return memories about someone else. That still happens for English questions over an all-English scope. Since v0.10.1 the pack is kept, with a warning, when the question or a variant has a non-English hint, a retrieved memory is in a script the question lacks, or the scope holds at least two non-English memories:
"warnings": ["entity_grounding_advisory: query entities [\"zorblax\"] not found in retrieved memories; pack kept (script_mismatch) - results may be about a different entity"]With diagnostics: "summary" or "full", diagnostics.filters_applied also lists
entity_grounding:advisory:<reason>.
Local time in the context
A non-English event captured with context.timezone gets a local-time bracket after its UTC one in
context_block, so the model reads the day the user meant:
[2026-09-23 20:30 UTC] [2026/09/24 (Thu) 02:00 Asia/Kolkata] प्रिया को चाय बहुत पसंद हैEnglish events render exactly as before.
Time phrases and time zones
temporal.timezone(IANA) makes calendar phrases such astoday,yesterdayandlast monthresolve on that zone's calendar; without it they resolve in UTC.IST→422 INVALID_BODYwithdetails.field: "temporal.timezone".temporal.natural_mode: "lenient"drops a phrase the server can't parse instead of failing with422 UNPARSEABLE_TEMPORAL, and adds atemporal_natural_ignored: …warning; no time filter is applied.temporal.refers_during { from, to, zone? }is a boost, not a filter: memories whoserefers_to(or capture time) falls in the window move to the front; nothing is added or removed. A malformed window →422 INVALID_BODY.
These work on /v1/recall and /v1/answer; /v1/answer does not return the pack's warnings[].
Answers
/v1/answer has no language or scripts field. It accepts and silently ignores any field it doesn't
know, so sending language there changes nothing. Put the
hints and variants on a /v1/recall call if you need them, and the question in the user's language.
Names, facts and enrichment of non-Latin text
These apply when enrichment is on:
- Entity ids keep non-Latin names apart. A name with any non-ASCII character gets a hashed id:
José →
ent_jose-0f66606aa13f1258, प्रिया →ent_u-0133576006acc0d5. v0.9.13 turned every such character into_, so different people in one script could share an id. Old ids still resolve on reads and forgets. See Facts. - Facts keep the original wording. Subjects and objects stay in the source script; a non-English
fact gets an English
glossused only as a search key, never shown in the context or citations. CORTEX_ENRICHMENT_SKIP_NON_EN=1(off by default) skips LLM enrichment for text in a non-Latin script; those events stay searchable as raw events. Latin-script Hinglish is always enriched.GET /v1/admin/metrics→enrichment_multilingualreportsnon_en_script_events_enriched,non_en_script_events_skipped,triples_rejectedand thefact_surfacecounters.
Text files in other encodings
A text blob in a legacy encoding (Windows-1252, Shift-JIS, UTF-16) must say so on upload
(Content-Type: text/plain; charset=windows-1252); otherwise a capture that uses it fails with
422 BLOB_TEXT_ENCODING_UNKNOWN rather than storing garbled text. See
Experience.
From the SDKs
The Python SDK (cortexdbai 0.12.0) passes these through: experience() takes language=,
scripts=, timezone= and refers_to=; recall() takes language=, scripts= and
query_variants=; capabilities() returns the list to gate on. See
Python SDK.