Between June 18 and July 23 I published five essays on this site about the memory system behind my assistant. I called it Recallatron. It lived inside M.O.T., my ticket tracker, as one SQLite file plus a graph kept in a JSONL file beside it. In September I replaced it with the first module of Rheo Stream, and the module is also called Recallatron. Same name, different system. Below, "the old store" means the first one and "the module" means the second.
The claim is this. The five essays were accurate about dates and wrong whenever they read a date as a design. The module is the same kind of object: a record of the order I built it in, and that order was set by one priority, get the migration working first. It was not a ruling on graphs versus vectors. If you came for that ruling, there isn't one on record, and this piece doesn't supply one.
I have made this argument once already, about the old store. In A boundary is a price I wrote: "The architecture is a faithful record of the order I wrote things in." This piece applies it to the module, and to the essays. The migration itself, 4,000 history entries and 523 curated memories moved on 2026-09-30, is told in Where the price sits and isn't retold here.
Each old claim below gets one of four verdicts: true then, still true, changed, or never quite true. The dates are first-commit dates from the repositories; I didn't trust my recollection. So each verdict puts a published sentence next to a commit.
The record, read by date
June: a ledger, injection, and one call
Memory is easy. Retrieval is hard. (2026-06-18) described three layers: an append-only conversation ledger in SQL with an FTS5 keyword index behind a chat_search tool, session digests written at close, and durable memory items. A gap of more than 2 hours opened a new session. The last 5 digests were injected at the start of every conversation.
True then. The module has no conversation ledger as a product feature. The old turns came across as owner-only history evidence, searchable by keyword and explicitly not memories. The bot now keeps its own local turn log, with the same 2-hour rule.
The same essay recorded a failure, and most of what came later was an answer to it. The model didn't go looking for history it couldn't see:
We'd been treating "the tool exists" as "the tool gets used."
The fix was boot-time injection. That fix is gone now; the failure it answered is not, and it gets the last section.
One call. The real prerequisite. (2026-06-23) collapsed four boot reads into one memory_context call with no-query defaults of 3 topic threads, 5 entities, every confirmed procedural note and 10 recent memory items. True in code then. The module has no bundle for the no-query case at all, so: changed. The essay's thesis, "Retrieval is downstream of store design," outlived the feature. In the module, the audience and purpose stored on each memory decide which candidates SQL lets through before anything is ranked.
Two smaller sentences in that essay expired fast. "Edges between entities are deferred" was true until typed edges landed on 2026-07-07. The nightly prune of unconfirmed candidates older than 30 days came out of cron on 2026-07-08, so that sentence was true for 15 days.
Give the model the tools, not the context (2026-06-27) said: "There is no embedding column. Search here is keyword, not meaning." True until the vectors landed eight days later. The same essay listed what wasn't built and said the quality claims were "reasoned, not measured."
July 5: what held
Vectors in the file you already have has the two claims that came through the rebuild unchanged. Both are statements about how a part behaves, and both crossed from SQLite to Postgres intact.
The first is that vectors are a derived index. "Losing the vectors costs recall, not correctness." In the module, embeddings are memory_embedding rows written by a job after the memory commits, deleted when a memory is corrected or superseded, and rebuilt by a rebuild command. A version stamp on the embed input forces a full rebuild when the composition changes. It is the same idea, moved to Postgres. The degrade contract came across too: if the dense arm can't contribute, the response says dense_available is false and hybrid answers with the lexical rows. It never refuses.
The second is Reciprocal Rank Fusion with constant 60, score = 1 / (60 + rank). The formula is unchanged. The 36-line lib/rrf.ts became hybrid.py.
Two things in that essay did not carry. sqlite-vec in the same file became pgvector with an HNSW cosine index in the workspace's Postgres database. And the default changed. The July 5 essay said mode defaulted to "fts", "so every existing caller behaves exactly as it did before." In the module the package default is hybrid, set per workspace.
One thing stayed that is easy to misread as a change. The embedding model is all-MiniLM-L6-v2 at 384 dimensions in both systems, run locally in both. Local embedding is a continuity. What did differ comes up in the floor section.
July 23: the dek that was never quite true
The dek of How I built the workers that keep a structured memory graph clean opens: "Recallatron keeps agent memory as an entity graph and episodic timelines, retrieved by name and structure rather than by vector similarity."
Put three facts next to it. Hybrid RRF over sqlite-vec had been in the old store since 2026-07-05, eighteen days before that dek. The bot started calling hybrid search on 2026-08-27. And in the logged tool calls, 36 of 37 entity_search calls used the default substring mode. The essay's own opening paragraph admitted the vector overlay, in the sentence just before it repeated the claim.
So the dek was true of the defaults and of how the agent actually behaved, and never true of the architecture. It described a habit and called it a design. The obvious question is why the default didn't flip on July 5, when the vectors landed. The recorded reason is the one quoted above: keep every existing caller unchanged. Nothing on record revisited it, and the default stayed fts until the old store's memory was switched off.
In the module the default is semantic for the first time: hybrid, plus a local cross-encoder that re-ranks the admitted candidates, merged on 2026-09-28. That says which arms run when the caller doesn't choose. It says nothing about the answers they return. There is no parity number between the old store and the module, by my choice, so there is no quality comparison to make. On 2026-09-27 I set the migration's goal as moving my data safely; proving the new retrieval matched the old was out of scope. The next day I dropped the planned query sets and overlap gates for informal spot checks, on the view that weak recall of ported content could be fixed later.
Two more July claims changed outright. "The graph is the record" is no longer true: the record is a Postgres row with links, and entities exist only as mentions. The rule "the LLM identifies, deterministic code executes" governed four nightly workers (resolution, dedup, autoconfirm, profile) that the module doesn't have. Dedup now proposes pairs at cosine 0.80 or above, from at most the 500 newest memories the caller may see, and a person decides on a web screen; nothing merges on a pair's strength. That moves dedup's cost from nightly model calls to my time, and a duplicate stays until I look. I haven't measured either side. Automatic memory comes only from explicitly stated evidence, and anything empty or ambiguous becomes a no-op. There is no unconfirmed pile waiting for a 0.9 gate and a 7-day age. When automatic capture went live on 2026-09-30, the first 52 real records it accepted produced 6 memories. That is one small sample, and I'm not calling it a rate.
One dated fact about those workers is new to this site: from 2026-08-13 to 2026-09-11 they ran dark without telling anyone, because the container they ran in had no claude binary. The July essay's own failures are in the July essay. The record doesn't give the outage as the reason the workers didn't come across, so neither do I.
The July essay's "The graph never forgets" holds in the sense that mattered, after a detour: nothing in the module expires on its own, and forget runs only when asked. The disuse-prune came out on 2026-07-08. The module's spec ratified a mandatory 365-day expiry on 2026-09-09. That was corrected on 2026-09-21 to an opt-in gate that is off by default.
What the verdicts sort into
Lined up, the verdicts sort by what kind of sentence each one was. The ones still true (the derived index, the degrade contract, RRF at 60) say how a part behaves. The ones that changed (the ledger, the boot bundle, "the graph is the record," the workers, the fts default) described the store's shape on the day they were written. The one that was never quite true read a default as a design.
What the module is now
A memory is a Postgres row. It has a kind (note, fact, decision or summary), a title, a body, a generated tsvector, a confidence, an origin (told, derived or migrated), a revision number, a superseded_by_id, and an invalidation reason if it was invalidated. Links are memory_link rows, derived_from or about, written when the memory is created and never edited.
Entities exist only as mentions. Each has a global UUID, two entities can share a name, and nothing merges them. No attributes, no entity-to-entity edges. There is no memory graph in the module today.
What the module has instead is a trust model. The old ledger schema had a chat id and nothing else about who a row was for. Every memory in the module carries an explicit audience, either the whole workspace or one member, and a purpose from a closed set of four. A memory derived from others gets the intersection of their audiences and purposes, never the union, and if the intersection is empty the derivation is refused.
Retrieval runs in this order. Candidates the caller isn't eligible to see are filtered out in SQL. The lexical arm is a tsvector query that, from three terms up, drops any term appearing in 15% or more of memories, a port of the old fts.ts logic. The dense arm embeds the bare query and searches by cosine with a floor of 0.30. RRF fuses the two. The permission walk then goes through the ranked list per link, in rank order, up to 500 candidates. The re-ranker only ever sees rows that passed, and the counts in the response are taken after the walk.
The agent reaches all of this through seven memory tools: recall, read, remember, derive, correct, supersede and forget. The old surface had 37 tool definitions, 6 of them for tickets. correct and supersede both check the revision number and fail for the loser of a race. Recall is on demand only.
Why the order was the order
When I was asked what the biggest shift was, my answer was: "sequencing mostly, get the migration working first." That is my whole stated view. The record adds three decisions that pinned parts of the order.
D1 put each workspace in its own Postgres database, with a small shared control plane. I weighed SQLite and reconfirmed Postgres on 2026-09-08. Isolation by database is what carries the permission model, so once D1 was decided, permissions had to be built at the foundation.
D11 made memory the first real module because it is "the smallest domain that can prove the module contract end to end." The point was that a contract defect and a domain defect would be distinguishable. That is a reason about the framework; memory was the domain that fit it.
D9 recorded that sqlite-vec and FTS5 don't exist in Postgres, so the port needed a real retrieval rework, contained behind an adapter so ranking could be revised later. That rework was going to happen whatever anyone thought about graphs.
Those three explain why permissions and retrieval were built first. None of them mentions the graph, and nothing on record ruled it out. The public ledger has only a scope line: release one supports entities as a global identity, memory-to-entity mentions and provenance links, and explicitly not attributes, entity-to-entity edges, traversal or merge. That says what release one holds. It doesn't say why.
For why the graph waits, the record has the one sentence above and a status: on 2026-09-22 I corrected a description that called leaving out the episodic ledger, entity attributes and merge, and typed entity relations "deliberate." All three are approved to plan, as 1c, 1d and 1e. "Approved to plan" is a status, not a commitment: there is no card and no sizing for any of them, and phase two hasn't started. What did come across is narrower. Every old memory-to-entity edge became a mention. Only the memory-to-memory and entity-to-entity verbs are waiting.
A number true of a corpus
The dense floor is the same claim at small scale. The old store's vector arm got a floor of 0.76 in L2 distance, about cosine 0.71, on 2026-09-18. It was measured on conversation turns: 37 relevant answers with a maximum of 0.733, 60 nonsense queries with a minimum of 0.599. The commit said the distributions overlap. Planning assumed the floor would carry over, since the model name was the same. On paper that is a reasonable assumption.
The build stopped on it. The old Node build and the shipped Python build of the same model differ in pooling and in query and passage prefixes, so they produce different embedding spaces. The corpus changed too, from conversation turns to memory records. The conversion was invalid and so were the borrowed populations. I chose to re-derive. The new floor is cosine 0.30, measured on the shipped provider against real text from the old store, by a stated rule: the lowest whole-percent floor that blocks at least 95% of off-topic queries while keeping every memory-shaped pair.
The two numbers also sit on different definitions of ground truth. The old method used turn adjacency, and a memory corpus has no turn adjacency. The 0.30 is a rule fitted to one sample of old text, with no held-out result or precision-recall curve behind it. The plan was for the migration to confirm the 30 on the migrated corpus. I have no record that it did.
The June failure is still open
The sequence is short. In June the model didn't retrieve on its own, so I injected context at boot. On 2026-08-27 the bot gained a recall protocol, because "digests are compact summaries that drop fine detail," and it told the model to search before claiming it didn't know. On 2026-09-30 the bot's injection fetches were removed and its memory moved to the module's remember, recall and read tools.
The mitigation was retired. The June finding was not. Today the only thing standing in for injection is text. The recallatron_recall description tells the agent that dense search returns its nearest rows, not its relevant ones, that a memory is a dated claim and not current truth, and to read the provenance before the items. That wording continues the old tool's own, which said vector mode "ALWAYS returns its nearest rows, so judge relevance yourself." Client guidance on when to recall exists and is not yet published.
Nothing injects at session start. Only Claude Code has the connection at all; desktop chat, claude.ai and the phone don't. The bot's recall has been restored, and as of 2026-10-01 it hasn't been confirmed with a real Telegram turn that it saves something and then recalls it. The bot has no automatic extraction yet, so it keeps a memory only when the model decides to call remember, which depends on the same thing June depended on: the model choosing to use a tool.
The test is easy to state, and I have run it once already. In June the answer was no, and the bot was rebuilt around that answer. The injection is gone now, and a tool description and a protocol stand where it stood. What they do to the answer, I can't say yet; the test is the same one. Does a recall call happen before the agent says it has no context?