Verzeta

Conversation memory and compaction

Every model has a fixed amount of text it can pay attention to at once: its context window. A long conversation eventually grows past that limit, and something has to give. This page explains what Verzeta Studio does when that happens, how to read the memory gauge in the message bar, and which controls you have.

The short version

How this differs from "compact" in Claude Code or Codex

If you've used coding assistants, you may know /compact as something that replaces your visible conversation with a summary: the old messages disappear from the transcript. Verzeta's compaction never touches what you see. The full transcript stays intact and scrollable forever; compaction only changes what is packaged and sent to the model behind the scenes.

The memory gauge in the message bar

Next to the Send button you'll find a small two-ring gauge. The two rings show the two things that can trigger an automatic memory refresh, and whichever ring fills first is what triggers it:

Each ring is colour-coded by how full it is: grey up to 35%, green to 65%, amber to 85%, orange to 95%, red beyond. Hover the gauge for both numbers in plain words, for example "Outer ring, Memory refresh: ~3 turns until the next refresh (5/8). Inner ring, Context window: 88% full." If the per-conversation refresh schedule is turned off, the outer ring is idle and only context pressure drives a refresh.

The gauge starts filling after the first reply in the conversation, and clicking it compacts right now. That is identical to typing /compact: never destructive, and repeatable as often as you like.

When auto-compaction kicks in

Three independent triggers can start a summary, and all three are careful never to fire repeatedly for the same content:

  1. Early warning (the main one). When the inner (context) ring reaches the amber zone (~70% full), a summary is generated before anything stops fitting, so the model never hits a point where older messages silently vanish without a summary already covering them.
  2. Conversation rhythm (configurable). Long team exchanges drift even when everything still fits, and agents slowly lose the thread. The "Refresh summary every N agent replies" setting (default 20, set 0 to disable) refreshes the summary on that rhythm so the team's working memory stays sharp. Set it higher for slow, deliberate conversations or lower for fast-moving ones.
  3. Backstop. If messages are already being left out of the model's view and no up-to-date summary covers them, one is generated immediately.

In every case the summary is generated in the background, so the conversation never waits for it. The reply you're watching streams normally; from the next message onward the model receives the summary. The newest ~15 messages are always sent verbatim, never summarised. The summary is added alongside the messages that still fit. It never replaces messages that would otherwise be sent.

The summary is rebuilt from scratch if the team roster changes or you move the conversation to a different project (so it never describes a stale cast of characters). Running tasks and background sub-agents are safe across compaction: their state lives outside the message history, and the summary explicitly preserves assignments, task status, and pending sub-agent runs.

Your controls

Control Where What it does
Dynamic compaction toggle Chat Settings Default on. Turn it off and Verzeta falls back to plain truncation: when the window fills, the oldest messages are left out of what the model sees (they stay visible to you).
Refresh summary every N agent replies Chat Settings Default 20. The conversation-rhythm trigger above. 0 = refresh only under context pressure.
Memory gauge / /compact Message bar / typed command Summarise now, on demand. Non-destructive, repeatable.
/flashmemory Typed command Destructive. Permanently deletes every message, the summary, and canvas state of the current conversation: a true fresh start. The conversation itself (title, members, settings) survives. You must confirm by typing /flashmemory confirm.
Context Window Chat Settings The size of the model's window Verzeta budgets against. Bigger = more conversation fits before compaction is needed, if your hardware and model support it.

Tips for long group chats

Embeddings and retrieval (RAG)

Separately from compaction, the app can pull relevant earlier messages and documents back into a turn even after they've scrolled out of the live window. This is retrieval (often called RAG), and it's off by default.

Retrieval runs in the background and never blocks your messages. If the embedder is slow or unreachable, the turn still sends, without the extra context.

Agent memory (saving facts to remember)

Compaction and retrieval are about this conversation's history. Agent memory is different: it lets an agent deliberately save a durable fact and recall it later, including in other conversations.

Team memory (shared across a project)

Inside a project or organization, the team can build a shared memory that any chat in that project can draw on.

What's next

Verzeta Studio guides