Skip to content

Your local model, about 2× better.

With veneta, eight open models from 12B to 70B scored 1.9 to 2.3 times higher on the memory benchmark than the same models without it. Measured on one machine, three runs each, every answer graded in code. And every lesson it keeps still waits for your yes.

Runs on your machine. Nothing leaves it.

Install with one line
curl -fsSL https://project.veneta.ai/install.sh | sh

The package goes live on 2026-11-24. The command will not change.

  • Apache-2.0
  • Linux · macOS · Windows 11
  • No account
  • Telemetry off by default

The console, after a turn: one lesson waiting for you, and the ledger that recorded it.

Measured

Eight models. Memory off, memory on. About 2× every time.

Twelve memory-ability cases: what did we decide last time, does a remembered figure lose to today's reading, does a correction persist, does a rejected lesson stay rejected, does the answer cite the record, does chatter stay out of an operational answer.

1.9–2.3×on all eight models
4 of 8at 12 out of 12
0regressions where memory can hurt
8 of 8up on 14 operational questions too
036912Gemma 4 31B (31B): 6 of 12 without veneta, 12 with veneta (2.0×)12Gemma 431B · 2.0×OTel 2.0 31B (31B): 5.7 of 12 without veneta, 12 with veneta (2.1×)12OTel 2.031B · 2.1×Gemma 4 12B (12B): 5 of 12 without veneta, 11.7 with veneta (2.3×)11.7Gemma 412B · 2.3×Qwen3.8 27B (27B): 5.7 of 12 without veneta, 12 with veneta (2.1×)12Qwen3.827B · 2.1×Mistral Small 3.2 24B (24B): 5 of 12 without veneta, 10.3 with veneta (2.1×)10.3Mistral Small24B · 2.1×EXAONE 4.0.1 32B (32B): 5 of 12 without veneta, 11.3 with veneta (2.3×)11.3EXAONE 432B · 2.3×GLM-4.7 Flash (30B-A3B): 5.7 of 12 without veneta, 10.7 with veneta (1.9×)10.7GLM-4.730B-A3B · 1.9×Llama 3.3 70B (70B): 5.7 of 12 without veneta, 12 with veneta (2.1×)12Llama 3.370B · 2.1×
Figure 1. Twelve memory-ability cases, eight open models, the memory layer switched off and on over the same judgment layer. Mean of three runs, dense recall, one machine. Source: campaign 3, 28–30 September 2026; the protocol and the data ship with v0.1.
Table view
ModelSizeWithoutWithRatio
Gemma 4 31B31B6122.00×
OTel 2.0 31B31B5.7122.11×
Gemma 4 12B12B511.72.34×
Qwen3.8 27B27B5.7122.11×
Mistral Small 3.2 24B24B510.32.06×
EXAONE 4.0.1 32B32B511.32.26×
GLM-4.7 Flash30B-A3B5.710.71.88×
Llama 3.3 70B70B5.7122.11×

Read it honestly: a model with no memory layer scores structurally zero on three of the six abilities, so the lower bars show what is impossible without the layer, not a rival's result. On the fourteen operational questions, where memory is optional, every model also rose (6.3–13.6 → 12.7–14.0 of 14) once the checks bind the critic. Small experiment: 14 + 12 cases, three repeats, one machine. A difference of 0.3 is noise.

Head to head with another memory layer

A widely used open-source memory layer, run as the recall arm under the same judgment layer: same machine, same models, same cases. Product names stay in the whitepaper and the paper.

On retrieval alone the other product is equal or slightly ahead. veneta pulls ahead on what happens around retrieval: the checks, the approval inbox, the ledger, and a search that stays under 10 ms at the 95th percentile. More products are being measured; this table will grow.

venetaOther product
14 operational questions, Gemma 4 31B14.014.0
14 operational questions, OTel 2.0 31B13.713.7
12 memory-ability cases12.011.7
16 long-horizon and bilingual cases, Gemma 4 31B1614.0
16 long-horizon and bilingual cases, OTel 2.0 31B1614.3
Needle recall, 720 queries98.8 %99.3 %
Memory search, median7 ms21–36 ms
Memory search, 95th percentile9 ms228–252 ms
Approval inbox, ledger, undoBuilt in—
Open source

The whole loop ships free. Not a trial of it.

Everything a single operator needs to run a governed memory is in the open-source package, under Apache-2.0. Nothing that ships here ever moves behind a paywall.

  • Five kinds of memory, each with its own lifetime
  • Lessons distilled from failed turns
  • Approval inbox, probation, tombstones
  • Hash-chained, append-only ledger
  • Undo and replay
  • Guard that scans every write
  • veneta-bench, with a no-memory control
  • CLI, console, OpenAI-compatible proxy
  • Adapters for Hermes Agent, OpenClaw, Paperclip
For teams

veneta Enterprise

The same core for many people: identity, roles, dual control, scale, compliance and support.

Talk to VENETA

One turn. Five steps. One of them is you.

This is what happens every time your agent asks the model something.

  1. 1

    Recall

    The rules, facts and past episodes that fit the question are pulled in, each stamped with when it was last read.

  2. 2

    Check

    Grounding, length, contamination, tool use. Deterministic checks run before the critic model does.

  3. 3

    Distil

    If the turn failed a check, at most two short rules are written from the failure. Not a transcript.

  4. 4

    Approve

    The rule waits as pending. You read it in the console and approve or reject it. A rejected rule never comes back.

  5. 5

    Ledger

    Every step above is one line that hashes the one before. Undo is another line. Nothing is deleted.

Where it sits

Point your agent at veneta instead of the model. That is the whole integration.

An OpenAI-compatible endpoint, one port over

Your agent keeps its config; only the base URL changes to port 4180. veneta forwards every call to your model and adds memory around it.

Your agentHermes · OpenClaw · Paperclip · any
OPENAI_BASE_URL →
venetalocalhost:4180memory.sqlite · ledger
→ unchanged
Your modelOllama · vLLM · LM Studio · llama.cpp

Works with what you already run

No new agent, no new model. veneta is the layer in between.

Agents
Hermes AgentAdapter included
OpenClawAdapter included
PaperclipAdapter included
Any OpenAI-compatible agentChange the base URL
Model servers
OllamaDefault target
vLLMTested with tool calling
LM StudioOpenAI-compatible server
llama.cppOpenAI-compatible server
Any OpenAI-compatible endpointPoint and go
Your data

Nothing phones home. Not even to us.

Hosted memory services keep your conversations on their servers. veneta keeps them in one SQLite file on your disk.

No account

Install, run, done. There is nothing to sign up for and no key to paste.

Telemetry off by default

Five tiers you can switch on. None of them carries content. One command shows exactly what a tier sends, another purges it.

One file you can open

Your memory is a SQLite database. Copy it, back it up, delete it, hand it to an auditor.

venetaHosted memory API
Where your memory livesYour diskTheir cloud
Your conversations leave your networkNeverOn every call
Usage data collectedNone by defaultTheir terms decide
Who approves what the model learnsYouNobody
Append-only ledger of every changeYesNo
Runs offlineYesNo

Which also means: when something breaks on your machine, we have no idea. You do.

We cannot see how it behaves for you. So tell us.

A project that collects nothing runs on people who write in. Every channel is read by the people who build veneta, and every report gets an answer within 48 hours.

Report a failure

veneta report

One command bundles the failing turn, the checks and the ledger lines, shows you the whole file, and sends nothing until you say yes.

Ask or propose

GitHub Issues and Discussions

Bugs, questions, ideas, in the open. Opens with the source on 2026-11-24. Until then, email.

Share your numbers

veneta bench --share

Ran the benchmark on your model? The aggregate we publish is the only picture of usage this project has.

Write to a person

[email protected]

A person reads it, and answers.

More about the community →

Questions people ask

What does "about 2×" mean?

On the memory benchmark (twelve cases, six abilities), eight open models scored 1.9 to 2.3 times higher with veneta than without it, mean of three runs. It is a measurement of memory ability, not of general intelligence: without a memory layer three of the six abilities are impossible by construction. The protocol and data ship with v0.1 so you can rerun it on your own model.

What is veneta?

A memory and judgment layer for the language model you run yourself. It sits between your agent and the model as an OpenAI-compatible proxy, recalls memory into every turn, checks the answer, distils a lesson, and lets you approve what is kept.

Is it an agent? Is it a model?

Neither. It runs under the agent you already use (Hermes Agent, OpenClaw, Paperclip or any OpenAI-compatible client) and in front of the model you already run (Ollama, vLLM, LM Studio, llama.cpp or any OpenAI-compatible endpoint).

Does anything leave my machine?

No. There is no account and no server of ours in the loop. Telemetry is off by default and never carries content; the five tiers, what each sends and how to switch it off or purge it are documented, and the aggregate we see is published.

How is this different from a memory framework?

A memory framework stores what the model said. veneta decides, with you, what is worth keeping: readings before memory, checks in code, lessons distilled from failures, a person as the last step, and a ledger of every decision.

Does it run on Windows 11?

Yes. Linux, macOS and Windows 11, with the same one-line install. The release gate is a clean install on all three.

What is the ledger for?

Every proposal, approval, rejection, block, reinforcement and undo is one line, and each line hashes the one before it. You can verify the chain and show it to whoever asks what your model learned and who allowed it.

What about veneta Enterprise?

The same core, distributed for many people: identity, roles, dual control, scale, compliance and support. Nothing that ships in the open-source project ever moves to paid.

What is the licence?

Apache-2.0. The names veneta, VENETA, TeleMem and the mascot follow a trademark policy: forks rename.

When?

v0.1 on 2026-11-24. Gates: a clean third-turn test on three operating systems, one external reproduction of the benchmark, and a 48-hour issue response.

About 2× on memory. And you keep the veto.

Free and open source. v0.1 on 2026-11-24.

Get started
project veneta

Governed memory for the model you run yourself. Free and open source for Linux, macOS and Windows 11.

© 2026 VENETA Inc. Apache-2.0. This site sets no cookies and runs no analytics.Made in Seoul. The veneta name and mascot follow a trademark policy.