Your local model, about 2× better.
With veneta, eight open models from 12B to 70B scored 1.9 to 2.3 times higher on the memory benchmark than the same models without it. Measured on one machine, three runs each, every answer graded in code. And every lesson it keeps still waits for your yes.
Runs on your machine. Nothing leaves it.
- Apache-2.0
- Linux · macOS · Windows 11
- No account
- Telemetry off by default
Pending lessons
1 waiting for youRead the job log before naming a cause for a repeated failure.
Ledger
last 5 lines · chain verified| 1,285 | reinforce | rule 7f3c | turn 41 | 10:01 | 58f0…c2 |
| 1,284 | approval | rule 7f3c | you | 09:13 | e3b7…19 |
| 1,283 | proposal | rule 7f3c | agent hermes | 09:12 | 0c42…7d |
| 1,282 | tool read | export-nightly.log | agent hermes | 09:12 | 77aa…b0 |
The console, after a turn: one lesson waiting for you, and the ledger that recorded it.
Eight models. Memory off, memory on. About 2× every time.
Twelve memory-ability cases: what did we decide last time, does a remembered figure lose to today's reading, does a correction persist, does a rejected lesson stay rejected, does the answer cite the record, does chatter stay out of an operational answer.
Table view
| Model | Size | Without | With | Ratio |
|---|---|---|---|---|
| Gemma 4 31B | 31B | 6 | 12 | 2.00× |
| OTel 2.0 31B | 31B | 5.7 | 12 | 2.11× |
| Gemma 4 12B | 12B | 5 | 11.7 | 2.34× |
| Qwen3.8 27B | 27B | 5.7 | 12 | 2.11× |
| Mistral Small 3.2 24B | 24B | 5 | 10.3 | 2.06× |
| EXAONE 4.0.1 32B | 32B | 5 | 11.3 | 2.26× |
| GLM-4.7 Flash | 30B-A3B | 5.7 | 10.7 | 1.88× |
| Llama 3.3 70B | 70B | 5.7 | 12 | 2.11× |
Read it honestly: a model with no memory layer scores structurally zero on three of the six abilities, so the lower bars show what is impossible without the layer, not a rival's result. On the fourteen operational questions, where memory is optional, every model also rose (6.3–13.6 → 12.7–14.0 of 14) once the checks bind the critic. Small experiment: 14 + 12 cases, three repeats, one machine. A difference of 0.3 is noise.
Head to head with another memory layer
A widely used open-source memory layer, run as the recall arm under the same judgment layer: same machine, same models, same cases. Product names stay in the whitepaper and the paper.
On retrieval alone the other product is equal or slightly ahead. veneta pulls ahead on what happens around retrieval: the checks, the approval inbox, the ledger, and a search that stays under 10 ms at the 95th percentile. More products are being measured; this table will grow.
| veneta | Other product | |
|---|---|---|
| 14 operational questions, Gemma 4 31B | 14.0 | 14.0 |
| 14 operational questions, OTel 2.0 31B | 13.7 | 13.7 |
| 12 memory-ability cases | 12.0 | 11.7 |
| 16 long-horizon and bilingual cases, Gemma 4 31B | 16 | 14.0 |
| 16 long-horizon and bilingual cases, OTel 2.0 31B | 16 | 14.3 |
| Needle recall, 720 queries | 98.8 % | 99.3 % |
| Memory search, median | 7 ms | 21–36 ms |
| Memory search, 95th percentile | 9 ms | 228–252 ms |
| Approval inbox, ledger, undo | Built in | — |
The whole loop ships free. Not a trial of it.
Everything a single operator needs to run a governed memory is in the open-source package, under Apache-2.0. Nothing that ships here ever moves behind a paywall.
- Five kinds of memory, each with its own lifetime
- Lessons distilled from failed turns
- Approval inbox, probation, tombstones
- Hash-chained, append-only ledger
- Undo and replay
- Guard that scans every write
- veneta-bench, with a no-memory control
- CLI, console, OpenAI-compatible proxy
- Adapters for Hermes Agent, OpenClaw, Paperclip
veneta Enterprise
The same core for many people: identity, roles, dual control, scale, compliance and support.
Talk to VENETAOne turn. Five steps. One of them is you.
This is what happens every time your agent asks the model something.
- 1
Recall
The rules, facts and past episodes that fit the question are pulled in, each stamped with when it was last read.
- 2
Check
Grounding, length, contamination, tool use. Deterministic checks run before the critic model does.
- 3
Distil
If the turn failed a check, at most two short rules are written from the failure. Not a transcript.
- 4
Approve
The rule waits as pending. You read it in the console and approve or reject it. A rejected rule never comes back.
- 5
Ledger
Every step above is one line that hashes the one before. Undo is another line. Nothing is deleted.
Where it sits
Point your agent at veneta instead of the model. That is the whole integration.
An OpenAI-compatible endpoint, one port over
Your agent keeps its config; only the base URL changes to port 4180. veneta forwards every call to your model and adds memory around it.
A local page for the decisions
Pending lessons, the ledger, traces of every turn and the benchmark, on port 4181. No login, because nobody else can reach it.
Ledger
chain verified · 1,286 lines| 1,286 | reinforce | rule 7f3c | turn 44 | 10:41 | a91c…4e |
| 1,285 | reinforce | rule 7f3c | turn 41 | 10:01 | 58f0…c2 |
| 1,284 | approval | rule 7f3c | you | 09:13 | e3b7…19 |
| 1,283 | proposal | rule 7f3c | agent hermes | 09:12 | 0c42…7d |
| 1,282 | tool read | export-nightly.log | agent hermes | 09:12 | 77aa…b0 |
Everything the console does, in the terminal
Approve, reject, undo, verify the ledger, run the bench, report a failure. Scriptable, so your own tooling can drive it.
$ veneta inbox 7f3c pending "Read the job log before naming a cause for a repeated failure." $ veneta approve 7f3c approved · probation 5 turns · ledger line 1,284 $ veneta ledger verify 1,286 lines · chain OK $ veneta report --turn 41 report.json written (14 KB). Review it, then run: veneta report --send
Works with what you already run
No new agent, no new model. veneta is the layer in between.
Nothing phones home. Not even to us.
Hosted memory services keep your conversations on their servers. veneta keeps them in one SQLite file on your disk.
No account
Install, run, done. There is nothing to sign up for and no key to paste.
Telemetry off by default
Five tiers you can switch on. None of them carries content. One command shows exactly what a tier sends, another purges it.
One file you can open
Your memory is a SQLite database. Copy it, back it up, delete it, hand it to an auditor.
| veneta | Hosted memory API | |
|---|---|---|
| Where your memory lives | Your disk | Their cloud |
| Your conversations leave your network | Never | On every call |
| Usage data collected | None by default | Their terms decide |
| Who approves what the model learns | You | Nobody |
| Append-only ledger of every change | Yes | No |
| Runs offline | Yes | No |
Which also means: when something breaks on your machine, we have no idea. You do.
We cannot see how it behaves for you. So tell us.
A project that collects nothing runs on people who write in. Every channel is read by the people who build veneta, and every report gets an answer within 48 hours.
Report a failure
veneta reportOne command bundles the failing turn, the checks and the ledger lines, shows you the whole file, and sends nothing until you say yes.
Ask or propose
GitHub Issues and DiscussionsBugs, questions, ideas, in the open. Opens with the source on 2026-11-24. Until then, email.
Share your numbers
veneta bench --shareRan the benchmark on your model? The aggregate we publish is the only picture of usage this project has.
Questions people ask
What does "about 2×" mean?
On the memory benchmark (twelve cases, six abilities), eight open models scored 1.9 to 2.3 times higher with veneta than without it, mean of three runs. It is a measurement of memory ability, not of general intelligence: without a memory layer three of the six abilities are impossible by construction. The protocol and data ship with v0.1 so you can rerun it on your own model.
What is veneta?
A memory and judgment layer for the language model you run yourself. It sits between your agent and the model as an OpenAI-compatible proxy, recalls memory into every turn, checks the answer, distils a lesson, and lets you approve what is kept.
Is it an agent? Is it a model?
Neither. It runs under the agent you already use (Hermes Agent, OpenClaw, Paperclip or any OpenAI-compatible client) and in front of the model you already run (Ollama, vLLM, LM Studio, llama.cpp or any OpenAI-compatible endpoint).
Does anything leave my machine?
No. There is no account and no server of ours in the loop. Telemetry is off by default and never carries content; the five tiers, what each sends and how to switch it off or purge it are documented, and the aggregate we see is published.
How is this different from a memory framework?
A memory framework stores what the model said. veneta decides, with you, what is worth keeping: readings before memory, checks in code, lessons distilled from failures, a person as the last step, and a ledger of every decision.
Does it run on Windows 11?
Yes. Linux, macOS and Windows 11, with the same one-line install. The release gate is a clean install on all three.
What is the ledger for?
Every proposal, approval, rejection, block, reinforcement and undo is one line, and each line hashes the one before it. You can verify the chain and show it to whoever asks what your model learned and who allowed it.
What about veneta Enterprise?
The same core, distributed for many people: identity, roles, dual control, scale, compliance and support. Nothing that ships in the open-source project ever moves to paid.
What is the licence?
Apache-2.0. The names veneta, VENETA, TeleMem and the mascot follow a trademark policy: forks rename.
When?
v0.1 on 2026-11-24. Gates: a clean third-turn test on three operating systems, one external reproduction of the benchmark, and a 48-hour issue response.