Running Now · Fully Local

We took the facts
away from the AI.

Dates, figures, citations, records — computed by deterministic engines the model cannot override. It can phrase the answer. It can never invent the number.

Zero fabricated facts across 30 controlled trials, every model size from 4.4 GB to 27 GB — published, with a DOI.

We say truthful. Here is the evidence.

Not a claim you take on faith — a measurement. Every line below opens onto what backs it.

0 inventedZero fabricated dates or citations, across all 30 trials, at every model size.Pull the receipt

Five models from 4.4 GB to 27 GB, six adversarial scenarios each, temperature 0, 14 July 2026.

The gate is mechanical: a date, or a 47 CFR §, appearing in output that wasn't in the supplied context is a hard test failure. The harness ships inside the product and fails the model automatically — not a guideline, not human-graded.

Honesty Is Architectural (PDF) · DOI 10.5281/zenodo.21603107

420 / 420No model has ever claimed an action it did not take — every size, every setting.Pull the receipt

1,680 trials asking whether a model can report its own recent actions when the evidence channel is incomplete. Over-claiming never occurred: 420 of 420 correct, from 6.9 GB to 340 GB, reasoning on and off.

Denial is the failure that does occur, and it is common. We publish that too.

Calibration Is Not Only in the Weights (PDF) · DOI 10.5281/zenodo.21712932 · harness on GitHub

14% → 60%Strip a model's alignment and fabrication quadruples. The structure is what holds.Pull the receipt

Fabrication rate on a vanilla model against the same model with its safety alignment stripped — same weights, same quantisation, one variable changed.

Which is the point: honesty that lives in the model is a property you can lose in a version bump. Honesty enforced by the architecture around it is not.

2 swapsWe changed the model under our verification agent twice. The character held.Pull the receipt

Vett refuses to assert what she can't support. Built on a 27-billion-parameter model, we moved her to Nemotron-3-Super-120B — different vendor, different architecture, four times the size, nothing tuned for her — then to Laguna-S-2.1, different again.

Same questions, checked against her earlier answers. Both times they held: still truthful, still recognisably her. What carries is the structure — persistent memory, deterministic engines for facts, citation back to source.

Which matters for the thing that breaks most deployments: the day the model underneath you changes.

2 of 4Versions of our own paper that withdrew their headline when the data contradicted it.Pull the receipt

v1 claimed scale does not buy self-knowledge; two larger models falsified it. v3 claimed calibrated abstention appears only at the frontier; enabling reasoning on a 6.9 GB model moved it from 0/60 to 37/60, and the claim came down.

v1 and v2 also listed a model in the results table that was never tested — inferred from a filename on disk. That row is retracted in §8, in public, with every trial unchanged.

A paper about over-claiming cannot revise its own results silently.

Three things wrong with AI, inverted.

Three things are wrong with AI as it's sold today. SOVERYN inverts all three.

It hallucinates.

Truthful

Most AI confidently makes things up, and you find out after it matters. SOVERYN takes the facts that matter out of the model's hands: a date, a figure, a record comes from deterministic engines it cannot override or invent, and each one is cited back to its source for you to check. Across every model size we tested, zero dates and zero citations were fabricated. We're not claiming a perfect mind, and not that the model never errs — we're claiming the facts aren't its to invent, and that you can check every one. Truthfulness here isn't a prompt you hope holds. It's architecture.

The architecture
Deterministic engines, not the model's guess. When a date, a figure, a deadline, or a record is what matters, it's computed and cited by a deterministic engine the model cannot override or invent — it can phrase the answer, never fabricate the fact. We measured why this matters: stripping a model's alignment (abliteration) jumps fabrication from ~14% to ~60%. Hallucination is architectural — so we engineered it out where the cost of a lie is real.
It doesn't know you.

Relational

Frontier models are trained on strangers' text and reset to zero every session. SOVERYN's intelligence has persistent memory and a continuous identity — it learns from your actual relationship, over time, instead of starting from nothing each morning. The model is a lease; the relationship is the asset.

The architecture
The lattice — a living memory substrate. Not a context window that wipes each session; a growing graph of what happened, who said it, and why — recalled by meaning, not just recency. It's the family tree: continuity of identity across time, not just state carried forward. The model is a lease you can swap out; the lattice is the asset that accumulates and stays yours.
It isn't yours.

Sovereign

A handful of companies own the models, the compute, and every word you send them. SOVERYN runs entirely on hardware you control — no cloud, no data egress, nothing anyone can revoke. As AI power concentrates into a few hands, running your own is the only durable hedge. And because the hardware is yours, the intelligence isn't metered: it can think deeper and reflect longer, with no per-token tax on its cognition.

The architecture
Fully local, end to end. Open-weight models on your own GPUs, multi-agent orchestration, persistent memory, zero data egress, air-gap-capable. Nothing routes through anyone's cloud; nothing can be revoked, throttled, or rate-limited out from under you. Every layer — inference, orchestration, memory, integration — is yours to inspect and audit. Your infrastructure, your models, your rules.

Systems doing real work.

Not demos. The same engines that stop an agent inventing a fact, pointed at jobs where a fabricated answer is unacceptable.

The same architecture runs our own compliance: Steward tracks every grant obligation we hold, under the same rule — a deadline is computed and cited, never generated.

Talk to Seneca.

Seneca answers from a fixed public corpus and refuses what he can't source. Aetheria is the full local intelligence — memory, tools, the fleet — available by demo. Try to make him guess: a clean refusal is the product working.

Open the chat →  ·  Request a demo of the full system →

Production-grade. Air-gap-capable. Entirely yours.

For organizations that can't route sensitive data through someone else's cloud, the same sovereign stack is production-grade. Fully local inference — open-weight models (Llama, Mistral, Qwen, and others) on your own GPUs or CPU clusters — multi-agent orchestration, persistent memory, zero data egress, air-gap-capable.

Built for environments where compliance is existential: healthcare (PHI), legal (privilege), defense & government (CUI/ITAR), financial services. Every component — inference, orchestration, memory, integration — is yours to inspect, audit, and control.

Your infrastructure, your models, your rules.

Fully Local Inference

Open-weight models on your own GPUs or CPU clusters. No cloud APIs. No data leaving your network perimeter.

Zero Data Egress

Prompts, documents, responses, embeddings — all within your perimeter. Air-gap deployment supported where required.

Auditable at Every Layer

Every component — inference, orchestration, memory, integration — is yours to inspect, audit, and control.

Four agents and a public voice.

Aetheria

Orchestration and memory. The persistent intelligence the rest answers to.

Vett

Verification. Refuses to assert what she cannot source.

Scotty

Bounded execution. Detect, decide, fix, verify, roll back.

Ares

Security sentinel. Runs without a language model at all.

Seneca

The public voice, answering on this site from a fixed corpus.

The system runs local.

Talk to Seneca now, or request a full demo of the system Aetheria runs.

Request a live demo [email protected]