Skip to content
Ultimus

04 of 04 · Center

The Ultimus Harness.

Sovereign Agentic Harness

The execution layer. Ultimus makes sense of all the context and executes. It is a local-first, model-agnostic agent harness that runs tasks, validates its own work against your test suites, and operates strictly within the policy boundaries you define. It cannot wander, it is completely controlled by the governance you design.

Everything upstream is context. Ultimus is what makes sense of it, and acts.

Ultimus is an agent harness. Local-first. Model-agnostic. Context flows in from your environment and from your feeds. You point it at a task and it runs the task. Then it checks itself.

  • It validates its own work against your test suites.
  • It operates strictly within the policy boundaries you define. Not its own. Yours. Those boundaries are the guardrails.
  • It cannot wander. It is completely controlled by the governance you design.

Sovereign means the harness stays in your hands. Swap the model underneath and the rules above it do not move.

What it costs

Measured from 61 real session logs, Ultimus on an open-weight model costs about 110× less per 1,000 output tokens than Claude Code on Opus 5. Ultimus works the way Codex and Claude Code do. It just costs a fraction as much to run. Read the benchmark, including what it does not prove.

Where it sits

The harness is the center. Internal context feeds and external feeds give it context. Context management governs what reaches it, and collects what it does.

Questions.

What does the harness do?

It makes sense of all the context and executes. Context flows in from your environment and from external feeds. Ultimus reasons over it, runs the task, and checks its own work.

What does local-first mean here?

Ultimus runs in your environment. The harness is sovereign. You keep control of where it runs.

Which models does it support?

It is model-agnostic. The harness does not tie you to one model provider.

How does Ultimus compare with Claude Code on cost?

In a benchmark of 61 real session logs, Ultimus on an open-weight model cost about 110× less per 1,000 output tokens than Claude Code on Opus 5. Cost is measured, and the energy figures are estimates. Only Claude Code was measured, not Codex.

Can it act outside its policy?

No. It cannot wander. It operates strictly within the policy boundaries you define, and it is completely controlled by the governance you design.

How does it check its own work?

It validates results against your test suites before it reports done.

Next module · Inside

Internal Context Feeds