Research · Arche 1.0

September 10, 2026 · 6 min read

Frontier intelligence you can hand real work.

Most AI answers questions about work. Arche 1.0 does the work — 300 clicks across 4 apps replaced by one sentence. It runs world-class open weights inside a harness built for privacy and safety, and it scores shoulder-to-shoulder with the biggest closed models on the benchmarks that measure real doing.

● Live in productionKimi K3 open weights1M-token contextSmart + Fast modes

Why this is a big deal

4 × #1

Best score on 4 agentic benchmarks — BrowseComp, SWE-Marathon, ProgramBench and τ³-Banking — against GPT-5.6 Sol, Claude Opus 5, Claude Fable 5 and Claude Opus 4.8.

1M tokens

A full project in one context window. Codebases, inboxes, research dossiers — Arche reads it all at once and acts on it, instead of forgetting page two.

Zero keys

The agent never holds your secrets. The sealed vault masks everything to sec_•••• — even from the model itself. Nothing sensitive acts without your yes.

Benchmarks: Arche 1.0 vs the frontier

Published technical results — Arche 1.0 (Kimi K3 at max reasoning effort) against current frontier models. Bold marks the best score per row. Arche takes the top spot on BrowseComp, SWE-Marathon, ProgramBench and τ³-Banking, and sits within a point or two of the leader almost everywhere else.

BenchmarkArche 1.0GPT-5.6 SolClaude Opus 5Claude Fable 5Claude Opus 4.8
GPQA Diamond93.594.193.892.691.0
Terminal-Bench 2.188.388.887.588.084.6
BrowseComp91.290.489.188.084.3
OSWorld-Verified84.883.084.285.083.4
SWE-Marathon42.039.041.035.040.0
FrontierSWE81.271.382.486.666.7
DeepSWE67.573.068.070.059.0
ProgramBench77.877.675.276.871.9
MCP-Atlas84.283.684.084.783.6
HLE (with tools)56.058.060.563.057.9
τ³-Banking33.433.030.226.827.6

Source: Moonshot AI Kimi K3 technical results; competitor scores as reported in the same release (each evaluated in its own vendor harness). Arche 1.0 serves these weights through the Belna privacy-first harness.

The harness: why Arche is more than its weights

A brilliant model with no hands is just a chatbot. The Belna agentic harness is what turns intelligence into finished work — safely:

✅ Approvals by default

Sensitive actions — apps, browsing, external writes — pause for your explicit yes, or an always-allow rule you can revoke anytime under Vault → Approved.

🔒 Sealed vault

Secrets are encrypted at rest and masked everywhere: chat, traces, model context. The agent only ever sees references like sec_••••. Reveal is for your eyes alone, on your device.

🧪 Sandboxed doing

Code, browser and computer use run contained with an allowlisted network. External integrations are read-only via your own per-request token — never stored, never logged.

👁 Visible trace

Every step lands in the Trace tab. Belna never pretends to browse, run code or read email unless the trace shows it. Check the work, don't just trust it.

Two speedsOne model, your choice
Smart mode
Deepest reasoning — research, builds, reviews
Fast mode
Snappy everyday answers at lower cost

Still open at the core

Arche 1.0 is built on Kimi K3 by Moonshot AI — fully open-source, open weights. Anyone can inspect the foundation we're standing on; there is no black-box base model. What Belna adds is everything around it: the agentic harness, the safety layer, the memory and vault, the Swedish-hosted product that keeps your data yours.

Openness is the point. Frontier capability shouldn't require handing your life to a closed box — and now, on the benchmarks that matter for real work, it doesn't have to.

Try Arche 1.0 free

Free starts with 20 credits. Pro $30/mo → 60 credits monthly + $50 gift card. Max $50/mo → 100 credits monthly + $100 gift card.

Get startedSee pricing

Home · Research · Pricing