Research

Results, models, and notes from the Belna team.

Cortex-S: a second hybrid confirms the signal

Sparse mixture-of-experts meets persistent recurrent state. A completely independent design, the same 2B-token protocol — and the same result: lower loss than the Transformer. Safety bounds included, training code included.

Read article

Active models

Everything running in production or under active investigation — switch tabs to explore.

Live in production

Arche 1.0

The personal agent that does the work — 300 clicks across 4 apps replaced by one sentence. Open Kimi K3 foundation, 1M-token context, Smart and Fast modes, wrapped in the Belna harness with approvals, a sealed vault, and a visible trace.

LiveKimi K3 open weights1M contextSmart + Fast

Read the Arche 1.0 story →

Research preview

STLM

Our 23.5M-parameter model that pins words to meaning: topology coordinates learned jointly with language modeling on WikiText-2. Lower perplexity than its baseline (80.89 vs 111.45), a 0.957 meaning correlation, and 35/35 on negation with axis guidance.

23.5M paramsPPL 80.89 vs 111.45Negation 35/35

Read the STLM story →

Research preview

Mini-SLA

One Transformer layer, looped four times around an explicit 10-slot memory with a symbolic negation router. At 11.3M params it beats its matched baseline on perplexity (63.89 vs 111.45) — and solves negation 35/35 with no help at inference.

11.3M paramsPPL 63.89 vs 111.45Autonomous negation 35/35

Read the SLA story →

Research preview

TAM v3

Our 101.8M-parameter hybrid: reduced-width attention running in parallel with a recurrent world-state, mixed by a learned gate. Lower loss and perplexity than a matched Transformer at 25M, 50M, and 100M — plus a real memory advantage at 256-token contexts.

101.8M params2B tokensNLL 2.698 · PPL 14.86

Read the 100M paper →

Research preview

Cortex-S

An independent cross-check: 8-expert sparse MoE with persistent recurrent state and full attention only every 6th layer. Same data, same hardware, same seed discipline — and it also beats the Transformer on loss, with auditable, bounded compute.

101.8M params2B tokensNLL 2.709 · PPL 15.02

Read the 100M paper →

Every paper ships its full training code inline — no hidden repositories. Replicate, remix, improve.

Home · Models · Research · Pricing