Ivonar Native ternary LLM family.

Built for extreme efficiency: 1.58-bit weights, Mamba-3 mixers and MLA attention.

48.4average of six benchmarks
8×smaller than FloatLM 390M
7.5×faster than its own FP16 path

Run Ivonar Nano on your own computer

A chat page, a terminal chat and an OpenAI-compatible API, on NVIDIA GPUs or the CPU.

curl -fsSL https://raw.githubusercontent.com/LuisOezdem/ivonar-inference/main/install.sh | sh

Four scales, one technical direction

Nano

Released

349M Dense

Trained on 100B tokens.

Mini

In planning

2.0B Dense

The first full production scale.

Medium

Coming later

A big step up in capability, at a size that stays practical.

High

Coming later

The flagship: Ivonar at full scale, built for the most demanding work.

one square = 250M parameters

Ivonar Nano, and how it scores

Against five published models from 390M to 700M, each scored as its authors report it.

Weights and kernels under Apache 2.0.

Zero-shot benchmarks6 models · 6 tasks
ModelIvonar NanoTriLM 390MFloatLM 390MQuantLM 4-bitBitNet b1.58LLaMA 700M
Parameters349M390M390M390M700M700M
Training tokens100B300B300B300B100B100B
WeightsTernaryTernaryFP164-bitTernaryFP16
ARC-Easy56.048.651.049.651.854.7
ARC-Challenge24.821.221.321.321.423.0
HellaSwag32.832.035.735.135.137.0
PIQA64.165.068.468.368.168.9
Winogrande51.952.251.853.755.254.8
BoolQ60.855.154.750.858.260.0
Average (6)48.445.747.246.548.349.7

Seven times faster, eight times smaller

About 1,430 tokens per second on an RTX 4060 Ti, from 95 MB of weights.

Inference speed, RTX 4060 Titokens / s
Ivonar Nano, native kernels1,430
TriLM 390M, TriRun250
QuantLM 4-bit, Marlin233
BitNet b1.58 700M, TriRun207
Ivonar Nano, FP16 path190
FloatLM 390M, FP16183

Our own measurement, one request at a time; BitNet weights from the 1bitLLM reproduction.

Nano decodes 950 tokens/s without CUDA graphs, and about 130 on a Ryzen 9 3900X CPU (24 without its kernels).

TriLM and BitNet run on TriRun, the fastest ternary GPU kernel we could run; QuantLM on Marlin; FloatLM in FP16.

Independent engineering, backed by Germany's AI ecosystem

Founded and led by Luis Oezdem, from architecture to training infrastructure and data pipelines.

Compute and ecosystem partners

hessian.AI

High-performance training compute

TU Darmstadt

Research and compute access

Berlin Startup Support

Early-stage growth programme

Real BitLinear training, hybrid sequence mixing

Weight footprintbits per parameter
FP1616.0
INT88.0
W1.581.58

W1.58A8 BitLinear Training

Ternary weights and 8-bit activations from the first training step, not quantised afterwards.

Mamba-3 and MLA Hybrid

Mamba-3 mixers for long context, MLA attention for precise recall: 18 and 6 blocks in Nano.

Packed 2-bit Inference

Weights packed at two bits, run by dedicated CUDA and CPU kernels.

A European ternary model, built to be used

Quality

Competitive at a tenth of the budget and a seventh of the size

Language and code capability without scaling up dense parameters.

Efficiency

Minimal memory footprint

Ternary weights and Mamba-3 mixers keep memory and compute low.

Business

Built for everyone

API access, subscriptions, and enterprise licences for teams that keep their data in house.

Talk to the team building it

Questions about the architecture, compute partnerships, or running Ivonar Nano can be sent directly.