Nano
Released349M Dense
Trained on 100B tokens.
Built for extreme efficiency: 1.58-bit weights, Mamba-3 mixers and MLA attention.
A chat page, a terminal chat and an OpenAI-compatible API, on NVIDIA GPUs or the CPU.
349M Dense
Trained on 100B tokens.
2.0B Dense
The first full production scale.
A big step up in capability, at a size that stays practical.
The flagship: Ivonar at full scale, built for the most demanding work.
one square = 250M parameters
Against five published models from 390M to 700M, each scored as its authors report it.
Weights and kernels under Apache 2.0.
| Model | Ivonar Nano | TriLM 390M | FloatLM 390M | QuantLM 4-bit | BitNet b1.58 | LLaMA 700M |
|---|---|---|---|---|---|---|
| Parameters | 349M | 390M | 390M | 390M | 700M | 700M |
| Training tokens | 100B | 300B | 300B | 300B | 100B | 100B |
| Weights | Ternary | Ternary | FP16 | 4-bit | Ternary | FP16 |
| ARC-Easy | 56.0 | 48.6 | 51.0 | 49.6 | 51.8 | 54.7 |
| ARC-Challenge | 24.8 | 21.2 | 21.3 | 21.3 | 21.4 | 23.0 |
| HellaSwag | 32.8 | 32.0 | 35.7 | 35.1 | 35.1 | 37.0 |
| PIQA | 64.1 | 65.0 | 68.4 | 68.3 | 68.1 | 68.9 |
| Winogrande | 51.9 | 52.2 | 51.8 | 53.7 | 55.2 | 54.8 |
| BoolQ | 60.8 | 55.1 | 54.7 | 50.8 | 58.2 | 60.0 |
| Average (6) | 48.4 | 45.7 | 47.2 | 46.5 | 48.3 | 49.7 |
About 1,430 tokens per second on an RTX 4060 Ti, from 95 MB of weights.
Our own measurement, one request at a time; BitNet weights from the 1bitLLM reproduction.
Nano decodes 950 tokens/s without CUDA graphs, and about 130 on a Ryzen 9 3900X CPU (24 without its kernels).
TriLM and BitNet run on TriRun, the fastest ternary GPU kernel we could run; QuantLM on Marlin; FloatLM in FP16.
Founded and led by Luis Oezdem, from architecture to training infrastructure and data pipelines.
Compute and ecosystem partners
High-performance training compute
Research and compute access
Early-stage growth programme
Ternary weights and 8-bit activations from the first training step, not quantised afterwards.
Mamba-3 mixers for long context, MLA attention for precise recall: 18 and 6 blocks in Nano.
Weights packed at two bits, run by dedicated CUDA and CPU kernels.
Language and code capability without scaling up dense parameters.
Ternary weights and Mamba-3 mixers keep memory and compute low.
API access, subscriptions, and enterprise licences for teams that keep their data in house.
Questions about the architecture, compute partnerships, or running Ivonar Nano can be sent directly.