Technical workMeasured, then explained

Results before adjectives.

We publish measured work so you can see what the products do and how we think. The first set covers large-model inference with Krasis.

Large models on single-GPU systems.

These results show three models running with Krasis. They are individual demonstrations on different systems, not a direct ranking between models or graphics cards.

01Measured result

Ornith 397B

1× Nvidia RTX PRO 6000

96 GB VRAM · 200 GB DDR4

2,354tokens/secPrefill
25.72tokens/secDecode
02Measured result

Step 3.7

1× Nvidia RTX PRO 6000

Single-GPU system

5,261tokens/secPrefill
55.4tokens/secDecode
03Measured result

Qwen 3.5 122B

1× Nvidia RTX 5090

Single-GPU system

3,692tokens/secPrefill
31.8tokens/secDecode

01

Prefill

Prefill is the work of reading and processing the input before a response begins. Faster prefill matters when prompts and context are large.

02

Decode

Decode is the generation of the response itself. Its token rate is the speed a person experiences once the model starts answering.

03

Reproducibility

Model configuration, prompt length and software versions affect results. Full test details will be published alongside future engineering notes.

Unsure where the bottleneck is?

Northloom offers focused, remote reviews of self-hosted inference systems and proposed hardware designs.

Ask an inference question