Ornith 397B
1× Nvidia RTX PRO 6000
96 GB VRAM · 200 GB DDR4
Technical workMeasured, then explained
We publish measured work so you can see what the products do and how we think. The first set covers large-model inference with Krasis.
Krasis / Benchmarks
These results show three models running with Krasis. They are individual demonstrations on different systems, not a direct ranking between models or graphics cards.
1× Nvidia RTX PRO 6000
96 GB VRAM · 200 GB DDR4
1× Nvidia RTX PRO 6000
Single-GPU system
1× Nvidia RTX 5090
Single-GPU system
Reading the results
01
Prefill is the work of reading and processing the input before a response begins. Faster prefill matters when prompts and context are large.
02
Decode is the generation of the response itself. Its token rate is the speed a person experiences once the model starts answering.
03
Model configuration, prompt length and software versions affect results. Full test details will be published alongside future engineering notes.
Your system
Northloom offers focused, remote reviews of self-hosted inference systems and proposed hardware designs.
Ask an inference question