Runtime is almost here: join TypeSafe AI, Cognition, DoorDash and more in SF. Limited seats left. Register now

K2 Horizon 0.9B

K2 Horizon 0.9B is the compact dense member of IFM's six-model K2 Horizon fleet: a 0.9B-class decoder-only reasoning model with a 128K-token context window, selectable reasoning effort, and tool calling, released as open weights under Apache 2.0 alongside its intermediate training checkpoints.

Model

K2 Horizon 0.9B is a 28-layer dense decoder-only transformer with grouped-query attention (32 query heads, 8 key-value heads), a 64,256-token vocabulary, and 0.9B parameters by IFM's count (1.08B including embeddings). Its original 8K context is extended to 131,072 tokens with YaRN RoPE scaling. The chat template emits a reasoning segment before each answer and exposes high, medium, and low reasoning-effort modes, and the model supports tool calling in JSON, XML, and typed-XML formats. The release checkpoint comes from multi-teacher on-policy distillation: math-and-code, STEM, and instruction-following RL experts were merged in weight space, then distilled into one student. IFM positions it for edge devices and lightweight tool use, reporting 48.5 on AIME 2026 and 79.9 on HumanEval+.

Intermediate checkpoints from pretraining and both context-extension stages are also published, and training logs are public on Weights & Biases; IFM says the training code and technical report will follow. The model is released under the Apache 2.0 license. Full details are in IFM's announcement and the model card.

Dedicated Endpoint

Docs

A Dedicated Endpoint is your own deployment of K2 Horizon 0.9B, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.

$ modal endpoint create --model IFM/K2-Horizon-0.9B

Related resources

Ship your first app in minutes.

Get Started

$30 / month free compute