Runtime is almost here: join TypeSafe AI, Cognition, DoorDash and more in SF. Limited seats left. Register now

GPT-OSS 120B

GPT-OSS 120B is OpenAI's open-weight reasoning model: a 117-billion-parameter Mixture-of-Experts model with 5.1B active parameters per token and a 128K-token context window, built for production agentic and high-reasoning workloads that fit on a single GPU.

Model

GPT-OSS 120B uses a 36-layer Mixture-of-Experts architecture with 128 experts, four of which are active per token. Its MoE weights ship in the 4-bit MXFP4 format so the full model fits in 80GB of GPU memory. The model exposes low, medium, and high reasoning effort levels, supports function calling and tool use, and is trained on OpenAI's harmony response format. Full details are in OpenAI's announcement and the model card.

Dedicated Endpoint

Docs

A Dedicated Endpoint is your own deployment of GPT-OSS 120B, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.

$ modal endpoint create --model openai/gpt-oss-120b

Related resources

Ship your first app in minutes.

Get Started

$30 / month free compute