PhoneLLM Alpha 1 is Pipecat's open-weights text model for phone and voice agents: a full-parameter fine-tune of NVIDIA Nemotron 3 Nano 30B-A3B with 30B total and 3.5B active parameters, a 256K-token context, accurate multi-turn tool calling with thinking disabled, and BSD 2-Clause weights.
PhoneLLM Alpha 1 is a full-parameter supervised fine-tune of NVIDIA's Nemotron 3 Nano 30B-A3B, trained by the Pipecat team at Daily with the NVIDIA NeMo framework. It keeps the base model's hybrid Mamba-Transformer mixture-of-experts architecture: 52 layers mixing Mamba-2, MoE and grouped-query attention blocks, 128 routed experts plus one shared expert with 6 active per token, 30B total parameters with 3.5B active, and a 262,144-token context. The fine-tune targets English-language phone agents that handle inbound customer-service calls and common outbound calling tasks, and is trained to call the right tools at the right time across long multi-turn conversations without thinking enabled, so the model answers without a reasoning delay; Pipecat recommends temperature 0 and thinking off, matching the training setup.
PhoneLLM Alpha 1 is released under the BSD 2-Clause license; as a derivative of Nemotron 3 Nano, redistribution must also carry the NVIDIA Nemotron Open Model License and attribution notices. Full details are in Pipecat's announcement and the model card.
A Dedicated Endpoint is your own deployment of PhoneLLM Alpha 1, on an inference stack that Modal tunes and autoscales. Bring fine-tuned weights if you have them. Compute is billed by the second, and only while it runs.
modal endpoint create --model pipecat-ai/phonellm-alpha-1