Modal logo

Server

class Server(object)

Server runs an HTTP server started in an @modal.enter method.

See the guide for more information.

Generally, you will not construct a Server directly. Instead, use the @app.server() decorator.

@app.server(port=8080, routing_region="us-east")
class MyServer:
    @modal.enter()
    def start_server(self):
        self.process = subprocess.Popen(["python3", "-m", "http.server", "8080"])

object_id

object_id(self)

Modal's internal ID for this Server instance.

logs

logs: ServerLogsManager

Access logs for a Server.

Use fetch() to read logs from a UTC time range, tail() to read the most recent logs, and stream() to follow new logs as they arrive.

See Also

logs.fetch

fetch(self, *, since, until=None, source=None, search_text="")

Fetch Server logs corresponding to the date range and filters.

Parameters

since datetime
Start date to fetch logs from. Must be in UTC or timezone-naive, which is interpreted as local time.
until datetime | None
Defaults to current date if None. Must be in UTC or timezone-naive, which is interpreted as local time.
source LogSource | None
Filter by source: 'stdout', 'stderr', or 'system'.
search_text str
Filter by search text. (Default is "" )

Yields

LogEntry objects in chronological order.

Usage

server = modal.Server.from_name("my-app", "web")

for entry in server.logs.fetch(
    since=datetime.now() - timedelta(minutes=25),
    source="stdout",
):
    print(entry.message, end="")

logs.tail

tail(self, entries=100, *, source=None)

Fetch the most recent Server logs.

Parameters

entries int
The number of log entries to return. (Default is 100 )
source LogSource | None
Filter by source: 'stdout', 'stderr', or 'system'.

Yields

LogEntry objects in chronological order.

Usage

server = modal.Server.from_name("my-app", "web")

for entry in server.logs.tail(20):
    print(entry.message, end="")

logs.stream

stream(self, timeout=None)

Stream new Server logs until the timeout is reached.

Parameters

timeout float | None
Number of seconds to wait between log entries before terminating the stream. By default, this will block until it is interrupted.

Yields

LogEntry objects as they arrive.

Usage

server = modal.Server.from_name("my-app", "web")

for entry in server.logs.stream(timeout=60):
    print(entry.message, end="")

sessions

sessions: ServerSessionsManager

Start and terminate sticky sessions on a Server decorated with @modal.sessioned().

sessions.start

start(self, idle_timeout=600)

Start a sticky session and return its ID and token.

Requests to the server URL that carry the returned token are routed to the same container until the session has had no connections for idle_timeout seconds or is terminated. A container won't be scaled down for as long as it holds a live session.

If no container has room for the session, the call waits for additional capacity. It will block for up to 25 minutes before giving up. To control the wait per call, use the HTTP API and apply your own retry policy.

Parameters

idle_timeout int
Seconds without an in-flight request before the session ends. (Default is 600 )

Usage

server = modal.Server.from_name("my-app", "MyServer")
server_url = server.get_url()
session = server.sessions.start(idle_timeout=600)
headers = {"Modal-Authorization": f"Bearer {session.token}"}

requests.get(server_url, headers=headers).raise_for_status()

server.sessions.terminate(session.token)

sessions.terminate

terminate(self, token)

Terminate a sticky session.

New requests to it will be rejected. The container continues serving other sessions.

Parameters

token str
The token of the ServerSessionCredentials to terminate.

Usage

server = modal.Server.from_name("my-app", "MyServer")
session = server.sessions.start()

server.sessions.terminate(session.token)

info

info(self, *, refresh=False)

Get an overview of a Server's resource requests, associated mounts, http config, etc.

This method performs a network request to populate this information if the Server handle is a remote lookup whose information has not yet been fetched (e.g. from Server.from_name(...)), or if refresh=True.

Parameters

refresh bool
Always perform a network request. Pass refresh=True to ensure that this method returns the most up to date information. (Default is False )

Returns

This returns a modal.types.ServerInfo dataclass.

get_url

get_url(self)

The URL for making requests to this Server.

update_autoscaler

update_autoscaler(self, *, target_concurrency=None, min_containers=None,
    max_containers=None, buffer_containers=None, scaleup_window=None,
    scaledown_window=None)

Override the current autoscaler behavior for this Server.

Unspecified parameters will retain their current value, i.e. either the static value from the @app.server() decorator, or an override value from a previous call to this method.

Subsequent deployments of the App containing this Server will reset the autoscaler back to its static configuration.

Parameters

target_concurrency float | None
Target number of concurrent requests per container. May be fractional, e.g. 1.5 to target three concurrent requests per two containers.
min_containers int | None
Minimum number of containers to keep running regardless of demand.
max_containers int | None
Limit on the number of containers that can be concurrently running.
buffer_containers int | None
Extra containers to scale up beyond current demand.
scaleup_window int | None
Seconds of sustained demand required before scaling up new containers.
scaledown_window int | None
Maximum duration (in seconds) idle containers wait before scaling down.

Returns

A ServerAutoscalerSettings dataclass which contains the current autoscaler settings of this Server after the call.

Usage

server = modal.Server.from_name("my-app", "Server")

# Always have at least 2 containers running, with an extra buffer of 2 containers
server.update_autoscaler(min_containers=2, buffer_containers=1)

# Limit this Server to avoid spinning up more than 5 containers
server.update_autoscaler(max_containers=5)

# Require 30 seconds of sustained demand before scaling up
server.update_autoscaler(scaleup_window=30)

# Adjust Server autoscaling to target 20 concurrent requests per replica
server.update_autoscaler(target_concurrency=20)

# Target three concurrent requests for every two containers
server.update_autoscaler(target_concurrency=1.5)

# Disable the Server autoscaling by setting target_concurrency to 0
server.update_autoscaler(target_concurrency=0)

hydrate

hydrate(self, client=None)

Synchronize the local object with its identity on the Modal server.

It is rarely necessary to call this method explicitly, as most operations will lazily hydrate when needed. The main use case is when you need to access object metadata, such as its ID.

from_name

from_name(cls, app_name, name, *, environment_name=None, client=None)

Reference a Server from a deployed App by its name.

This is a lazy method that defers hydrating the local object with metadata from Modal servers until the first time it is actually used.

from_id

from_id(cls, server_id, *, client=None)

Reference a Server from a deployed or running App by its ID.

This is a lazy method that defers hydrating the local object with metadata from Modal servers until the first time it is actually used.

Parameters

server_id str
The ID of the server.
client _Client | None
Modal client instance for this session.

Usage

server = modal.Server.from_id("fu-456")

stats

stats(self, *, since=None, until=None, container=None)

Return statistics for a modal Server.

The default time range is the most recent hour. The maximum time range is 7 days.

Parameters

since datetime | None
The beginning of the time range, inclusive. If omitted, this defaults to an hour before until. Values without a timezone are interpeted as local time.
until datetime | None
The end of the time range, exclusive. If omitted, this defaults to current time. Values without a timezone are interpeted as local time.
container str | None
If passed in, the stats are computed for only this container. Default None.

Returns

A ServerStats object