Server
class Server(object)Server runs an HTTP server started in an @modal.enter method.
See the guide for more information.
Generally, you will not construct a Server directly.
Instead, use the @app.server() decorator.
@app.server(port=8080, routing_region="us-east")
class MyServer:
@modal.enter()
def start_server(self):
self.process = subprocess.Popen(["python3", "-m", "http.server", "8080"])object_id
object_id(self)Modal's internal ID for this Server instance.
logs
logs: ServerLogsManagerAccess logs for a Server.
Use fetch()
to read logs from a UTC time range, tail()
to read the most recent logs, and stream()
to follow new logs as they arrive.
See Also
modal app logs: CLI access to logs for an App.
logs.fetch
fetch(self, *, since, until=None, source=None, search_text="")Fetch Server logs corresponding to the date range and filters.
Parameters
Yields
LogEntry objects in chronological order.
Usage
server = modal.Server.from_name("my-app", "web")
for entry in server.logs.fetch(
since=datetime.now() - timedelta(minutes=25),
source="stdout",
):
print(entry.message, end="")logs.tail
tail(self, entries=100, *, source=None)Fetch the most recent Server logs.
Parameters
Yields
LogEntry objects in chronological order.
Usage
server = modal.Server.from_name("my-app", "web")
for entry in server.logs.tail(20):
print(entry.message, end="")logs.stream
stream(self, timeout=None)Stream new Server logs until the timeout is reached.
Parameters
Yields
LogEntry objects as they arrive.
Usage
server = modal.Server.from_name("my-app", "web")
for entry in server.logs.stream(timeout=60):
print(entry.message, end="")sessions
sessions: ServerSessionsManagerStart and terminate sticky sessions on a Server decorated with @modal.sessioned().
sessions.start
start(self, idle_timeout=600)Start a sticky session and return its ID and token.
Requests to the server URL that carry the returned token are routed to the same container until the
session has had no connections for idle_timeout seconds or is terminated. A container won't be scaled down
for as long as it holds a live session.
If no container has room for the session, the call waits for additional capacity. It will block for up to 25 minutes before giving up. To control the wait per call, use the HTTP API and apply your own retry policy.
Parameters
Usage
server = modal.Server.from_name("my-app", "MyServer")
server_url = server.get_url()
session = server.sessions.start(idle_timeout=600)
headers = {"Modal-Authorization": f"Bearer {session.token}"}
requests.get(server_url, headers=headers).raise_for_status()
server.sessions.terminate(session.token)sessions.terminate
terminate(self, token)Terminate a sticky session.
New requests to it will be rejected. The container continues serving other sessions.
Parameters
Usage
server = modal.Server.from_name("my-app", "MyServer")
session = server.sessions.start()
server.sessions.terminate(session.token)info
info(self, *, refresh=False)Get an overview of a Server's resource requests, associated mounts, http config, etc.
This method performs a network request to populate this information if the Server handle is
a remote lookup whose information has not yet been fetched (e.g. from Server.from_name(...)),
or if refresh=True.
Parameters
Returns
This returns a modal.types.ServerInfo
dataclass.
get_url
get_url(self)The URL for making requests to this Server.
update_autoscaler
update_autoscaler(self, *, target_concurrency=None, min_containers=None,
max_containers=None, buffer_containers=None, scaleup_window=None,
scaledown_window=None)Override the current autoscaler behavior for this Server.
Unspecified parameters will retain their current value, i.e. either the static value
from the @app.server() decorator, or an override value from a previous call to this method.
Subsequent deployments of the App containing this Server will reset the autoscaler back to its static configuration.
Parameters
Returns
A ServerAutoscalerSettings dataclass which contains the current autoscaler settings of
this Server after the call.
Usage
server = modal.Server.from_name("my-app", "Server")
# Always have at least 2 containers running, with an extra buffer of 2 containers
server.update_autoscaler(min_containers=2, buffer_containers=1)
# Limit this Server to avoid spinning up more than 5 containers
server.update_autoscaler(max_containers=5)
# Require 30 seconds of sustained demand before scaling up
server.update_autoscaler(scaleup_window=30)
# Adjust Server autoscaling to target 20 concurrent requests per replica
server.update_autoscaler(target_concurrency=20)
# Target three concurrent requests for every two containers
server.update_autoscaler(target_concurrency=1.5)
# Disable the Server autoscaling by setting target_concurrency to 0
server.update_autoscaler(target_concurrency=0)hydrate
hydrate(self, client=None)Synchronize the local object with its identity on the Modal server.
It is rarely necessary to call this method explicitly, as most operations will lazily hydrate when needed. The main use case is when you need to access object metadata, such as its ID.
from_name
from_name(cls, app_name, name, *, environment_name=None, client=None)Reference a Server from a deployed App by its name.
This is a lazy method that defers hydrating the local object with metadata from Modal servers until the first time it is actually used.
from_id
from_id(cls, server_id, *, client=None)Reference a Server from a deployed or running App by its ID.
This is a lazy method that defers hydrating the local object with metadata from Modal servers until the first time it is actually used.
Parameters
Usage
server = modal.Server.from_id("fu-456")stats
stats(self, *, since=None, until=None, container=None)Return statistics for a modal Server.
The default time range is the most recent hour. The maximum time range is 7 days.
Parameters
Returns
A ServerStats object