Skip to content
EN
English 简体中文 soon 日本語 soon

Modal

Serverless GPU compute from Python

Visit official site

What Modal is

Modal is a serverless platform for AI and data workloads. You define functions in Python and it runs them in the cloud, with GPU scheduling, automatic scaling, fast cold starts, custom container images, cloud storage mounts, logging and HTTPS endpoints.

Billing follows actual CPU and GPU usage time, and a monthly free compute allowance is described on the site.

What you can do with it

  • Serve a model behind an HTTPS endpoint
  • Run batch inference or data processing
  • Fine-tune models on rented GPUs
  • Schedule recurring compute jobs

Who it is for

  • Developers deploying AI inference
  • Teams avoiding GPU infrastructure work
  • Startups wanting pay-per-use compute

What to watch out for

  • Costs scale with runtime; long GPU jobs get expensive quickly
  • Cold start and queue behaviour vary by GPU class and region
  • Data and model artefacts live on the platform unless you manage them
  • Confirm current GPU availability and free credit terms

Pros & cons

✓ What we like

  • No infrastructure to operate
  • Fine-grained billing
  • GPU access from a Python function

! What to watch out for

  • Cost grows with runtime
  • Platform data residency
  • GPU availability varies

FAQ

What is it for?

Running AI inference, batch jobs and fine-tuning as serverless cloud functions.

Do I manage servers?

No. The platform handles scheduling and scaling.

How is it billed?

By actual CPU and GPU usage time, with a monthly free allowance described on the site.

Last reviewed: 2026-09-17

More AI coding tools tools

View all →

How we review