Modal
Serverless GPU cloud for ML inference, agent sandboxes, and batch jobs.
Executive Summary
Serverless GPU cloud for ML inference, agent sandboxes, and batch jobs.
Use Cases
- ML inference
- LLM inference
- Agent sandboxes
- Batch processing
- ML training
Features
Visibility
- Billing and Usage Monitoring: Monitor compute usage and costs through a dedicated dashboard.
- OpenTelemetry Integration: Export Modal logs to any OpenTelemetry-compatible observability provider.
Intelligence
- Instant Autoscaling: Automatically scales GPU resources from 0 to 1000+ based on demand, with no commitments.
- Sub-second Cold Starts: Rapid deployment and execution of ML workloads with minimal startup latency.
- Multi-cloud Orchestration: Routes workloads across clouds and regions in real time to optimize performance and availability.
Support
- Academic Program Credits: Unlock up to $10k of credits for academic projects to supercharge research.
Technical Specifications
- Architecture
- Serverless GPU cloud platform that autoscales from 0 to 1000+ GPUs, routes workloads across clouds and regions in real time, and offers sub-second cold starts for ML workloads.
- Deployment
- SaaS
- API Available
- Yes
Infrastructure
- Multi-cloud (AWS, GCP, Azure)
Integrations
- OpenTelemetry
Security & Compliance
Certifications: SOC 2 Type II, HIPAA (Enterprise plans with BAA)
Pricing
- Model
- Pay-as-you-go (usage-based)
- Starting Price
- Usage-based, see pricing page for rates
- Target Customer
- SMB,Mid-Market,Enterprise
- Contract Type
- Usage-based, billed per second/minute
- Free Trial
- Yes (no credit card required)
About Modal
Modal is a serverless compute platform that provides high-performance AI infrastructure. It abstracts away the complexities of GPU infrastructure management, making it easy for developers to run compute-intensive workloads like ML inference, fine-tuning, and AI agent sandboxes.