Modular

AI inference tools from kernels to serving

Contact sales

Is Modular right for you?

Good for

  • AI engineering teams deploying and optimizing inference workloads
  • Tools spanning kernel and serving layers
  • Hosted and customer-cloud deployment options

Keep in mind

  • Hardware support, deployment licensing and workload costs need technical evaluation.
  • Cloud serving is usage-based by tokens or GPU time; BYOC and enterprise arrangements depend on deployment. Trial shared endpoints are offered.

Choose a plan for your work.

Contact sales

Explore plans

Pricing and access

Contact sales

Paid Inquire

See plans on Modular
More about Modular

Modular provides an AI inference stack spanning GPU kernels, model execution and cloud serving. Teams can use hosted endpoints or explore their own deployment arrangements across supported hardware.

Available on

Api

Alternatives to Modular

Choose around the work you need to do.

AgentOps

Trace and debug AI agents.

Explore

AI/ML API

Access multiple AI models through one API

Explore

AI21

Optimize agent systems with model tuning and routing.

Explore

Arcade

Control the actions AI agents take.

Explore