How Physical Intelligence runs remote, real-time, robotic inference on Modal.
語言: en
索引內容摘要
All posts Back Customer Stories April 8, 2026 •5 minute read Real-time inference for robots at Physical Intelligence Physical Intelligence (Pi) is building a general-purpose robotic intelligence system capable of operating any robot across any task. Their core model—a Visual-Language-Action (VLA) architecture—takes visual observations, natural-language instructions, and the robot’s proprioceptive state, then outputs motor commands for the next fraction of a second. Every arm movement in their system flows through this closed loop of continuous inference. To evaluate progress, Pi doesn’t just rely on simulation. Every model revision must be validated on real robots performing real tasks. That means thousands of inference cycles running 24/7 across a growing fleet of robots.Running this compute on Modal simplified operations and enabled rapid experimentation with larger models, while only adding 10-15ms of network overhead.Designing for the constraints of real-time latencyFor standard local inference, Pi models are trained in the cloud then downloaded locally to run using on-board GPUs attached to each robotic system. This provides a reliable system with good performance and straight…
首次發現: 最近檢查: 內容更新:
文章站內連結
Real-time, multi-node inference for Runway Characters
Modal is proud to power real-time inference for Runway Characters.
語言: en
索引內容摘要
All posts Back Customer Stories March 26, 2026 •3 minute read Runway chooses Modal to power real-time inference for Runway Characters Today, we're announcing that Runway is partnering with Modal to power real-time inference for Runway Characters.Runway Characters is a real-time video agent API that lets developers, startups, enterprises and consumers build fully custom conversational characters. These video agents can have any appearance and any visual style, with full control over voice, personality, knowledge and actions. Built on Runway’s general world model, GWM-1, Characters generates expressive digital personas from a single image, with zero fine-tuning required. Thousands of organizations are already using Characters, including Fortune 10 technology companies, major Hollywood studios, global advertising agencies and gaming companies, with use cases ranging from customer support and internal training to experiential advertising and immersive game worlds. Characters represents the first step toward a future of online interaction built around real-time video rather than text.This kind of continuous, expressive, low-latency video generation held across extended conversations and…
Learn how Chai Discovery moves seamlessly from ML experimentation to production antibody pipelines with Modal.
語言: en
索引內容摘要
All posts Back Customer Stories January 15, 2026 •5 minute read Seamless computational bio at Chai Discovery Greta Workman Product Marketing @gretaworkman Chai Discovery is a frontier drug discovery company using machine learning to design new medicines. Their mission is to develop a flexible, ML-driven platform that can adapt to new biological targets and experimental data, accelerating discovery across diseases and modalities—the “computer-aided design suite for molecules”.Too often, infrastructure is the bottleneck for research and discovery. By building on Modal, Chai can scale experiments seamlessly, keep data consistent, and run the same workflows from research through production.The challenge: complex, bursty bio workloadsChai’s machine learning pipelines combine diverse models, large biological datasets, and GPU-heavy computations. Each experiment can differ dramatically in scale, from small protein structure tests to full antibody design campaigns, and must scale from one run to thousands overnight. The workloads are heterogeneous and bursty, with frequent precomputation steps each with shifting hardware demands.Running all this on traditional cloud infrastructure would ha…
Learn how Reducto used GPU memory snapshotting and flexible autoscaling to build fast multi-model pipelines.
語言: en
索引內容摘要
All posts Back Customer Stories November 19, 2025 •6 minute read How Reducto improved enterprise-scale document processing latency by 3x Reducto is on a mission to turn messy, real world documents into structured data. Their platform ingests and processes millions of PDFs, spreadsheets, and slide decks per day. Reducto helps companies ranging from AI-native startups to the world’s largest enterprises and hedge funds, particularly in sensitive verticals like finance, legal, insurance, and healthcare. Their customers bring demanding workloads: millions of pages at once, highly variable traffic, and strict latency requirements for real-time user interactions. To maintain a seamless experience, Reducto needed elastic infrastructure for their multi-model pipelines. That’s where Modal came in. The challenge: spiky, high-stakes workloads Early on, Reducto ran their product on manually provisioned EC2 instances. As usage grew, they adopted Kubernetes to manage their compute. It worked for a while, but problems soon emerged as Reducto kept growing: Scaling a monolith: With dozens of models in production, Reducto had to scale all of them together, even though usage patterns varied. Over-prov…
How Decagon and Modal made real-time voice AI possible, combining fine-tuned small models with a re-engineered inference runtime for sub-second latency.
語言: en
索引內容摘要
All posts Back Customer Stories November 13, 2025 •5 minute read How Decagon shipped real-time voice AI on Modal Richard Gong Member of Technical Staff Timothy Feng Member of Technical Staff Cyrus Asgari Research Engineer at Decagon This post was written in conjunction with the Decagon team. Decagon provides a unified platform to build, optimize, and scale AI agents that deliver concierge-level customer experiences across every channel. Earlier this year, they launched Decagon Voice, enabling teams to build fast, intelligent voice AI agents tailored to each brand’s tone, personality, and performance needs. Voice is among the most technically demanding frontiers in AI. It requires sub-second latency, natural turn-taking, and seamless responsiveness, all while maintaining conversational quality. Meeting these constraints in production and at massive scale demanded breakthroughs in both model accuracy and inference performance. To make it possible, the Decagon and Modal teams partnered closely on two fronts: Improving model accuracy: Using advanced supervised fine-tuning (SFT) and reinforcement learning (RL) techniques, the team trained a suite of compact open-source models that deliv…
During a single weekend event, Lovable users built 250,000 new applications, all running in isolated development environments. Lovable used Modal to generate 1 million code sandboxes—with 20,000 running concurrently at peak—over just 48 hours.
語言: en
索引內容摘要
All posts Back Customer Stories July 7, 2025 •4 minute read How Modal powered 250,000 Lovable app creations in a weekend Lovable is one of the fastest growing startups in history, having hit $75M ARR a mere 7 months after they launched. Lovable is on a mission to build the “last piece of software,” enabling anyone to create fully functional applications without writing code. Users enter a prompt describing the app they want, and Lovable generates the code, spins up the application, and enables real-time visual edits. Lovable uses Modal Sandboxes at massive scale to run LLM-generated code for thousands of apps in safe and isolated sandboxes. The problem: performance and reliability at scale In June 2025, Lovable had a major promotional weekend event coming up, in collaboration with Anthropic, OpenAI, and Google. This event was set to dramatically increase user activity. Unfortunately, Lovable’s existing code sandbox provider—a distributed cloud VM platform—raised concerns around scalability and operational risk. The team wasn’t confident the provider could scale fast enough to meet the expected demand, and with only this single vendor in place, Lovable was exposed to high risk of se…
首次發現: 最近檢查: 內容更新:
文章站內連結
“We’re actively saving 2 engineers’ worth of ongoing time”
Quora is building Poe, a platform where anyone can deploy a public AI chatbot. Quora uses Modal Sandboxes at scale to safely run LLM-generated code in the context of user chats.
語言: en
索引內容摘要
All posts Back Customer Stories June 30, 2025 •3 minute read How Quora uses Modal to run thousands of Python sandboxes simultaneously Quora is a Q&A platform where users can ask, answer, and peruse questions on a variety of topics. With 400 million monthly unique visitors, it’s an invaluable contributor to the world’s knowledge-sharing. Quora uses Modal Sandboxes to securely execute LLM-generated code in Poe, their AI chatbot platform. The team shipped months earlier using Modal rather than building in-house. They’re also saving an ongoing 2 engineers’ worth of infrastructure maintenance time! Hello, Poe In 2023, Quora launched Poe, an AI chatbot platform where anyone can deploy a public chatbot. With millions of monthly active users, Poe is the default destination for many AI builders to experiment with different models. Quora has since raised $75M to keep expanding Poe. A code interpreter for Poe Many of the LLM bots in Poe can generate code, and users expected to run that code in Poe rather than copy-pasting it to their editors. The Quora team needed a way to safely execute code in Poe in a completely isolated way, keeping that code separate from both the main Quora infrastructu…
首次發現: 最近檢查: 內容更新:
文章站內連結
“Modal makes it easy to write code that runs on 100s of GPUs in parallel, transcribing podcasts in a fraction of the time.”
Learn how Substack sped up their developer iteration cycles by moving ML training and deployment to Modal from AWS SageMaker.
語言: en
索引內容摘要
All posts Back Customer Stories May 20, 2024 •3 minute read Why Substack moved their AI and ML pipelines to Modal Substack is a popular online platform for writers to publish newsletters, with over 17k writers and $300M in paid subscription volume. Substack employs ML for various purposes, including spam detection, newsletter recommendations, audio transcription, sentiment analysis, and image generation. For nearly all these models Substack has moved both training and deployment from AWS SageMaker to Modal. The challenges deploying AI and ML before Modal Previously, Substack’s training and deployment pipelines were built on AWS SageMaker and orchestrated with Airflow. Adding or updating models was a slow and painful process for a few reasons. One, the developer experience on SageMaker was convoluted. Engineers had to navigate to the SageMaker product, create a notebook, specify machine requirements, and wait for that machine to turn on, all before a single line of code could be written. Not to mention the difficulty of juggling multiple remote environments—from the Jupyter notebooks to the SageMaker training machines to the final production infra. Two, collaborating was difficult. …
Find out how Suno uses Modal to scale inference and batch pre-processing to thousands of GPUs.
語言: en
索引內容摘要
All posts Back Customer Stories February 21, 2024 •3 minute read How Suno shaved 4 months off their launch timeline with Modal Suno uses Modal to scale inference and batch pre-processing to thousands of GPUs. With Modal, Suno was able to bring a state-of-the-art music generation model to market four months early instead of hiring a team of engineers to build and maintain infrastructure. About Suno Suno is a music generation app that can make any song you describe. Enter a simple text description—like “a deep house song about serverless infra”—and Suno makes you a song complete with vocals in seconds. Suno’s users include Grammy-winning artists, but the core user base is people experiencing making music for the first time. Microsoft recently announced they’ve partnered with Suno to bring song generation capabilities to Copilot, their AI chatbot! Avoiding past infrastructure pain Prior to starting Suno, all four founders worked at Kensho, an AI tech startup for financial data. They had personally spent significant amounts of time setting up and managing Kubernetes clusters to support their data-heavy workloads—so when they started working on Suno, they knew exactly what they did not …
This example demonstrates how to run OpenCode remotely and connect to it from your local terminal or browser.
語言: en
索引內容摘要
Run OpenCode in a Modal Sandbox This example demonstrates how to run OpenCode remotely and connect to it from your local terminal or browser. Combine self-hosted OpenCode with serving a big, smart model on Modal and you’ve got “coding agents at home”! Coding agents are most useful when they have context and tools. By default, this script clones the Modal examples repo and gives the agent access to your Modal credentials, so it can run and debug examples (including this one!). Meta. Set up OpenCode on Modal import argparse import os from pathlib import Path import modal MINUTES = 60 HOURS = 60 * MINUTES OPENCODE_PORT = 4096 DEFAULT_GITHUB_REPO = "modal-labs/modal-examples" First, we define a Modal container Image with OpenCode installed. def define_base_image() -> modal.Image: image = ( modal.Image.debian_slim() .apt_install("curl", "git", "gh") .run_commands("curl -fsSL https://opencode.ai/install | bash") .env({"PATH": "/root/.opencode/bin:${PATH}"}) ) # We also bring the global default OpenCode configuration along for the ride. CONFIG_PATH = Path("~/.config/opencode/opencode.json").expanduser() if CONFIG_PATH.exists(): print("🏖️ Including config from", CONFIG_PATH) image = image…
Explore open models you can run and customize on Modal.
語言: en
索引內容摘要
Model Library Explore open models you can run and customize on Modal. Language & vision models 8 models Z.ai GLM 5.3 A natively multimodal model for coding, agents, and long-context workloads. Language 753B MoE · 1M context Z.ai GLM 5.3 Flash A natively multimodal model for coding, agents, and long-context workloads. Vision + language 320B MoE · 1M context Qwen Qwen3.8-Max The open-weights, text-only Max-class model for coding and long-horizon agentic work. Language 2.4T MoE · Up to 1M context Moonshot AI Kimi K3 Open frontier intelligence for deep reasoning and knowledge work. Vision + language 2.8T MoE · 1M context Thinking Machines Inkling NVFP4 An open-weights multimodal model with controllable reasoning across text, images, and audio. Multimodal 975B MoE · 1M context DeepSeek DeepSeek V4.1 Flash A strong MoE model with a 1M-token context, native visual understanding, and an 3.9× smaller KV cache than V4 Flash. Vision + language 552B MoE · 1M context OpenAI GPT-OSS 120B OpenAI's open-weight reasoning model with configurable reasoning effort and native tool use, sized to fit on a single GPU. Language 117B MoE · 128K context Google Gemma 4 31B IT Google DeepMind's largest open Ge…
Modal provides high-performance, scalable computing infrastructure. It's our vision of a cloud accessible to all developers.
語言: en
索引內容摘要
Runtime, the conference for engineers running AI in production. Oct. 1 in SF Register now Product Solutions Resources CustomersPricingDocs Log In Sign Up Making cloud developmentwork like magic We're hiring! Our founders Erik Bernhardsson and Akshat Bubna, pictured with Mugi (Akshat's dog). Not pictured: Erik's cats Lance and Stan. We started Modal with the goal to make it easier to iterate and ship applications for data, AI, and machine learning. In order to deliver the developer experience we wanted, we went deep and built our own infrastructure — including our own custom file system, container runtime, scheduler, container image builder, and much more.Customers use Modal for a wide range of use cases, including Generative AI inference, LLM fine-tuning, computational biotech, and media processing. We're focused on developer experience at our core, letting companies ship value faster without having to think about infrastructure. In a few lines of code, we let you scale from zero to thousands of CPUs or GPUs. And since pricing is entirely usage-based, you only pay for the time your code is running.Our team is based out of New York, Stockholm, and San Francisco. It includes creators…
Runtime, the conference for engineers running AI in production. Oct. 1 in SF Register now Product Solutions Resources CustomersPricingDocs Log In Sign Up Software as a Service Agreement Effective May 2026. This Software as a Service Agreement (the "Agreement") is between the entity named below ("Customer") and Modal Labs, Inc., a Delaware corporation ("Modal"). This Agreement consists of these terms, each order form for Services that has been executed by Modal and Customer (each an "Service Order") and all exhibits and amendment of any of the foregoing. Customer consents to this Agreement by executing a Service Order or creating an account on the Service. SCOPE OF SERVICE AND RESTRICTIONS Access and Scope of Service. Subject to receipt of the applicable Fees with respect to the service specified in the corresponding Service Order (the "Service(s)"), Modal will make the Service available to Customer as set forth in this Agreement and the Service Order. Subject to Customer's compliance with the terms and conditions of the Agreement and the Service Order, Customer may access and use the Service during the period specified in the Service Order. Any such use of the Service by Customer i…
This example finetunes the Flux.1-dev model on images of a pet (by default, a puppy named Qwerty) using a technique called textual inversion from the “Dreambooth” paper. Effectively, it teaches a general image generation model a new “proper noun”, allowing for the personalized generation of art and photos. We supplement textual inversion with low-rank adaptation (LoRA) for increased efficiency during training.
語言: en
索引內容摘要
Fine-tune Flux on your pet using LoRA This example finetunes the Flux.1-dev model on images of a pet (by default, a puppy named Qwerty) using a technique called textual inversion from the “Dreambooth” paper. Effectively, it teaches a general image generation model a new “proper noun”, allowing for the personalized generation of art and photos. We supplement textual inversion with low-rank adaptation (LoRA) for increased efficiency during training. It then makes the model shareable with others — without costing $25/day for a GPU server— by hosting a Gradio app on Modal. It demonstrates a simple, productive, and cost-effective pathway to building on large pretrained models using Modal’s building blocks, like GPU-accelerated Modal Functions for compute-intensive work, Volumes for storage, and Web Functions for serving. And with some light customization, you can use it to generate images of your pet! You can find a video walkthrough of this example on the Modal YouTube channel here. Imports and setup We start by importing the necessary libraries and setting up the environment. from dataclasses import dataclass from pathlib import Path import modal Building up the environment Machine le…
Runtime, the conference for engineers running AI in production. Oct. 1 in SF Register now Product Solutions Resources CustomersPricingDocs Log In Sign Up Privacy Policy Last updated May 17th, 2023. Privacy Policy of Modal Labs Modal Labs operates the modal.com website, which provides a cloud computing service. This page is used to inform website visitors regarding our policies with the collection, use, and disclosure of Personal Information if anyone decided to use our Service, the modal.com website. If you choose to use our Service, then you agree to the collection and use of information in relation with this policy. The Personal Information that we collect are used for providing and improving the Service. We will not use or share your information with anyone except as described in this Privacy Policy. The terms used in this Privacy Policy have the same meanings as in our Terms and Conditions, which is accessible at Website URL, unless otherwise defined in this Privacy Policy. Information Collection and Use For a better experience while using our Service, we may require you to provide us with certain personally identifiable information, including but not limited to your name, phon…
View and follow events from Modal on Luma. AI infrastructure that developers love
語言: en
索引內容摘要
(function () { var applied = []; function define(target, name, value) { Object.defineProperty(target, name, { value: value, writable: true, configurable: true, enumerable: false }); applied.push(name); } // Chrome 93. Unguarded in a timezone module and in the react-markdown chain. if (!Object.hasOwn) { define(Object, "hasOwn", function hasOwn(target, key) { if (target === null || target === undefined) { throw new TypeError("Cannot convert undefined or null to object"); } return Object.prototype.hasOwnProperty.call(Object(target), key); }); } // Chrome 92. Unguarded throughout web-vitals. function at(index) { var target = Object(this); var length = target.length >>> 0; var relative = Math.trunc(Number(index)) || 0; var resolved = relative < 0 ? length + relative : relative; return resolved < 0 || resolved >= length ? undefined : target[resolved]; } if (!Array.prototype.at) { define(Array.prototype, "at", at); } if (!String.prototype.at) { define(String.prototype, "at", at); } // Chrome 85. Unguarded in query-string's parser, so it runs on any URL with // a query — which is how this first reached us. if (!String.prototype.replaceAll) { define(String.prototype, "replaceAll", function …
GuideExamplesReference Releases Playground Search documentation / Log In Sign Up Modal Documentation Modal provides a serverless cloud for engineers and researchers who want to build compute-intensive applications without thinking about infrastructure. Run generative AI models, large-scale batch workflows, job queues, and more, all faster than ever before. Get Started Guide Everything you need to know to run code on Modal. Dive deep into all of our features and best practices. Examples Powerful applications built with Modal. Explore guided starting points for your use case. Reference Technical information about the Modal API. Quickly refer to basic descriptions of various programming functionalities. Playground Interactive tutorials to learn how to start using Modal. Run serverless cloud functions from your browser. Guide Everything you need to know to run code on Modal. Dive deep into all of our features and best practices. Examples Powerful applications built with Modal. Explore guided starting points for your use case. Reference Technical information about the Modal API. Quickly refer to basic descriptions of various programming functionalities. Playground Interactive tutorials …
Runtime, the conference for engineers running AI in production. Oct. 1 in SF Register now Product Solutions Resources CustomersPricingDocs Log In Sign Up Modal Training Train more, configure less Launch more experiments and training jobs. Spin up single-node experiments or scale to multi-node GPU training instantly. Get Started Read the docs “Modal lets us deploy new ML models in hours rather than weeks. We use it across spam detection, recommendations, audio transcription, and video pipelines, and it’s helped us move faster with far less complexity.” Mike Cohen, Head of AI & ML Engineering “Modal's user-friendly interface and efficient tools have truly empowered our team to navigate data-intensive tasks with ease, enabling us to achieve our project goals more efficiently.” Karim Atiyeh, Co-Founder & CTO DEFINE IN CODE NATIVE STORAGE SUB-SECOND STARTUP Modal Training Where researchers can run experiments, not ops Define in code Define your training function with Modal’s SDK. Easily keep ML dependencies and GPU requirements in sync with application code. 01 02 03 04 05 06 07 08 09 10 11 12 13 14 Native storage Ingest training data from anywhere: Modal’s distributed Volumes, cloud bu…
The document outlines Modal’s security and privacy commitments.
語言: en
索引內容摘要
Security and privacy at Modal The document outlines Modal’s security and privacy commitments. Application security (AppSec) AppSec is the practice of building software that is secure by design, secured during development, secured with testing and review, and deployed securely. We build our software using memory-safe programming languages, including Rust (for our worker runtime and storage infrastructure) and Python (for our API servers and Modal client). Software dependencies are audited by Github’s Dependabot. We make decisions that minimize our attack surface. Most interactions with Modal are well-described in a gRPC API, and occur through modal, our open-source command-line tool and Python client library. We have automated synthetic monitoring test applications that continuously check for network and application isolation within our runtime. We use HTTPS for secure connections. Modal forces HTTPS for all services using TLS (SSL), including our public website and the Dashboard to ensure secure connections. Modal’s client library connects to Modal’s servers over TLS and verify TLS certificates on each connection. All user data is encrypted in transit and at rest. All public Modal …
首次發現: 最近檢查: 內容更新:
頁面站內連結
Voice chat with LLMs Build an interactive voice chat app
QuiLLMan is a complete voice chat application built on Modal: you speak and the chatbot speaks back!
語言: en
索引內容摘要
QuiLLMan: Voice Chat with Moshi QuiLLMan is a complete voice chat application built on Modal: you speak and the chatbot speaks back! At the core is Kyutai Lab’s Moshi model, a speech-to-speech language model that will continuously listen, plan, and respond to the user. Thanks to bidirectional websocket streaming and Opus audio compression, response times on good internet can be nearly instantaneous, closely matching the cadence of human speech. You can find the demo live here. Everything — from the React frontend to the model backend — is deployed serverlessly on Modal, allowing it to automatically scale and ensuring you only pay for the compute you use. This page provides a high-level walkthrough of the GitHub repo. Code overview Traditionally, building a bidirectional streaming web application as compute-heavy as QuiLLMan would take a lot of work, and it’s especially difficult to make it robust and scale to handle many concurrent users. But with Modal, it’s as simple as writing two different classes and running a CLI command. Our project structure looks like this: Moshi Websocket Server: loads an instance of the Moshi model and maintains a bidirectional websocket connection with …