SEARWEB SITE INDEX

Modal · Indexed content

modal.com

Explore internal pages, articles and content excerpts discovered from this site’s public sources.

Internal links
43
Articles
9
Last indexed
20‏/9‏/2026، 9:30:17 م
43 indexed items
ArticleInternal link

Real-time robot control running on Modal with 10–15 ms latency.

https://modal.com/blog/physical-intelligence-runs-real-time-remote-inference-for-robotic-control-on-modal

Open original page

How Physical Intelligence runs remote, real-time, robotic inference on Modal.

Language: en
Indexed excerpt

All posts Back Customer Stories April 8, 2026 •5 minute read Real-time inference for robots at Physical Intelligence Physical Intelligence (Pi) is building a general-purpose robotic intelligence system capable of operating any robot across any task. Their core model—a Visual-Language-Action (VLA) architecture—takes visual observations, natural-language instructions, and the robot’s proprioceptive state, then outputs motor commands for the next fraction of a second. Every arm movement in their system flows through this closed loop of continuous inference. To evaluate progress, Pi doesn’t just rely on simulation. Every model revision must be validated on real robots performing real tasks. That means thousands of inference cycles running 24/7 across a growing fleet of robots.Running this compute on Modal simplified operations and enabled rapid experimentation with larger models, while only adding 10-15ms of network overhead.Designing for the constraints of real-time latencyFor standard local inference, Pi models are trained in the cloud then downloaded locally to run using on-board GPUs attached to each robotic system. This provides a reliable system with good performance and straight

Discovered: Last checked: Content changed:
ArticleInternal link

Real-time, multi-node inference for Runway Characters

https://modal.com/blog/runway-chooses-modal-to-power-real-time-inference-for-runway-characters

Open original page

Modal is proud to power real-time inference for Runway Characters.

Language: en
Indexed excerpt

All posts Back Customer Stories March 26, 2026 •3 minute read Runway chooses Modal to power real-time inference for Runway Characters Today, we're announcing that Runway is partnering with Modal to power real-time inference for Runway Characters.Runway Characters is a real-time video agent API that lets developers, startups, enterprises and consumers build fully custom conversational characters. These video agents can have any appearance and any visual style, with full control over voice, personality, knowledge and actions. Built on Runway’s general world model, GWM-1, Characters generates expressive digital personas from a single image, with zero fine-tuning required. Thousands of organizations are already using Characters, including Fortune 10 technology companies, major Hollywood studios, global advertising agencies and gaming companies, with use cases ranging from customer support and internal training to experiential advertising and immersive game worlds. Characters represents the first step toward a future of online interaction built around real-time video rather than text.This kind of continuous, expressive, low-latency video generation held across extended conversations and

Discovered: Last checked: Content changed:
ArticleInternal link

ML‑driven molecular design

https://modal.com/blog/seamless-computational-bio-at-chai-discovery

Open original page

Learn how Chai Discovery moves seamlessly from ML experimentation to production antibody pipelines with Modal.

Language: en
Indexed excerpt

All posts Back Customer Stories January 15, 2026 •5 minute read Seamless computational bio at Chai Discovery Greta Workman Product Marketing @gretaworkman Chai Discovery is a frontier drug discovery company using machine learning to design new medicines. Their mission is to develop a flexible, ML-driven platform that can adapt to new biological targets and experimental data, accelerating discovery across diseases and modalities—the “computer-aided design suite for molecules”.Too often, infrastructure is the bottleneck for research and discovery. By building on Modal, Chai can scale experiments seamlessly, keep data consistent, and run the same workflows from research through production.The challenge: complex, bursty bio workloadsChai’s machine learning pipelines combine diverse models, large biological datasets, and GPU-heavy computations. Each experiment can differ dramatically in scale, from small protein structure tests to full antibody design campaigns, and must scale from one run to thousands overnight. The workloads are heterogeneous and bursty, with frequent precomputation steps each with shifting hardware demands.Running all this on traditional cloud infrastructure would ha

Discovered: Last checked: Content changed:
ArticleInternal link

3x latency decrease for document processing

https://modal.com/blog/reducto-case-study

Open original page

Learn how Reducto used GPU memory snapshotting and flexible autoscaling to build fast multi-model pipelines.

Language: en
Indexed excerpt

All posts Back Customer Stories November 19, 2025 •6 minute read How Reducto improved enterprise-scale document processing latency by 3x Reducto is on a mission to turn messy, real world documents into structured data. Their platform ingests and processes millions of PDFs, spreadsheets, and slide decks per day. Reducto helps companies ranging from AI-native startups to the world’s largest enterprises and hedge funds, particularly in sensitive verticals like finance, legal, insurance, and healthcare. Their customers bring demanding workloads: millions of pages at once, highly variable traffic, and strict latency requirements for real-time user interactions. To maintain a seamless experience, Reducto needed elastic infrastructure for their multi-model pipelines. That’s where Modal came in. The challenge: spiky, high-stakes workloads Early on, Reducto ran their product on manually provisioned EC2 instances. As usage grew, they adopted Kubernetes to manage their compute. It worked for a while, but problems soon emerged as Reducto kept growing: Scaling a monolith: With dozens of models in production, Reducto had to scale all of them together, even though usage patterns varied. Over-prov

Discovered: Last checked: Content changed:
ArticleInternal link

65% Latency reduction

https://modal.com/blog/decagon-case-study

Open original page

How Decagon and Modal made real-time voice AI possible, combining fine-tuned small models with a re-engineered inference runtime for sub-second latency.

Language: en
Indexed excerpt

All posts Back Customer Stories November 13, 2025 •5 minute read How Decagon shipped real-time voice AI on Modal Richard Gong Member of Technical Staff Timothy Feng Member of Technical Staff Cyrus Asgari Research Engineer at Decagon This post was written in conjunction with the Decagon team. Decagon provides a unified platform to build, optimize, and scale AI agents that deliver concierge-level customer experiences across every channel. Earlier this year, they launched Decagon Voice, enabling teams to build fast, intelligent voice AI agents tailored to each brand’s tone, personality, and performance needs. Voice is among the most technically demanding frontiers in AI. It requires sub-second latency, natural turn-taking, and seamless responsiveness, all while maintaining conversational quality. Meeting these constraints in production and at massive scale demanded breakthroughs in both model accuracy and inference performance. To make it possible, the Decagon and Modal teams partnered closely on two fronts: Improving model accuracy: Using advanced supervised fine-tuning (SFT) and reinforcement learning (RL) techniques, the team trained a suite of compact open-source models that deliv

Discovered: Last checked: Content changed:
ArticleInternal link

Powering AI app generation at scale

https://modal.com/blog/lovable-case-study

Open original page

During a single weekend event, Lovable users built 250,000 new applications, all running in isolated development environments. Lovable used Modal to generate 1 million code sandboxes—with 20,000 running concurrently at peak—over just 48 hours.

Language: en
Indexed excerpt

All posts Back Customer Stories July 7, 2025 •4 minute read How Modal powered 250,000 Lovable app creations in a weekend Lovable is one of the fastest growing startups in history, having hit $75M ARR a mere 7 months after they launched. Lovable is on a mission to build the “last piece of software,” enabling anyone to create fully functional applications without writing code. Users enter a prompt describing the app they want, and Lovable generates the code, spins up the application, and enables real-time visual edits. Lovable uses Modal Sandboxes at massive scale to run LLM-generated code for thousands of apps in safe and isolated sandboxes. The problem: performance and reliability at scale In June 2025, Lovable had a major promotional weekend event coming up, in collaboration with Anthropic, OpenAI, and Google. This event was set to dramatically increase user activity. Unfortunately, Lovable’s existing code sandbox provider—a distributed cloud VM platform—raised concerns around scalability and operational risk. The team wasn’t confident the provider could scale fast enough to meet the expected demand, and with only this single vendor in place, Lovable was exposed to high risk of se

Discovered: Last checked: Content changed:
ArticleInternal link

“We’re actively saving 2 engineers’ worth of ongoing time”

https://modal.com/blog/quora-case-study

Open original page

Quora is building Poe, a platform where anyone can deploy a public AI chatbot. Quora uses Modal Sandboxes at scale to safely run LLM-generated code in the context of user chats.

Language: en
Indexed excerpt

All posts Back Customer Stories June 30, 2025 •3 minute read How Quora uses Modal to run thousands of Python sandboxes simultaneously Quora is a Q&A platform where users can ask, answer, and peruse questions on a variety of topics. With 400 million monthly unique visitors, it’s an invaluable contributor to the world’s knowledge-sharing. Quora uses Modal Sandboxes to securely execute LLM-generated code in Poe, their AI chatbot platform. The team shipped months earlier using Modal rather than building in-house. They’re also saving an ongoing 2 engineers’ worth of infrastructure maintenance time! Hello, Poe In 2023, Quora launched Poe, an AI chatbot platform where anyone can deploy a public chatbot. With millions of monthly active users, Poe is the default destination for many AI builders to experiment with different models. Quora has since raised $75M to keep expanding Poe. A code interpreter for Poe Many of the LLM bots in Poe can generate code, and users expected to run that code in Poe rather than copy-pasting it to their editors. The Quora team needed a way to safely execute code in Poe in a completely isolated way, keeping that code separate from both the main Quora infrastructu

Discovered: Last checked: Content changed:
ArticleInternal link

“Modal makes it easy to write code that runs on 100s of GPUs in parallel, transcribing podcasts in a fraction of the time.”

https://modal.com/blog/substack-case-study

Open original page

Learn how Substack sped up their developer iteration cycles by moving ML training and deployment to Modal from AWS SageMaker.

Language: en
Indexed excerpt

All posts Back Customer Stories May 20, 2024 •3 minute read Why Substack moved their AI and ML pipelines to Modal Substack is a popular online platform for writers to publish newsletters, with over 17k writers and $300M in paid subscription volume. Substack employs ML for various purposes, including spam detection, newsletter recommendations, audio transcription, sentiment analysis, and image generation. For nearly all these models Substack has moved both training and deployment from AWS SageMaker to Modal. The challenges deploying AI and ML before Modal Previously, Substack’s training and deployment pipelines were built on AWS SageMaker and orchestrated with Airflow. Adding or updating models was a slow and painful process for a few reasons. One, the developer experience on SageMaker was convoluted. Engineers had to navigate to the SageMaker product, create a notebook, specify machine requirements, and wait for that machine to turn on, all before a single line of code could be written. Not to mention the difficulty of juggling multiple remote environments—from the Jupyter notebooks to the SageMaker training machines to the final production infra. Two, collaborating was difficult.

Discovered: Last checked: Content changed:
ArticleInternal link

4 months faster to launch

https://modal.com/blog/suno-case-study

Open original page

Find out how Suno uses Modal to scale inference and batch pre-processing to thousands of GPUs.

Language: en
Indexed excerpt

All posts Back Customer Stories February 21, 2024 •3 minute read How Suno shaved 4 months off their launch timeline with Modal Suno uses Modal to scale inference and batch pre-processing to thousands of GPUs. With Modal, Suno was able to bring a state-of-the-art music generation model to market four months early instead of hiring a team of engineers to build and maintain infrastructure. About Suno Suno is a music generation app that can make any song you describe. Enter a simple text description—like “a deep house song about serverless infra”—and Suno makes you a song complete with vocals in seconds. Suno’s users include Grammy-winning artists, but the core user base is people experiencing making music for the first time. Microsoft recently announced they’ve partnered with Suno to bring song generation capabilities to Copilot, their AI chatbot! Avoiding past infrastructure pain Prior to starting Suno, all four founders worked at Kensho, an AI tech startup for financial data. They had personally spent significant amounts of time setting up and managing Kubernetes clusters to support their data-heavy workloads—so when they started working on Suno, they knew exactly what they did not

Discovered: Last checked: Content changed:
PageInternal link

Careers

https://modal.com/careers

Open original page
Discovered: Last checked: Content changed:
PageInternal link

Serve your own LLM API

https://modal.com/docs/examples/llm_inference

Open original page
Discovered: Last checked: Content changed:
PageInternal link

slack

https://modal.com/slack

Open original page
Discovered: Last checked: Content changed:
PageInternal link

Terms

https://modal.com/legal/terms

Open original page

AI infrastructure that developers love.

Language: en
Indexed excerpt

Runtime, the conference for engineers running AI in production. Oct. 1 in SF Register now Product Solutions Resources CustomersPricingDocs Log In Sign Up Software as a Service Agreement Effective May 2026. This Software as a Service Agreement (the "Agreement") is between the entity named below ("Customer") and Modal Labs, Inc., a Delaware corporation ("Modal"). This Agreement consists of these terms, each order form for Services that has been executed by Modal and Customer (each an "Service Order") and all exhibits and amendment of any of the foregoing. Customer consents to this Agreement by executing a Service Order or creating an account on the Service. SCOPE OF SERVICE AND RESTRICTIONS Access and Scope of Service. Subject to receipt of the applicable Fees with respect to the service specified in the corresponding Service Order (the "Service(s)"), Modal will make the Service available to Customer as set forth in this Agreement and the Service Order. Subject to Customer's compliance with the terms and conditions of the Agreement and the Service Order, Customer may access and use the Service during the period specified in the Service Order. Any such use of the Service by Customer i

Discovered: Last checked: Content changed:
PageInternal link

About

https://modal.com/company

Open original page

Modal provides high-performance, scalable computing infrastructure. It's our vision of a cloud accessible to all developers.

Language: en
Indexed excerpt

Runtime, the conference for engineers running AI in production. Oct. 1 in SF Register now Product Solutions Resources CustomersPricingDocs Log In Sign Up Making cloud developmentwork like magic We're hiring! Our founders Erik Bernhardsson and Akshat Bubna, pictured with Mugi (Akshat's dog). Not pictured: Erik's cats Lance and Stan. We started Modal with the goal to make it easier to iterate and ship applications for data, AI, and machine learning. In order to deliver the developer experience we wanted, we went deep and built our own infrastructure — including our own custom file system, container runtime, scheduler, container image builder, and much more.Customers use Modal for a wide range of use cases, including Generative AI inference, LLM fine-tuning, computational biotech, and media processing. We're focused on developer experience at our core, letting companies ship value faster without having to think about infrastructure. In a few lines of code, we let you scale from zero to thousands of CPUs or GPUs. And since pricing is entirely usage-based, you only pay for the time your code is running.Our team is based out of New York, Stockholm, and San Francisco. It includes creators

Discovered: Last checked: Content changed:
PageInternal link

Transcribe speech with Kyutai STT Stream transcripts at the speed of speech

https://modal.com/docs/examples/streaming_kyutai_stt

Open original page

This example demonstrates the deployment of a streaming audio transcription service with Kyutai STT on Modal.

Language: en
Indexed excerpt

Stream transcriptions with Kyutai STT This example demonstrates the deployment of a streaming audio transcription service with Kyutai STT on Modal. Kyutai STT is an automated speech recognition/transcription model that is designed to operate on streams of audio, rather than on complete audio files. See the linked blog post for details on their “delayed streams” architecture. Setup We start by importing some basic packages and the Modal SDK. import asyncio import base64 import time from pathlib import Path import modal Then we define a Modal App and an Image with the dependencies of our speech-to-text system. app = modal.App(name="example-streaming-kyutai-stt") stt_image = ( modal.Image.debian_slim(python_version="3.12") .uv_pip_install( "moshi==0.2.9", "fastapi==0.116.1", "huggingface-hub==0.33.5", "julius==0.2.7" ) .env({"HF_XET_HIGH_PERFORMANCE": "1"}) ) One dependency is missing: the model weights. Instead of including them in the Image or loading them every time the Function starts, we add them to a Modal Volume. Volumes are like a shared disk that all Modal Functions can access. For more details on patterns for handling model weights on Modal, see this guide. MODEL_NAME = "kyu

Discovered: Last checked: Content changed:
PageInternal link

GPU Glossary

https://modal.com/gpu-glossary

Open original page

A glossary of terms related to GPUs.

Language: en
Indexed excerpt

GPU Glossary GPU Glossary Terminal Light green Light Deploy on GPUs TABLE OF CONTENTS Home - README Device Hardware - CUDA (Device Architecture) Streaming Multiprocessor SM Core Special Function Unit SFU Load/Store Unit LSU Warp Scheduler CUDA Core Tensor Core Tensor Memory Accelerator TMA Streaming Multiprocessor Architecture Texture Processing Cluster TPC Graphics/GPU Processing Cluster GPC Register File L1 Data Cache Tensor Memory GPU RAM Device Software - CUDA (Programming Model) Streaming ASSembler SASS Parallel Thread eXecution PTX Compute Capability Thread Warp Warpgroup Cooperative Thread Array Kernel Thread Block Thread Block Grid Thread Hierarchy Memory Hierarchy Registers Shared Memory Global Memory CUDA Tile Programming Model Host Software - CUDA (Software Platform) CUDA C++ (programming language) NVIDIA GPU Drivers nvidia.ko CUDA Driver API libcuda.so NVIDIA Management Library NVML libnvml.so nvidia-smi CUDA Runtime API libcudart.so CUDA Graphs NVIDIA CUDA Compiler Driver nvcc NVIDIA Runtime Compiler NVIDIA CUDA Profiling Tools Interface CUPTI NVIDIA Nsight Systems CUDA Binary Utilities cuBLAS cuDNN CUTLASS CuTe CuTe DSL Performance - Performance Bottleneck Roofline Mo

Discovered: Last checked: Content changed:
PageInternal link

Voice chat with LLMs Build an interactive voice chat app

https://modal.com/docs/examples/llm-voice-chat

Open original page

QuiLLMan is a complete voice chat application built on Modal: you speak and the chatbot speaks back!

Language: en
Indexed excerpt

QuiLLMan: Voice Chat with Moshi QuiLLMan is a complete voice chat application built on Modal: you speak and the chatbot speaks back! At the core is Kyutai Lab’s Moshi model, a speech-to-speech language model that will continuously listen, plan, and respond to the user. Thanks to bidirectional websocket streaming and Opus audio compression, response times on good internet can be nearly instantaneous, closely matching the cadence of human speech. You can find the demo live here. Everything — from the React frontend to the model backend — is deployed serverlessly on Modal, allowing it to automatically scale and ensuring you only pay for the compute you use. This page provides a high-level walkthrough of the GitHub repo. Code overview Traditionally, building a bidirectional streaming web application as compute-heavy as QuiLLMan would take a lot of work, and it’s especially difficult to make it robust and scale to handle many concurrent users. But with Modal, it’s as simple as writing two different classes and running a CLI command. Our project structure looks like this: Moshi Websocket Server: loads an instance of the Moshi model and maintains a bidirectional websocket connection with

Discovered: Last checked: Content changed:
PageInternal link

All examples

https://modal.com/docs/examples

Open original page

How to run LLMs, Stable Diffusion, data-intensive processing, computer vision, audio transcription, and other tasks on Modal.

Language: en
Indexed excerpt

Featured Getting started Hello, world Simple web scraper Serving Web Functions Large language models (LLMs) Deploy an OpenAI-compatible LLM service with vLLM Cut Ministral 3 cold start times by 10x with snapshots Maximize tokens per second in batch processing with vLLM Serve an ultra-low-latency chatbot with SGLang Deploy Nemotron 3 Serve Inkling-Small with SGLang Serve DeepSeek-V4-Flash Efficient LLM Finetuning with Unsloth Run a multimodal RAG chatbot to answer questions about PDFs Fine-tune an LLM to replace your CEO Deploy a stateless MCP with FastMCP Images, video, & 3D Edit images with Flux Kontext Fine-tune Wan2.1 video models on your face Run Flux fast with torch.compile Fine-tune Flux with LoRA Animate images with LTX-Video Generate video clips with LTX-Video Run Stable Diffusion with a CLI, API, and web UI Audio Deploy a Moshi voice chatbot Stream transcripts at the speed of speech using Kyutai STT Make music with ACE-Step Generate speech with Chatterbox Run high throughput batched transcription with Whisper Fine-tune Whisper to recognize new words Real-time communication (WebRTC) Serverless WebRTC WebRTC quickstart with FastRTC Computational biology Design protein binder

Discovered: Last checked: Content changed:
PageInternal link

Transcribe speech in batches with Whisper Turn audio bytes into text at scale

https://modal.com/docs/examples/batched_whisper

Open original page

In this example, we demonstrate how to run dynamically batched inference for OpenAI’s speech recognition model, Whisper, on Modal. Batching multiple audio samples together or batching chunks of a single audio sample can help to achieve a 2.8x increase in inference throughput on an A10G!

Language: en
Indexed excerpt

Fast Whisper inference using dynamic batching In this example, we demonstrate how to run dynamically batched inference for OpenAI’s speech recognition model, Whisper, on Modal. Batching multiple audio samples together or batching chunks of a single audio sample can help to achieve a 2.8x increase in inference throughput on an A10G! We will be running the Whisper Large V3 model. To run any of the other HuggingFace Whisper models, simply replace the MODEL_NAME and MODEL_REVISION variables. Setup Let’s start by importing the Modal client and defining the model that we want to serve. from typing import Optional import modal MODEL_DIR = "/model" MODEL_NAME = "openai/whisper-large-v3" MODEL_REVISION = "afda370583db9c5359511ed5d989400a6199dfe1" Define a container image We’ll start with Modal’s baseline debian_slim image and install the relevant libraries. image = ( modal.Image.debian_slim(python_version="3.11") .uv_pip_install( "torch==2.5.1", "transformers==4.47.1", "huggingface-hub==0.36.0", "librosa==0.10.2", "soundfile==0.12.1", "accelerate==1.2.1", "datasets==3.2.0", ) .env({"HF_XET_HIGH_PERFORMANCE": "1", "HF_HUB_CACHE": MODEL_DIR}) ) model_cache = modal.Volume.from_name("hf-hub-cac

Discovered: Last checked: Content changed:
PageInternal link

Articles

https://modal.com/articles

Open original page

Short form articles

Language: en
Indexed excerpt

Runtime, the conference for engineers running AI in production. Oct. 1 in SF Register now Product Solutions Resources CustomersPricingDocs Log In Sign Up Modal Articles All Posts Serverless GPUs Frameworks and Tools LLMs Image and Video Models Audio Models AI Agents Data Infrastructure Embedding Models December 16, 2025 Launch a chatbot that runs inference on Modal using the Vercel AI SDK How to build a chatbot running open source models on Modal with a Vercel AI SDK and AI Elements frontend.November 3, 2025 How to deploy vLLM A step by step tutorial of deploying vLLM, a popular open-source LLM inference engine, on GPUs in the cloud.November 1, 2025 Choosing between Whisper variants: faster-whisper, insanely-fast-whisper, WhisperX Compare the top open-source Whisper variants for ASR tasks.October 31, 2025 The Best Jupyter Notebooks Products in 2025 Compare the top cloud Jupyter notebook products available in 2025, focusing on five vendors: Google Colab, AWS SageMaker, Deepnote, Databricks, and Modal Notebooks.October 30, 2025 Top embedding models on the MTEB leaderboard Overview of the top-ranking embedding models on the MTEB leaderboardOctober 20, 2025 Top 5 serverless GPU provide

Discovered: Last checked: Content changed:
PageInternal link

Deploy OpenCode agents in a cloud Sandbox

https://modal.com/docs/examples/opencode_server

Open original page

This example demonstrates how to run OpenCode remotely and connect to it from your local terminal or browser.

Language: en
Indexed excerpt

Run OpenCode in a Modal Sandbox This example demonstrates how to run OpenCode remotely and connect to it from your local terminal or browser. Combine self-hosted OpenCode with serving a big, smart model on Modal and you’ve got “coding agents at home”! Coding agents are most useful when they have context and tools. By default, this script clones the Modal examples repo and gives the agent access to your Modal credentials, so it can run and debug examples (including this one!). Meta. Set up OpenCode on Modal import argparse import os from pathlib import Path import modal MINUTES = 60 HOURS = 60 * MINUTES OPENCODE_PORT = 4096 DEFAULT_GITHUB_REPO = "modal-labs/modal-examples" First, we define a Modal container Image with OpenCode installed. def define_base_image() -> modal.Image: image = ( modal.Image.debian_slim() .apt_install("curl", "git", "gh") .run_commands("curl -fsSL https://opencode.ai/install | bash") .env({"PATH": "/root/.opencode/bin:${PATH}"}) ) # We also bring the global default OpenCode configuration along for the ride. CONFIG_PATH = Path("~/.config/opencode/opencode.json").expanduser() if CONFIG_PATH.exists(): print("🏖️ Including config from", CONFIG_PATH) image = image

Discovered: Last checked: Content changed:
PageInternal link

Deploy a TTS API with Chatterbox Serve text-to-speech with Chatterbox to generate natural audio from text

https://modal.com/docs/examples/chatterbox_tts

Open original page

This example demonstrates how to deploy a text-to-speech (TTS) API using the open source model Chatterbox Turbo on Modal.

Language: en
Indexed excerpt

Create a Chatterbox TTS API on Modal This example demonstrates how to deploy a text-to-speech (TTS) API using the open source model Chatterbox Turbo on Modal. Chatterbox Turbo is a state-of-the-art TTS model that can generate natural, expressive speech that rivals proprietary models. Prompts can include paralinguistic tags like [chuckle], [sigh], and [gasp]. Chatterbox also support voice cloning by passing a short (about 10 seconds) audio prompt of the target voice. Check out Resemble AI’s website or the Chatterbox Github repo for more details. Setup Import modal, the only required local dependency. import modal Define a container image We start with Modal’s baseline debian_slim image and install the required packages. chatterbox-tts: The TTS model library fastapi: Web framework for creating the API endpoint “peft”: Required for properly loading the model image = modal.Image.debian_slim(python_version="3.10").uv_pip_install( "chatterbox-tts==0.1.6", "fastapi[standard]==0.124.4", "peft==0.18.0", ) We’ll also use Chatterbox’s provided set of voice prompts which you can download here. Unzip the file and upload it to a modal.Volume called chatterbox-tts-voices with the following CLI co

Discovered: Last checked: Content changed:
PageInternal link

Batch

https://modal.com/products/batch

Open original page

AI infrastructure that developers love.

Language: en
Indexed excerpt

Runtime, the conference for engineers running AI in production. Oct. 1 in SF Register now Product Solutions Resources CustomersPricingDocs Log In Sign Up Modal Batch Run Batch jobs with 1 line of code Spawn 1 million jobs in seconds. Powered by Modal's hyper-elastic compute infrastructure. Get Started Read the docs “With the old in-house systems, we'd have to tune number of workers, instance size, parallelization strategy, all this stuff, which was very time-consuming and not directly generating business value. Modal magically handled all that.” Samarth Goel, ML Engineer “Processing external quantum mechanical datasets comes with unique challenges. Jobs can fail in numerous ways—from low-level errors thrown by the underlying analysis packages, to transient issues communicating with our storage server. Modal's retry mechanism and batching primitives have made our data pipeline much more robust.” Liz Decolvenaere, Quantum Chemical Engineer “Modal's autoscaling capabilities give us the best of both worlds. We get the power of a massive GPU cloud when we need it, without the complexity of managing spot instances or cloud-specific infrastructure.” Georg Kucsko, CTO DEFINE IN CODE BUILT-

Discovered: Last checked: Content changed:
PageInternal link

Events

https://modal.com/events

Open original page

View and follow events from Modal on Luma. AI infrastructure that developers love

Language: en
Indexed excerpt

(function () { var applied = []; function define(target, name, value) { Object.defineProperty(target, name, { value: value, writable: true, configurable: true, enumerable: false }); applied.push(name); } // Chrome 93. Unguarded in a timezone module and in the react-markdown chain. if (!Object.hasOwn) { define(Object, "hasOwn", function hasOwn(target, key) { if (target === null || target === undefined) { throw new TypeError("Cannot convert undefined or null to object"); } return Object.prototype.hasOwnProperty.call(Object(target), key); }); } // Chrome 92. Unguarded throughout web-vitals. function at(index) { var target = Object(this); var length = target.length >>> 0; var relative = Math.trunc(Number(index)) || 0; var resolved = relative < 0 ? length + relative : relative; return resolved < 0 || resolved >= length ? undefined : target[resolved]; } if (!Array.prototype.at) { define(Array.prototype, "at", at); } if (!String.prototype.at) { define(String.prototype, "at", at); } // Chrome 85. Unguarded in query-string's parser, so it runs on any URL with // a query — which is how this first reached us. if (!String.prototype.replaceAll) { define(String.prototype, "replaceAll", function

Discovered: Last checked: Content changed:
PageInternal link

Create custom art of your pet

https://modal.com/docs/examples/diffusers_lora_finetune

Open original page

This example finetunes the Flux.1-dev model on images of a pet (by default, a puppy named Qwerty) using a technique called textual inversion from the “Dreambooth” paper. Effectively, it teaches a general image generation model a new “proper noun”, allowing for the personalized generation of art and photos. We supplement textual inversion with low-rank adaptation (LoRA) for increased efficiency during training.

Language: en
Indexed excerpt

Fine-tune Flux on your pet using LoRA This example finetunes the Flux.1-dev model on images of a pet (by default, a puppy named Qwerty) using a technique called textual inversion from the “Dreambooth” paper. Effectively, it teaches a general image generation model a new “proper noun”, allowing for the personalized generation of art and photos. We supplement textual inversion with low-rank adaptation (LoRA) for increased efficiency during training. It then makes the model shareable with others — without costing $25/day for a GPU server— by hosting a Gradio app on Modal. It demonstrates a simple, productive, and cost-effective pathway to building on large pretrained models using Modal’s building blocks, like GPU-accelerated Modal Functions for compute-intensive work, Volumes for storage, and Web Functions for serving. And with some light customization, you can use it to generate images of your pet! You can find a video walkthrough of this example on the Modal YouTube channel here. Imports and setup We start by importing the necessary libraries and setting up the environment. from dataclasses import dataclass from pathlib import Path import modal Building up the environment Machine le

Discovered: Last checked: Content changed: