Build with AI

How to build an AI subtitle generator for your business with Replit

Build a subtitle studio in one Replit workspace, from the first message to the day people start captioning in it. Transcription on the viewer’s own machine, a timeline for fixing and styling the lines, translation on your own AI key with a limit per account, and exports as subtitle files or a captioned video, with everything saved to a database set up for you and published from the same tab.

September 2026 · 42 min read · Updated September 2026

Replit

$ Build me a subtitle studio: videos transcribed in the viewer’s own browser into timed captions, a timeline for fixing and styling the lines, translation into other languages on my own key, and export as a subtitle file or a captioned video. Add a database for the videos and captions, and keep each person’s work to themselves.

  • Database provisioned in the workspace
  • Speech model and editor built
  • Ready for you to publish
You describe it, Replit builds it
Start here

Overview & core architecture

An AI subtitle studio turns a spoken video into timed, editable, translated captions, and it does the expensive part, the transcription, on the visitor’s own computer rather than on a server you pay by the minute.

The usual version of this is a service. You upload a video, a server somewhere listens to it, and captions come back with a bill attached, priced by the minute of audio. Every product in the comparison further down works that way, and the minute is the line they all meter.

This version downloads a speech model into the browser once and runs it there, in the background, while the page stays usable. A minute of transcription costs you nothing, and only the finished text and the video the person chose to upload reach your database. What you build instead is everything around that moment: a library, a timeline editor, caption styling, translation through your own AI key, and an export that produces a subtitle file or a captioned video.

The honest trade is written into every page of this guide. The model that runs in a browser is the smallest one, its accuracy sits below what the rented services return, and the first run downloads it, which the app says can take thirty to sixty seconds. A larger model buys accuracy at the cost of a longer download. Deciding where you sit on that line is the first decision below.

The timeline editor is the product

A raw transcript is a start, and nobody ships one. The screen where a person drags a caption to the right second, splits a line that runs too long and sees the result over the video is what they judge the app on, and it is where most of the build time goes.

Transcription minutes cost nothing

The speech model runs on the visitor’s own processor, so a thousand hours of video transcribed this month costs you the same as one. What you pay for is storage for the videos people keep and the translation calls that leave your app.

Translation is the only metered line

Translating a finished track goes through your own AI key at the provider’s per-token price, a few cents for an hour of speech at the default model’s rates. It is the one line that grows with use, so the app logs every call and caps each account.

What you are building

Essential subtitle studio components

Six building blocks make up the studio, and the transcription is the one to prove first. Each is something you can ask your AI coding tool to build or rework in plain words.

01

In-browser speech-to-text

A speech model that downloads once and runs on the visitor’s own machine, in the background, so the page stays usable while it listens. It returns lines with a start and an end time, which is what everything else on this list is built on. No server hears the audio and nobody bills you for the minute.

02

Timeline caption editor

Every line on a track under the video, draggable to retime, editable in place, splittable when it runs too long and mergeable when it is choppy. Undo and autosave, because a caption editor is used for hours at a stretch.

03

Caption styling & live preview

Font, size, colour, background and position, shown over the actual video rather than in a settings form. What the person sees here is exactly what the export produces.

04

One-click translation

The finished track sent to your AI provider a batch of lines at a time, with every timestamp kept, and progress shown while it runs in the background. Each translated track becomes its own export.

05

Subtitle file & captioned video export

Three plain subtitle formats for uploading to a video platform, and a captioned video recorded in the browser for the platforms that cannot take a separate file. Export is where a subtitle tool either feels finished or does not.

06

Video library, quotas & admin console

Every upload with what has been done to it, hourly limits per action so one account cannot run up your translation bill, a demo account that wipes itself, and an admin view over users, videos and usage.

Build vs buy

Own the subtitle studio or rent it by the minute

Most subtitle apps charge you based on audio minutes processed. They set monthly limits, charge per team member, and bill extra when you run out. Running the AI inside the user’s browser removes these minute limits entirely instead of just reducing them: that is the main difference to keep in mind.

Build your own

Own your platform completely. Transcription runs on your users’ devices, videos stay in your own database, and your only recurring cost is optional translation. With modern AI coding tools, building this takes a few weeks instead of months.

  • Zero minute costs: Transcribe unlimited audio because processing happens directly on the user’s device
  • Cheaper translations: Pay wholesale rates for AI translations and easily set monthly spending caps per user
  • 100% custom branding: Keep your own brand, logins, and design without third-party restrictions
  • Full feature control: Decide your own caption styles, export formats, and limits without paying for plan upgrades
  • Complete data ownership: Store videos and captions in your database with full control over privacy and deletion
  • Total code freedom: Own all source code outright and host your application on any server you choose

Rent the minutes

Rev · Descript · Kapwing · Happy Scribe

Renting an existing tool gives you higher accuracy, an editor ready to use today, and automatic software updates. However, you pay a fee for every minute of video processed, plus extra monthly fees per user.

  • Higher accuracy out of the box: Powerful cloud models offer higher accuracy with no initial setup or load time
  • Ready to use today: Launch immediately, with options for human proofreading when accuracy matters most
  • Zero maintenance: External companies handle all system updates, model upgrades, and bug fixes
  • Pay-per-minute pricing: Every provider meters audio minutes, and many charge additional fees per user seat
  • Use-it-or-lose-it plans: Monthly minute allowances expire whether your team actually uses them or not
  • Restrictive free tiers: Free plans only cover occasional usage, forcing active teams onto expensive paid tiers
RevFree for "45 AI transcription & caption minutes/month (English only)", then Essentials at "$25.49 per seat/month" billed annually or "$29.99 per seat/month" monthly, with "5,000 AI transcription & caption minutes/seat/month"

The clearest example of the seat-plus-allowance meter. A single person captioning under 45 minutes a month never pays, and a team pays per seat every month whether or not anybody captioned anything. The Pro tier is "$47.99 per seat/month" annually for "10,000 verbatim AI transcription minutes/seat/month", and human captions start at "$1.99 /min.", which is the rate to compare the browser model against when accuracy is the whole point.

rev.com · checked September 2026

DescriptFree with "60 minutes (1 hr) / month" at 720p with a watermark, then Hobbyist at $16 a month billed annually ($24 monthly) for "10 media hours / month", Creator at $24 annually ($35 monthly) for "30 media hours / month", and Business at $50 annually ($65 monthly) for "40 media hours / month", per person

An editor with transcription inside it, priced by media hours per person per month. The allowance is generous for one creator and resets whether it was used or not, so the bill is a function of headcount rather than of output. Descript is the row to read if what you actually want is a video editor that happens to caption.

descript.com · checked September 2026

KapwingFree for "Up to 50 minutes" of auto-subtitling with exports capped at "4 minutes" and a watermark, then Pro at "$16" per member a month billed annually ("$24 billed monthly") for "Up to 1,000 minutes per month" of subtitles and "Up to 500 minutes per month" of translation, and Business at "$50" annually ("$64 billed monthly") for 4,000 and 2,000 minutes

Two meters at once, per member: subtitle minutes and translation minutes are separate allowances, and the free tier caps the exported video at four minutes. It is the closest product to this template in shape, which is what makes it the fairest comparison. Everything this template does for one account with no allowance, Kapwing does per member with two.

kapwing.com · checked September 2026

Happy ScribeA "10-minute free trial", then Basic at "$17 / month" ($8.50 annually) for "120 minutes of AI Transcription, Subtitling, and Translation per month", Pro at "$29 / month" ($19 annually) for 600 minutes, Business at "$89 / month" ($59 annually) for 6,000 minutes, and top-ups at "$0.20/min"

The row with the cleanest per-minute figure: once the allowance is gone, each extra minute is "$0.20/min", so an hour of video is $12 in top-ups. Multiply that by the hours your users would actually caption in a month and compare it to zero, then remember the honest half of the comparison: the rented model is more accurate than the one that fits in a browser tab.

happyscribe.com · checked September 2026

Rule of thumb: if you caption a handful of your own videos a month, do not build this. Rev’s 45 free minutes or Descript’s free hour covers you, and the rented models are more accurate than the one that runs in a browser. If you caption for clients, for a team or for other people’s uploads, the minutes are the bill, and owning the studio turns a meter into a fixed cost. The honest middle case is a product whose users need high accuracy on difficult audio, and there the right answer is to build the studio and swap in a larger speech model, accepting a longer first download in return.

No dev needed

Why build with Replit

Skip the speech-recognition engineering. Describe the library, the editor and the translation you want, say who may use them, and every video, caption track and export is saved under its owner’s account without you writing server code.

Replit’s Agent handles the database, the access rules, and the hosting from one chat, in the same workspace the app ends up living in:

The build loop
1

Describe

Tell the Agent what to build, in plain language.

2

Watch

It writes the code, sets up the database, and shows the app running live.

3

Try it

Use the real app in the preview rather than a mockup.

Publish

Take it live on Replit’s own hosting, or ask for the next change.

Loop back to Describe

The build and the place it ends up running are the same workspace throughout, so there’s no separate hosting account to set up later.

One workspacebuilds, runs, and hosts it

Replit is the one tool here that also deploys what it builds. Publishing takes the same project live on Replit’s own infrastructure, with a working domain, uptime monitoring, and security scanning included.

Managed Postgres with 20GB included free

Ask the Agent to add a database and it creates the schema and wires your app to it. What you get is a real, fully-managed SQL database rather than a mocked one.

Up to 10 Agent sessions in parallel (Pro)

Core allows up to 2 parallel Agent sessions and Pro allows up to 10, so more than one part of the app can be worked on at the same time.

What it costs

Pay a developer, or do it with AI

When you own your subtitle app, there are no per-minute processing fees. Your only real costs are constructing the code, video storage, and optional translations. Here is a clear breakdown of hiring a developer versus building it yourself using AI.

Hire a developer

Custom build, from scratch
Developer
~$11k-$42k
Supabase (backend)
Free tier · $25/mo (Pro plan)*
Hosting
$0 free tier
AI translation
Per token on your own key
Build time
~210 hrs of their work

~$11k-$42k to build, then from $25/mo after launch

Our ~210-hour estimate, costed against the rate survey linked below, whose bands run from $45-$75/hr for North American contractors up through $100-$150+/hr for senior US developers, before the 20-40% an agency adds, which brackets the range at roughly $50/hr and $200/hr. Most of those hours are the speech model and the timeline, and neither shows in a screenshot. Then read the translation row: it is the only line that grows with use, and at the default model’s rates an hour of speech costs a few cents.

Build it with Replit

From scratch, with Replit
Replit
Free (daily credits) to $25/month (Core) or $100/month (Pro)
Database (built-in Postgres)
Free to start · 20GB included
Hosting (Replit Deployments)
Billed separately, on top of the plan
Your time
~101 hrs

Free to try the idea, ~$25-$100/month on Core or Pro while you build a real one, then whichever plan (plus any deployment cost) you keep using

Replit’s plan price and its credit grant are the same number, not a subscription plus a separate credit purchase: Core is $25/month for $25 of monthly credits (or $20/month billed annually), Pro is $100/month for $100 of monthly credits (or $95/month annually). Once you publish, Replit bills hosting through its own Deployments separately, on top of whichever plan you’re on. Budget for it as a second line, not folded into the $25 or $100.

* Unlike most apps in this catalogue, this one keeps large files: the videos people upload live in storage until you delete them. Supabase Free includes 1 GB of file storage and 5 GB of egress a month, and pauses a project after a week without activity. Pro, from $25/mo, includes 100 GB of storage and 250 GB of egress, then charges $0.0213 per GB stored and $0.09 per GB served, and keeps a daily backup for 7 days. A retention rule that deletes a video a month after its last export is the cheapest decision on this page.

Prices and rates from supabase.com, developex.com and replit.com, checked September 2026.

Plan first

Decide before you build

Six decisions to make before writing code. Three of them decide what your users’ videos cost you to keep, and one decides how long they wait for the first caption.

01

Which speech model, and how long will people wait for it?

The smallest model downloads in under a minute and gets ordinary speech mostly right. Larger ones are more accurate and take longer to arrive on the first visit. Decide now which you ship and what the screen says while it loads, because a silent wait looks like a broken app.

02

How long a video will you accept?

The transcription runs in a browser tab with finite memory, and the template already warns itself above thirty minutes of audio. Decide the ceiling by length or by file size, say it before the upload rather than after, and refuse politely. A limit stated up front reads as a feature, and a tab that dies at 80% reads as a broken product.

03

Do you keep the videos, and for how long?

Uploaded videos are the only large thing in this app, and storage is the one bill that grows quietly. Decide whether a video lives until the person deletes it, expires a month after its last export, or is never kept once the captions exist. Build the deletion now rather than after the storage bill arrives.

04

Which languages, and who pays for translation?

Transcription costs $0, but AI translation costs money for every translated word. Decide upfront which languages to support and set strict usage limits per account (for example, 5 free translations per month). Adding a hard limit now prevents surprise bills later when user activity spikes.

05

Is video storage public or private?

By default, uploaded videos are accessible to anyone with the link so the browser can process exports smoothly. Switching to private storage with expiring links keeps user content secure, but requires setting up server-side rendering for video exports. Decide on your privacy model before your first user uploads sensitive content.

06

Who gets an account, and what does a demo get?

Open sign-up, invitations, or a demo account that anyone can try. The template gives demo accounts their own role and wipes them after an hour of inactivity, which is the right shape for a public trial. Whatever you choose, limits attach to the account rather than the visitor, because accounts are free to create.

Approaches

Comparing your build options

Building a box that shows a transcript is fast. The hard part is everything behind it: a speech model running in the visitor’s browser, a timeline that stays in step with the video, translation that keeps every timestamp, and a captioned file at the end. Here are three ways to build the exact same product.

~210 hrsBuilding by hand

Getting a speech model to run inside a browser tab, without freezing the page, is where the first weeks go. Then the caption you dragged has to stay in step with the video, the translated line has to land at the same second as the original, and the captioned file has to come out of a browser that was never designed to render one.

~155 hrsGeneric UI starter kit

A kit gives you a login page, a file table and a settings screen. It has never heard of a caption, a timestamp or a speech model, so the editor, the transcription and the export, which are the whole product, start from nothing.

~101 hrsAI-powered development with Replit

You ask the Agent for a piece at a time and it writes the database, the server and the editor in the same workspace that ends up hosting the finished app. The check that matters: sign in as a second account on the published address and confirm the first account’s videos never appear.

Interactive calculator

Estimate your exact build timeframe

Customize your feature list below to see how build time changes. If you only need subtitles in one language, or nobody will export a captioned video, uncheck those rows to reduce the estimate.

What your subtitle studio needs

Your estimate

101 hrs

start to finish

Based on the 7 of 7 features you’ve selected, plus ~21h of groundwork. Toggle any on the left to watch the number move, and open the groundwork row to untick what you have already, such as a database that is already running or going live if you are only building a mock-up for now.

A rough estimate, not a quote. Real time depends on how much you customize and how clean your data is.

Setting up your workspace

Let’s set up the tools you need

Replit runs entirely in the browser, and it’s the one tool here that also hosts what you build, so no GitHub account is required first. Before step 01: an account and a plan. A database comes later, the moment your app actually needs one, and GitHub whenever you want a copy of the code outside Replit.

1

Replit account

Cost: Free

Sign up and you land in a workspace with an Agent chat, the code, and a live preview side by side, with nothing to install.

Sign up for Replit
2

Replit subscription

Cost: Free (daily credits), then $25/month (Core) or $100/month (Pro)

Starter’s free daily credits are enough to try an idea, not to finish one. Core is $25/month billed monthly, or $20/month billed annually, for $25 of monthly credits and up to 2 parallel Agent sessions. Pro is $100/month monthly, or $95/month annually, for $100 of monthly credits, more collaborators, and access to the strongest available models. The price you pay and the credits you get are the same number on both plans, so you have no separate subscription-plus-credits split to work out.

Compare Replit plans
3

Replit database (Postgres)

Cost: Free to start · 20GB included

Every Replit app includes its own managed Postgres database with 20GB of free storage. Ask the Agent to add one and it creates the schema and connects your app to it, with no separate account to create anywhere else.

Replit’s built-in database, Replit docs
Optional
4

GitHub connection

Cost: Free

Not needed to start, and not needed as an undo either, because Replit checkpoints the whole workspace as the Agent works. Connect a repository from the Git pane, free on every plan, and a copy of the real code lives outside Replit under your own account. Worth doing once the project is one you would hate to lose.

Using the Git pane, Replit docs

The first two are all you need to start. Everything here stays inside the one browser tab, the app included once you publish it, and the GitHub copy is the one deliberate exception.

Step by step

Build your subtitle studio, one Agent message at a time

Nothing installs and nothing gets connected from outside. The Agent writes the files, runs the commands, and provisions the database as it goes. One thing to keep in mind throughout: the speech model runs in the viewer’s browser, not in this workspace, so the server here stays small.

  1. 01

    Start the app, ask for the database, and set the rules

    The database comes free with the workspace, so it goes in the opening message. So do the two constraints that matter most on this build.

    PromptSet up the project and the database
    Set up a React 18 + Vite + TypeScript app with Tailwind, and add a Postgres database to this Repl, the one Replit provisions, not anything outside. Write a short notes file at the project root saying this is a subtitle studio, fixing the vocabulary as videos, tracks, captions and exports, and recording two standing rules: transcription always happens in the viewer’s browser and is never sent to this server, and the only call that leaves the app is translation through one server route. Keep the database connection details in Replit’s Secrets rather than in the code.

    The rules are worth writing into the notes file rather than a single message. A server-side transcription service is a natural suggestion for any assistant, and it would put back the cost this build exists to avoid.

  2. 02

    Prove the speech model in the live preview

    One throwaway page, one captioned clip. Do it before anything depends on it, and roll back to a checkpoint if it goes sideways rather than unpicking it.

    PromptGet the speech model running in the browser
    Add @huggingface/transformers and build a single page that proves in-browser transcription works: pick a local video, press a button, and get a list of lines with start and end times. Run onnx-community/whisper-tiny in a Web Worker so the page stays responsive, with device set to webgpu and a fallback to wasm, decode the audio in the browser and feed it to the model in thirty-second chunks, and show the model download and the transcription as two separate progress bars. Then tell me how long the first load takes in the preview and how it behaves on a twenty-minute file.

    Check this in the live preview rather than trusting the description. A speech model in a browser tab is exactly the kind of thing that looks fine in code and behaves differently in a real tab.

  3. 03

    Add sign-in and scope every row to its owner

    Replit’s Postgres is a plain database with no login wired into it, so the Agent writes both halves: sessions on the server, and policies the server’s identity feeds.

    PromptAdd sign-in, tables, roles and access rules
    Add email-and-password sign-up, login, logout and a server-side session, with a profile row per account. Then the tables. Videos owned by one account with the file’s storage path, duration and detected language. Subtitle tracks, one per video and language, holding captions as timed segments plus a status and progress for translation. Usage records, an activity log, a settings table and a rate-limits table keyed by account and action. Store uploaded videos in the workspace’s object storage rather than the database, in a folder per account, and serve them only through a route that checks the owner. Add a roles table giving each account one of visitor, user, admin or demo, in its own table, never on the account record. Then turn on row-level security across every table: have the server set the current account and role as a session-local setting at the start of each request, write the policies against that setting so a person reaches only their own rows and an admin reaches everything, and show me how to verify a second account comes back empty for the first account’s videos.

    There is no ready-made login handing the database a user id here, so the server is what knows who is asking. Build that first and the policies have something to read.

  4. 04

    Build the library and the editor

    The screen that carries the product. Check the preview after each change here rather than stacking several and finding out which one broke the timeline.

    PromptBuild the library and the editor
    Build the library: an upload that stores the video and creates its row, a list of the signed-in person’s videos with their status, and a detail page that runs the transcription from step 02 and saves the result as the video’s first track. Then the editor: the video playing above a horizontal timeline of captions, each draggable to retime and editable in place, with split and merge, keyboard shortcuts, undo and redo, and autosave shortly after the last change. Drive all of it from one piece of caption state so the timeline, the list and the preview can never disagree, and add a styling panel whose changes show over the video immediately.

    This is the step where checkpoints earn their keep. A timeline that felt right two messages ago and does not now is a one-click problem rather than an afternoon.

  5. 05

    Add translation and the exports

    Translation is the first and only call that leaves the app, so it gets a route with the key handling and logging built in. Then the four ways a finished track gets out.

    PromptAdd translation and export
    Add one server route that calls my AI provider with the key held in Replit’s Secrets, refuses any caller without a session, and logs every call with the account, the video and the token counts. Add a "Translate" action on a track: send its captions to that route ten at a time with their timestamps, ask for the same lines in the chosen language with the timestamps untouched, return immediately, write progress to the new track as batches finish, and poll it from the editor. Then exports: .srt, .vtt and .txt from any track with no AI involved, and a captioned-video export in the browser that draws the styled captions over the video on a canvas and records a WebM, served the video through our own route so it stays private.
  6. 06
    Destination

    Add limits, admin and the demo account, then publish

    Finish with the pieces that keep the bill and the trial honest, then publish, checking who the app is visible to before that first release, since Publishing is what sets it.

    PromptAdd limits, admin, demo, and publish
    Add the last three pieces. Rate limits of ten uploads, twenty transcriptions, thirty translations and fifty exports per account per hour, counted in the database rather than in one process’s memory, with quota costs read from the settings table. An admin console over users, videos, usage and settings, restricted to the admin role. And a demo role that any sign-up ending in @demo.com receives automatically, with a job that wipes that account’s videos, tracks, logs and usage after an hour of inactivity. Then help me test the whole flow: sign up as a demo account and caption a short clip, translate it, export a .srt and a captioned video, then sign up as a second real account and confirm it sees none of the first account’s videos. When it holds up, walk me through Publishing: which deployment type fits an app whose heavy work happens in the viewer’s browser, and who the app should be visible to.
Authentication & security

Protecting user content and translation keys

A subtitle studio manages three critical assets: your users’ videos, their saved captions, and your paid translation keys. Follow these essential security rules before launching.

Secure login & password protection

There is no hosted identity service in this workspace, so sign-up, sign-in, sessions and password resets are code on your own server rather than a product you switch on. Use a well-known library rather than writing password handling yourself, and treat the session as the thing that decides everything else in this list.

User access levels & safe demo mode

Visitor, user and admin, plus a demo role that any sign-up with a demo address receives automatically. A demo account can do everything a real one can, and a scheduled job wipes everything it did after an hour of inactivity, which is how a public trial stays harmless.

Private user data protection

A video, its tracks, its usage records and its activity log belong to one account, and an admin reaches all of them. The policies apply that on every read and write, so a screen that forgets to filter still cannot show one person another person’s captions.

Automatic security checks

Your server declares who is asking by setting a session variable on the connection before each query, and the row-level policies read it. Two things follow. The declaration has to be set on every request rather than once at startup, because connections are reused between people. And a policy reading a variable nobody set does not complain, it just admits everybody.

AI access for logged-in users only

The translation route refuses a caller without a session, so an anonymous visitor cannot spend your key by finding the address. Check this first on any route you add that calls a model, because an open AI endpoint is an invoice rather than a data leak.

Abuse protection & spending limits

Ten uploads, twenty transcriptions, thirty translations and fifty exports an hour per account. An autoscaling deployment runs more than one copy of your server under load, so a counter held in a variable is per copy and resets whenever a new one starts, which is precisely when the app is busiest. Count in Postgres, where every copy sees the same number.

Private video & file protection

Uploaded videos are files rather than rows, and they need their own access check. Serve them through a route that confirms the owner rather than handing out a storage link that works for anyone holding it. The captioned-video export reads the video through that same route, so it stays private without breaking.

Paid API key security

The provider key lives in Replit’s Secrets tool, readable by your server and absent from anything the browser downloads. Then the part people skip: the caption text goes to your provider for translation. Say so to whoever signs in, and read your provider’s terms on what it does with what it receives.

Data backups & recovery

The Agent checkpoints as it works (files, configuration and optionally the database) so a bad change is a click back rather than an afternoon. Check what your plan retains for this workspace’s Postgres before you rely on it, and take your own dump on a schedule if the answer is thinner than you assumed, because a lost afternoon of caption edits cannot be regenerated.

PromptCheck who can see what
Review the access rules (row-level security policies) on every table. For each one, tell me in simple terms who can view, add, edit, and delete records, confirm that people can only reach their own data while the right roles can reach more, and flag anything left open that shouldn’t be.

Paste this into the chat before launch so the Agent checks nobody can see data they shouldn’t.

One rule outranks everything above it: the database connection and your AI provider key belong in Secrets only: never in anything the browser downloads, and never in a repository. If either gets out, treat it as compromised and rotate it the same day.

Workflow rules

What speeds the build, and what slows it

Speeds the build

  • One small, specific request per message, checked in the live preview before the next one
  • Letting the Agent provision the database from the chat instead of wiring one up by hand
  • Running two Agent sessions in parallel on unrelated parts of the app, once your plan allows it
  • Rolling back to a checkpoint the moment a change goes wrong, instead of unpicking it by hand
  • Reviewing what Publishing changed, meaning the domain, who can reach the app, and the machine it runs on, before the first release

Slows the build

  • Asking for the whole app in one message instead of one piece at a time
  • Building screens for data that isn’t in the database yet
  • Running unrelated Agent sessions against the same files at the same time
  • Letting several risky changes stack up before checking whether any of them actually broke something
  • Publishing without checking who the app is visible to first
Keeping a safe copy

Connecting GitHub (Optional)

Every change in Replit is saved automatically without any technical setup. Checkpoints are the day-to-day undo. You only need to link GitHub if you want a private copy under your own control or plan to hand the codebase over to external developers.

Every milestone is already saved

Replit’s Agent creates a checkpoint automatically at key points as it works: a full snapshot of the files, the configuration, and even the AI conversation itself, not just the code.

Checkpoints and rollbacks, Replit docs

Rolling back restores the whole workspace

One click returns your project to an earlier checkpoint (files and configuration together, and optionally the database), which is broader than a typical code-only undo, so a rollback after real data has changed is worth a second look before you confirm it.

GitHub keeps a copy outside Replit

Connect a repository from the Git pane, free on every plan, and stage, commit, and push changes back to GitHub with a click, or pull in anything changed outside Replit.

Using the Git pane, Replit docs

It’s also how an existing project gets in

Point Replit at a public repository’s URL for a fast import, or use the guided import for a private one. Either way, Replit detects the stack and installs everything on its own.

Import from a provider, Replit docs

You rarely type git commands

The Git pane’s buttons cover staging, committing, and pushing. If you’d rather type them yourself, the workspace Shell stays in sync with whatever the pane just did.

Going live

Going live without external hosting

Skip third-party hosting and external database setup. Your database, backend, and domain live in the same Replit workspace. Just click Publish to go live.

HostBest forNotesFree tier
AutoscaleMost subtitle studiosGrows with traffic and shrinks to nothing when nobody is captioning. It suits this build, because the speech model runs in the viewer’s browser and the server only ever stores files, serves them back and runs the occasional translation. Confirm the hourly limits are counted in the database, since more than one copy of the server can be running.Metered - billed with your plan
Reserved VMAlways-warm responseDedicated compute that never sleeps, so nobody waits on a cold start before their library opens. Less urgent here than on apps that do their heavy work on the server, and worth it once a team uses the studio all day.By machine size - billed with your plan
ScheduledRetention and the demo sweepRuns on a timer rather than answering requests, which is the right shape for two jobs this build has: deleting videos past their retention date, and wiping demo accounts after an hour of inactivity. It publishes separately from the app.Metered - billed with your plan
StaticNot this appFiles only, with no server behind them. Closer to viable here than on most builds, since transcription genuinely runs client-side, but accounts, stored videos, translation and the access rules all need a server.Metered - billed with your plan

All four are Replit rather than a third party, so the choice is shape rather than vendor, and Autoscale suits most of these apps. One thing specific to this build: the speech model does not touch your deployment at all, because the browser fetches it from its own home on the first visit, so the server stays light no matter how much video people caption. Review who the app is visible to before that first release.

The first transcription downloads the speech model, and on a slow connection that looks like nothing is happening. The template says thirty to sixty seconds and gives up after two minutes. Before you publish, try it on a phone on mobile data rather than at your desk, keep the progress message on screen the whole time, and decide what the app says if the download fails rather than letting a spinner run forever.

Database & backend

Your database is already part of the workspace

Nothing to connect. Captions and usage are small rows of text, so the database stays tiny. The videos people upload are the large thing, and they belong in the workspace’s object storage rather than in the database.

ServiceBest forNotesFree tier
Replit PostgresData, built inManaged Postgres with 20GB included free, provisioned from the same chat that builds the app, and since it is ordinary Postgres the access policies from step 03 and the rate-limits table both work exactly as they would anywhere else. What it does not include is a hosted identity service, which is why your own server declares the account before any policy can act on it. Keep the videos in object storage and serve them through your own route, so a link to one video only works for its owner.Free to start · 20GB included
AI workflows

Add AI capabilities in one simple step

Securely route your AI API keys through a lightweight serverless function. Use simple prompts to automatically translate a subtitle track, detect the spoken language, and tidy the lines before export.

Translate a finished track

The feature people buy a subtitle tool for. Send the captions to your provider a batch at a time, keep every timestamp, and show progress while it runs, because a long video takes a while and a blank screen looks broken.

PromptTranslate a finished track
Add one server-side "ai" function that talks to my AI provider, with the key held in server secrets and never anywhere the browser can reach, and make it refuse any caller who is not signed in. Then add a "Translate" action on a subtitle track: send the lines to that function ten at a time with their timestamps, ask for the same lines in the target language with the timestamps untouched, and save the result as a new track linked to the original. Return immediately, write progress to the track row as batches finish, and have the editor poll it so the person sees a bar rather than a spinner. Log every call with the account, the video and the token counts.

Detect the language from the audio, not the file name

Guessing the spoken language from a file name works until somebody uploads recording_final_v2.mp4. Use the first minute of the transcript instead, and let the person correct it before anything else runs.

PromptDetect the language from the audio, not the file name
Change language detection to work from the first thirty seconds of the transcript rather than from the file name. Send that text to the ai function and ask for a two-letter language code, show the result in the editor as an editable field with a note saying it was detected, and use the person’s choice for translation from then on. Fall back to the file name only when no transcript exists yet.

Tidy the lines before export

A raw transcript has no punctuation worth the name, runs sentences together and keeps every um. One pass through the model turns it into captions a person would have written, and the editor keeps the original in case it overreaches.

PromptTidy the lines before export
Add a "Tidy captions" action that sends the track to the ai function and asks for the same lines with punctuation and capitalisation fixed, filler words removed, and any line longer than 42 characters split at a natural break, keeping every start and end time and never changing the meaning. Show the result as a proposal the person accepts or rejects line by line, and keep the original track untouched until they accept.

Chapters and a summary from the transcript

The description, the chapter markers and the short summary a video platform asks for are all in the transcript already. Drafting them is the chore that stops people publishing.

PromptChapters and a summary from the transcript
Add a panel on the export screen that sends the finished transcript to the ai function and returns a short summary, five to eight chapter markers with timestamps, and a handful of tags. Present all of it as editable drafts the person can copy, and let them regenerate any one part on its own without redoing the rest.

A cheap model is plenty for most of the prompts above. Save Pro’s stronger models for the one or two spots where the extra reasoning actually pays for itself. Add any provider key through Replit’s own Secrets tool rather than hard-coding it, and keep every AI feature behind one server-side function so a single key covers the whole app.

Ready-made option

Get a head start with our template

Everything above starts from an empty folder, and it does not have to. The same studio exists already built: the speech model runs, the timeline works, translation runs in the background and the exports come out, so your time goes on the languages, the styles and the name on the door.

AI Subtitle Studio

The exact subtitle studio this guide builds, packaged so you can open it, point it at your own backend, and make it yours from there. A subtitle workspace that takes a video from spoken words to a finished file. Transcribe it locally with no per-minute cost, fix and style the lines on a timeline, translate into other languages, and export a captioned video - so one upload reaches a much wider audience.

React 18ViteTypeScriptTailwind CSSSupabase
Out of the box

The key benefits of starting with a template

The speech model, the timeline editor, translation and the exports already work. Behind them, the roles, the hourly limits, the demo account and the admin console are done too.

Building the core from scratch

~101 hrs

Opening the template, already built

~1 hr

Pay Once, Own Forever. Build exactly what your team needs without renting a monthly SaaS subscription, paying per-seat fees, or dealing with platform lock-in.

Transcription in the browser, already wired

The speech model loads in a background worker, uses the graphics chip where the browser allows it and falls back where it does not, and chunks long audio so the page stays usable. That is the part a first attempt gets wrong, and it is done.

A timeline editor with undo and autosave

Captions on a track under the video, retimed by dragging, edited in place, with keyboard shortcuts, undo and autosave, plus a styling panel whose preview is what the export produces.

Translation in the background, progress on screen

Send a track to your own AI key and the function works through it in batches while the editor shows progress. Each translated track saves as its own version and exports on its own.

Four export paths

Three plain subtitle formats generated on the server with no AI involved, and a captioned video recorded in the browser for the platforms that cannot take a separate file.

Roles, limits and an admin console

Visitor, user and admin roles plus a self-resetting demo role, 31 access policies holding each account to its own videos, hourly rate limits on uploads, transcriptions, translations and exports, and an admin view over users, videos, usage and settings.

Customer story

From founders who build on our templates

We needed a working product in front of users fast. I started from one of these templates instead of a blank repo, customized it in our AI tool, and shipped in days - not the weeks it usually takes.
Jeevan ThomasJeevan ThomasFounder & CEO, Hado.ai
Got questions?

Common questions

The transcription minute is, yes. The speech model downloads into the visitor’s browser on first use and runs on their own processor, so no service bills you for the audio. What you pay for is your Replit plan, the object storage for the videos people keep and, if you offer it, translation through your own AI key, which costs tokens rather than minutes.

Less accurate than the rented services, and honestly so. The model that fits in a browser is the smallest one, and it gets clear speech mostly right and difficult audio less so, which is why the timeline editor exists. A larger model is a one-line swap and buys accuracy at the cost of a longer download on the first visit.

Yes, to object storage in this workspace, so it can sit in the library, play under the editor and be exported later. The transcription itself still happens in the browser rather than on your server. Decide how long uploads live, because they are the one large thing this app keeps.

A browser tab has finite memory, so this is a decision rather than a fixed number. The template chunks audio into thirty-second pieces and warns itself above thirty minutes. Short talks and lessons are comfortable on ordinary hardware. Set a ceiling you can defend, say it before the upload, and fail politely.

The speech model transcribes many spoken languages, and translation into any target language goes through your own AI key. One honest note: the template guesses the spoken language from the file name rather than the audio, so the editor lets a person correct it, and the AI section shows the prompt that fixes it properly.

Yes. The browser draws the styled captions over the video and records the result as a WebM file, with no server rendering. On this build the video is read through your own route, so the export works without making anybody’s video public, and it is slower than a desktop editor because it plays the video through in real time.

No. Tables, sessions, access rules and the server routes behind them all get written from plain-language prompts, and the database they write into is created inside this workspace when you ask. Nothing needs opening anywhere else first.

Less than a subtitle service, because the metered minute is missing. What is left is your Replit plan. Core starts at $25/month, or $20/month billed annually, with 20GB of database included, plus object storage for the videos you keep, Publishing billed separately on top, and a translation bill on your own AI key that runs to cents per hour of speech at the default model’s rates.

Nobody, unless you build that on purpose. Every video, track and usage record belongs to one account, the database enforces it on every query, and the videos themselves are served through a route that checks the owner rather than by a link anyone could open.

Not without an account, but the template makes accounts cheap to try: any sign-up with a demo address gets a demo role, can do everything a real account can, and is wiped after an hour of inactivity. That is the right shape for a public trial of a tool that spends your translation key.

Nothing here is proprietary. Captions live in ordinary PostgreSQL and export as standard subtitle files, videos come out of object storage as the files they went in as, and a standard database dump gives you everything in a form any Postgres host accepts.

Yes. The replit.app subdomain it starts on can be replaced with your own through Publishing, HTTPS included. Do it before you invite anyone, because a tool that asks people to upload their own videos should look like it belongs to you.

No. You describe what you want in the chat, and the Agent handles the rest: the code, the database, and a live version of the app right there in the workspace. The setup section above covers the account and plan you need first.

Replit hosts it. Publishing takes the same project live on a Replit domain, or your own if you connect one, with monitoring and access controls included, so you never open a separate hosting account.

Starter’s daily credits and a paid plan’s monthly grant both refill on their own schedule. Hitting either limit mid-build doesn’t touch what you’ve already made. Move up a plan for more headroom right away, or wait it out.

References

Sources checked September 2026
  1. 01Pricing (AI minutes per seat, human captions), Rev. rev.com
  2. 02Pricing (media hours per person), Descript. descript.com
  3. 03Pricing (subtitle and translation minutes per member), Kapwing. kapwing.com
  4. 04Pricing (minutes per month, top-up rate), Happy Scribe. happyscribe.com
  5. 05Transformers.js documentation (running models in the browser). huggingface.co
  6. 06whisper-tiny model card (the browser speech model). huggingface.co
  7. 07API pricing (per-token rates for the default translation model), OpenAI. developers.openai.com
  8. 08Pricing (Pro plan, storage, egress, backups), Supabase. supabase.com
  9. 09Web developer hourly rates 2026 (freelance and agency benchmarks). developex.com
  10. 10Pricing (Starter, Core, Pro), Replit. replit.com
  11. 11Built-in database, Replit docs. docs.replit.com
  12. 12Publishing overview, Replit docs. docs.replit.com
  13. 13Deployment types, Replit docs. docs.replit.com
  14. 14Checkpoints and rollbacks, Replit docs. docs.replit.com
  15. 15Using the Git pane, Replit docs. docs.replit.com
  16. 16Import from a provider, Replit docs. docs.replit.com
  17. 17Secrets, Replit docs. docs.replit.com

This guide is general information. Third-party prices, plan limits and market rates are quoted from the sources above and were last checked on the date shown. Vendors change them without notice, so confirm before you budget. Transcription accuracy, speed and the practical file limit depend on the visitor’s own device and browser, so treat any performance expectation here as a starting point to test rather than a specification. Build hours and the cost estimates derived from them are our own estimates, not quotes. Replit is a product of Replit, Inc. Verify current capabilities and pricing before relying on them.