Short vertical videos with big captions that light up word by word are everywhere now. A lot of people watch with the sound off, and Reels or TikToks get re-shared to WhatsApp Status and YouTube Shorts, where the platform’s own auto-captions don’t come along. So small businesses and creators want the captions burned into the video file.

The usual way to get them is to open each clip in CapCut, run auto-captions, fix the typos, choose a style and export. That’s fine for one video. For a shop that posts 20 clips a month, or for someone editing videos for several clients, it turns into hours of tapping on a phone.

In this guide you’ll build a small tool that runs on your own computer and:

  • transcribes every video in a folder with faster-whisper (open-source Whisper, runs offline)
  • saves the words with timestamps to a plain TSV file you can correct in any text editor or spreadsheet
  • builds an ASS subtitle file with short 2–4 word chunks and the current word highlighted in yellow
  • burns the captions into a new MP4 with FFmpeg, in batch

Software cost: $0. No API key, no subscription, and the videos stay on your machine. Build time is about 45 minutes.

If you just want text transcripts or .srt files for meetings, read Turn Meetings and WhatsApp Voice Notes into Searchable Notes with Whisper instead. This post is about the finished, styled video.

What you need

ItemCostNotes
Python 3.10+$0Windows, macOS or Linux
faster-whisper$0pip install faster-whisper "av<19" (MIT licence)
FFmpeg built with libass$0Ubuntu/Debian apt install ffmpeg, macOS brew install ffmpeg, Windows: the “full” build from gyan.dev
A bold font$0Montserrat from fonts.google.com (SIL Open Font License)
A computer$0 if you already have oneA GPU is optional. Short clips are fine on a CPU
Disk space~0.5–3 GBFor the Whisper model, depending on size

Check FFmpeg has subtitle support:

ffmpeg -hide_banner -filters | grep -E " ass | subtitles "

You should see both ass and subtitles. If nothing shows up, your FFmpeg was built without libass and you need a different build.

Optional paid alternatives (if your computer is too slow), prices as of October 2026, so check before you rely on them:

  • Groq whisper-large-v3-turbo: $0.04 per hour of audio, billed at a minimum of 10 seconds per request. It supports word timestamps (timestamp_granularities=["word"] with response_format="verbose_json").
  • OpenAI whisper-1: $0.006 per minute of audio, with word timestamps (response_format="verbose_json", timestamp_granularities=["word"]). But OpenAI has deprecated it and plans to remove it from the API on February 26, 2027, together with gpt-4o-transcribe and gpt-4o-mini-transcribe. The recommended replacement, gpt-transcribe ($0.0045/min), is a good transcriber, but OpenAI’s docs send you to whisper-1 when you need word timestamps, and this tool needs them. Don’t build a client workflow on whisper-1 now.

A 60-second Reel through Groq costs well under a cent, so even 100 clips a month is a few cents. The local route still wins on privacy, and you’re not depending on someone’s API or their deprecation schedule.

Step 1: Set up the project

mkdir captioner && cd captioner
python3 -m venv .venv && source .venv/bin/activate   # Windows: .venv\Scripts\activate
pip install faster-whisper "av<19"

mkdir in out fonts

Why pin av<19? With faster-whisper 1.2.1 (the latest release in October 2026) and PyAV 19, transcription crashes with TypeError: open() got an unexpected keyword argument 'metadata_errors'. PyAV 18 works. If a newer faster-whisper release fixes this, you can drop the pin.

Download Montserrat from Google Fonts, unzip it, and copy static/Montserrat-Bold.ttf into fonts/. Use the static file, not the variable font. libass often fails to match the bold weight in a variable font and quietly falls back to a default font such as DejaVu Sans.

Layout:

captioner/
├── in/          # raw videos go here (plus the .tsv and .ass files the script makes)
├── out/         # captioned videos come out here
├── fonts/       # Montserrat-Bold.ttf
├── captions.py
└── burn.sh

Step 2: Choose a Whisper model

faster-whisper downloads the model on first run. The main trade-off is speed versus accuracy:

ModelDownloadWhen to use it
small~480 MBClear English speech. Fast on any recent laptop CPU
medium~1.5 GBIndonesian or mixed-language content on CPU, if you can wait
large-v3-turbo~1.6 GBThe best balance if you have an NVIDIA GPU. Usable on CPU for short clips
large-v3~3 GBMaximum accuracy, slowest

Speed depends heavily on your hardware, so don’t trust anyone’s benchmark, including mine. Time a 60-second clip on your own machine with each model. For short-form video, even a model that runs at 1× real time is fine: a 45-second Reel takes 45 seconds while you get on with something else.

Accuracy on Indonesian, and especially on campur Indonesian–English with brand names, is clearly worse on small than on the large models. Test with your own voice and your own product names before you pick.

Step 3: The caption script

Save this as captions.py:

import argparse, csv, subprocess
from pathlib import Path

HEADER = """[Script Info]
ScriptType: v4.00+
PlayResX: {w}
PlayResY: {h}
WrapStyle: 0
ScaledBorderAndShadow: yes

[V4+ Styles]
Format: Name, Fontname, Fontsize, PrimaryColour, SecondaryColour, OutlineColour, BackColour, Bold, Italic, Underline, StrikeOut, ScaleX, ScaleY, Spacing, Angle, BorderStyle, Outline, Shadow, Alignment, MarginL, MarginR, MarginV, Encoding
Style: Cap,{font},{size},&H00FFFFFF,&H00FFFFFF,&H00000000,&H64000000,-1,0,0,0,100,100,0,0,1,{outline},2,2,60,60,{margin},1

[Events]
Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text
"""
HIGHLIGHT = r"{\c&H00FFFF&}"   # ASS colours are BGR: 00FFFF = yellow
NORMAL = r"{\c&HFFFFFF&}"

def video_size(path: Path) -> tuple[int, int]:
    out = subprocess.run(
        ["ffprobe", "-v", "error", "-select_streams", "v:0",
         "-show_entries", "stream=width,height", "-of", "csv=p=0:s=x", str(path)],
        capture_output=True, text=True, check=True).stdout.strip()
    w, h = out.split("x")[:2]
    return int(w), int(h)

def ts(t: float) -> str:                       # ASS time: H:MM:SS.cc
    cs = int(round(t * 100))
    return f"{cs // 360000}:{cs // 6000 % 60:02d}:{cs // 100 % 60:02d}.{cs % 100:02d}"

def transcribe(video: Path, model_name: str, lang: str | None, prompt: str | None):
    from faster_whisper import WhisperModel   # imported here so TSV-only runs are instant
    model = WhisperModel(model_name, device="auto", compute_type="int8")
    segments, info = model.transcribe(
        str(video), language=lang, word_timestamps=True,
        vad_filter=True, initial_prompt=prompt, beam_size=5)
    words = [(round(w.start, 2), round(w.end, 2), w.word.strip())
             for seg in segments for w in seg.words if w.word.strip()]
    print(f"  language={info.language} words={len(words)}")
    return words

def load_tsv(path: Path):
    with open(path, encoding="utf-8") as f:
        return [(float(r[0]), float(r[1]), r[2]) for r in csv.reader(f, delimiter="\t")
                if len(r) == 3 and r[2].strip()]

def save_tsv(path: Path, words):
    with open(path, "w", newline="", encoding="utf-8") as f:
        csv.writer(f, delimiter="\t").writerows(words)

def chunks(words, max_words: int, max_chars: int):
    cur = []
    for start, end, text in words:
        line = " ".join(w[2] for w in cur + [(0, 0, text)])
        if cur and (len(cur) >= max_words or len(line) > max_chars
                    or start - cur[-1][1] > 0.6):          # pause = new chunk
            yield cur
            cur = []
        cur.append((start, end, text))
        if text[-1] in ".?!,":                              # punctuation = new chunk
            yield cur
            cur = []
    if cur:
        yield cur

def build_ass(words, w: int, h: int, args) -> str:
    size = round(min(w, h) * 0.08)
    margin = round(h * 0.25) if h > w else round(h * 0.08)  # clear of TikTok/Reels UI
    out = [HEADER.format(w=w, h=h, font=args.font, size=size,
                         outline=max(3, size // 14), margin=margin)]
    for ch in chunks(words, args.max_words, args.max_chars):
        texts = [t.replace("{", "(").replace("}", ")").replace("\\", "/") for _, _, t in ch]
        if args.upper:
            texts = [t.upper() for t in texts]
        for i in range(len(ch)):
            start = ch[0][0] if i == 0 else ch[i][0]
            end = ch[i + 1][0] if i + 1 < len(ch) else ch[-1][1]
            end = max(end, start + 0.05)
            line = " ".join(HIGHLIGHT + t + NORMAL if j == i else t
                            for j, t in enumerate(texts))
            out.append(f"Dialogue: 0,{ts(start)},{ts(end)},Cap,,0,0,0,,{line}")
    return "\n".join(out) + "\n"

def main():
    p = argparse.ArgumentParser()
    p.add_argument("video", type=Path)
    p.add_argument("--model", default="small")
    p.add_argument("--lang", default=None, help="e.g. id, en; omit to auto-detect")
    p.add_argument("--prompt", default=None, help="brand names and spellings to bias toward")
    p.add_argument("--font", default="Montserrat")
    p.add_argument("--max-words", type=int, default=3)
    p.add_argument("--max-chars", type=int, default=18)
    p.add_argument("--upper", action="store_true")
    p.add_argument("--retranscribe", action="store_true")
    args = p.parse_args()

    tsv = args.video.with_suffix(".tsv")
    if tsv.exists() and not args.retranscribe:
        words = load_tsv(tsv)                 # your corrected version wins
        print(f"  using edited {tsv.name}")
    else:
        words = transcribe(args.video, args.model, args.lang, args.prompt)
        save_tsv(tsv, words)
    w, h = video_size(args.video)
    args.video.with_suffix(".ass").write_text(build_ass(words, w, h, args), encoding="utf-8")
    print(f"  wrote {args.video.with_suffix('.ass').name} ({w}x{h})")

if __name__ == "__main__":
    main()

What it does, in order:

  1. Transcribes with word timestamps. vad_filter=True skips silence, which cuts down Whisper’s habit of inventing text over quiet parts or background music.
  2. Saves a TSV: one line per word, start end word. This is the file you correct. On the next run the script sees the TSV and uses your corrected version instead of transcribing again.
  3. Splits the words into chunks of at most 3 words / 18 characters, and starts a new chunk after a pause or a comma or full stop. Short chunks are what make the “TikTok style” readable at a glance.
  4. Writes one subtitle event per word. Each event shows the whole chunk with the current word in yellow. That’s the word-by-word highlight.
  5. Scales the font and position to the video. On vertical video the captions sit about a quarter of the way up, so they don’t hide under the like/comment buttons and the description text.

Try it on one clip:

cp ~/Downloads/promo.mp4 in/
python captions.py in/promo.mp4 --lang id --upper \
  --prompt "Kopi Senja, es kopi susu gula aren, GoFood, ShopeeFood"
  language=id words=112
  wrote promo.ass (1080x1920)

Step 4: Correct the transcript (the step that matters)

Open in/promo.tsv:

0.42	0.81	Halo
0.81	1.2	semua,
1.38	1.9	hari
1.9	2.22	ini
2.22	2.6	kopi
2.6	3.1	senjah

Fix senjah → Senja, save, and run the same command again. It reads the TSV, skips transcription and rebuilds promo.ass in under a second. You can merge two words into one line, or delete a line for a filler word like “eee”. Leave the timestamps alone unless a word is clearly in the wrong place.

If you edit the TSV in Excel or Google Sheets, export it as tab-separated text again, and check that the decimal separator is still a dot. Spreadsheets set to Indonesian locale will happily turn 1.38 into 1,38 or a date. A plain text editor (Notepad, VS Code) is safer.

The --prompt option is Whisper’s initial_prompt. It nudges the model toward the spellings you give it. It helps with brand names, but it doesn’t guarantee them, so you still need to look over the transcript.

Want to see the result before you burn it in? Open the video in VLC and drag the .ass file onto it, or play it with mpv in/promo.mp4 --sub-file=in/promo.ass.

Step 5: Burn the captions in, in batch

Save this as burn.sh:

#!/usr/bin/env bash
set -euo pipefail
LANG_CODE="${1:-id}"
shopt -s nullglob nocaseglob
for f in in/*.mp4 in/*.mov; do
  name="$(basename "${f%.*}")"
  [[ -f "out/$name.mp4" ]] && { echo "skip $name"; continue; }
  echo "== $name"
  python captions.py "$f" --lang "$LANG_CODE" --upper
  ffmpeg -y -loglevel error -i "$f" \
    -vf "ass=in/$name.ass:fontsdir=fonts" \
    -c:v libx264 -crf 20 -preset medium -pix_fmt yuv420p \
    -c:a aac -b:a 160k -movflags +faststart \
    "out/$name.mp4"
done
chmod +x burn.sh
./burn.sh id       # or: ./burn.sh en

The order to work in:

  1. Drop all the raw clips into in/.
  2. Run ./burn.sh once. Every clip gets a TSV and a first captioned version.
  3. Watch the outputs. For any clip with mistakes, fix its TSV, delete that clip’s file from out/, and run ./burn.sh again. Only the deleted clips are redone, using your corrected TSVs.

Notes on the FFmpeg flags:

  • -crf 20 is high quality for social media. Instagram and TikTok re-compress your upload anyway, so going lower than 18 just makes bigger files.
  • -pix_fmt yuv420p stops “file not supported” errors on some phones and players. This flag alone doesn’t convert iPhone HDR footage properly, though. HDR clips can come out washed-out, so set the iPhone camera to record in SDR for client work.
  • -c:a aac re-encodes the audio so .mov files with unusual audio still produce a valid MP4.
  • Filenames: keep them simple (promo-01.mp4). Spaces, colons, commas and apostrophes in a path break FFmpeg’s filter syntax.

On Windows, run the loop in Git Bash or WSL, or translate it to a PowerShell foreach. The FFmpeg command is the same.

Honest cost and time breakdown

Scenario: a café that posts 20 Reels a month, each about 45 seconds, mostly spoken in Indonesian.

ItemMonthly cost
faster-whisper, FFmpeg, fonts$0
Electricity for a few minutes of CPU per clipNegligible
Optional: Groq API instead, 15 min of audio~$0.01
Total$0 – ~$0.01

Time per clip (estimates, so measure your own):

  • Manual in a phone editor (auto-captions, fixing words, styling, export): roughly 10–15 minutes
  • This tool: about 1 minute to drop the clip in, 2–5 minutes reviewing and fixing the TSV, and processing that runs unattended

At 20 clips that’s about 3–5 hours down to about 1–2 hours a month. The bigger gain is consistency: every clip has exactly the same font, size, colour and position, which a brand cares about more than you’d expect.

Making money with it

None of the figures here are promises. They’re how to think about pricing.

1. Captioning as part of a social media package. Many UMKM already pay someone (often a freelancer or a relative) to post on Instagram and TikTok. Burned-in captions are an easy add-on to that service, because the client sees the difference straight away on the video. Offer it per clip or as a monthly bundle.

2. Batch work for creators and coaches. People who record many talking-head clips in one session (coaches, teachers, property agents, preachers, product reviewers) need 10–30 clips captioned at a time. Your edge over doing it by hand in CapCut is turnaround: a batch is done the same day, in a consistent style.

3. Caption gigs on Fiverr/Upwork. Search “video captions” or “TikTok subtitles” on Fiverr and look at what the sellers with reviews actually charge, and how many reviews they have. It’s a crowded market, so compete on language skills (Indonesian, Javanese mixed in, bilingual content) rather than on price.

How to price: work it out from your review time per clip (realistically 3–10 minutes with corrections) plus revisions, not from the $0 software cost. If a client sends 30 clips a month, charge a monthly fee that pays you well for 2–4 hours, and include one round of corrections.

Realistic expectations: the first client usually comes from showing a before/after of their own video, not from a portfolio. “I captioned your last Reel, here’s the file” converts much better than describing the tool.

Reality check: what does NOT work

Thinking AI captions don’t need review. Whisper gets most clear speech right, but it regularly misspells brand names, mishears slang, and sometimes “hears” words over background music or silence. A typo burned into a client’s ad can’t be fixed after it’s posted. The TSV review step is the product, so don’t skip it.

Music-heavy clips. If a clip has loud background music or no speech at all, Whisper may produce garbage or repeat a phrase over and over. Caption the voice-over before the music is mixed in if you can, or leave those clips out.

Ignoring what the platforms already give away. TikTok and Instagram both have built-in auto-captions, and CapCut does styled captions with a few taps. If someone posts two videos a month, they don’t need you or this tool. The value is in volume, consistency, correct names, and captions that stay with the video when it’s re-shared to WhatsApp Status or other platforms.

Squashed or misplaced captions on phone footage. Some phone videos store rotation as metadata. ffprobe then reports 1920×1080 for a clip that plays as 1080×1920, and the captions come out the wrong size. Fix it by re-encoding once (ffmpeg -i in.mov -c:v libx264 -crf 18 -c:a aac fixed.mp4) and captioning fixed.mp4. FFmpeg bakes the rotation in by default.

The wrong font appears. If libass can’t find “Montserrat” it quietly falls back to a default font. Check that fonts/Montserrat-Bold.ttf exists and the path in fontsdir= is right. Run FFmpeg without -loglevel error once to see font warnings.

Too many words per line. It’s tempting to raise --max-words to 6 so there are fewer caption changes. On a phone held upright, 6 bold words either wrap onto two lines or shrink to an unreadable size. Keep 2–4 words for vertical video. Landscape YouTube videos can take more.

Using the large model on an old laptop for long videos. For 30–90 second clips, a CPU is fine. For a 40-minute podcast you want burned-in captions on, large-v3 on an old CPU may take longer than the video itself. Use small/medium, a GPU, or the paid API for that.

Music and footage rights. Captioning a video doesn’t change who owns it. Only process videos the client owns or has licensed, and don’t add trending songs to client ads unless you know they’re allowed commercially.

Final checklist

  • ffmpeg -filters shows ass; Montserrat-Bold.ttf in fonts/
  • Timed a 60-second clip with small and medium on your machine
  • --prompt filled with the client’s brand and product names
  • Every TSV reviewed before the final burn
  • Output played on an actual phone, captions clear of the app’s buttons
  • Raw clips and corrected TSVs backed up, so you can re-render in a different style later

Caption your own last five videos first. If the review step takes you a few minutes per clip and the results look as good as what you’d make in CapCut, you have a tool that saves you hours, and a demo you can show to the next business owner who posts videos with sound off.