Plenty of Indonesian YouTubers would like English subtitles, and plenty of English-language creators would like to reach the Indonesian, Malaysian or Spanish-speaking audience. Most never do it. Hiring a human translator for every upload is expensive, and YouTube’s free auto-translated captions are run through machine translation from auto-generated captions, so they often turn names, slang and product terms into nonsense.

That gap is a small, real service you can sell: clean, reviewed subtitle files in another language, delivered as an .srt the creator uploads in YouTube Studio. AI does the first draft. You do the part clients actually pay for: checking it, fixing terms, and making sure every line can be read in time.

In this guide you’ll build a Python tool that:

  • reads an .srt file and sends the cues to Gemini 3.5 Flash-Lite in batches
  • gets back structured JSON with exactly one translation per cue, so timing never drifts
  • uses a glossary so names and brand terms stay consistent
  • re-wraps lines to a maximum of 42 characters and flags cues that are too fast to read
  • prints the token cost of every run

AI cost: about $0.01 per 10 minutes of video. Build time: about 45 minutes.

What you need

ItemCostNotes
Python 3.10+$0Windows, macOS or Linux
Gemini API key$0 to startaistudio.google.com → “Get API key”
google-genai, pydantic, srt$0pip install
A source .srt file$0From YouTube Studio, Whisper, or the creator
Fluency in the target languageNot optionalYou review every file before it ships

Pricing as of October 2026 (check ai.google.dev/pricing before you rely on it, because Google changes models and prices often):

  • Gemini 3.5 Flash-Lite (gemini-3.5-flash-lite, paid tier): $0.30 per 1M input tokens, $2.50 per 1M output tokens. This is the default in the script.
  • Gemini 3.6 Flash (gemini-3.6-flash, paid tier): $0.75 per 1M input, $3.75 per 1M output until Dec 31, 2026, then $1.50 / $7.50. Worth trying only if Flash-Lite struggles with a language pair.
  • Free tier: rate-limited, fine for testing. Don’t use it for client work (see Reality check).

A note on Gemini 2.5: older tutorials use gemini-2.5-flash. Since September 2026 Google only serves the 2.5 models to projects that already used them, so a new API key can’t call them. Use a 3.x model.

Where the source subtitles come from

  • The creator’s own captions. In YouTube Studio: Subtitles → pick the video → next to the language, open the ⋮ menu → Download → .srt. Auto-generated captions can usually be downloaded the same way, but fix obvious mistakes before you translate them. Errors in the source end up in every language.
  • Whisper. If there are no captions, transcribe the audio first. Our guide on batch captioning with faster-whisper produces .srt files for free on your own machine, and the Groq Whisper guide does it for about $0.04 per hour of audio.

How the cost works out

People speak roughly 130–160 words per minute in a talking-head video, so 10 minutes is about 1,300–1,600 words. With JSON overhead, the prompt, the glossary and a few lines of context per batch, a 10-minute file comes to roughly 6,000 input tokens and 3,500 output tokens.

On Gemini 3.5 Flash-Lite:

  • Input: 6,000 × $0.30/1M = $0.0018
  • Output: 3,500 × $2.50/1M = $0.0088
  • Total: about $0.01 per 10 minutes, or about $0.06 per hour of video

Gemini 3.x models “think” before answering, and thinking tokens are billed as output. Flash-Lite can’t switch thinking off completely, but the script sets thinking_level to minimal (the lowest setting). Raise it and the cost can go up several times with no clear gain for this job. The script prints real token counts, so you don’t have to trust these estimates.

Step 1: Set up the project

mkdir subtrans && cd subtrans
python3 -m venv .venv && source .venv/bin/activate   # Windows: .venv\Scripts\activate
pip install google-genai pydantic srt

export GEMINI_API_KEY="paste-your-key-here"           # Windows PowerShell: $env:GEMINI_API_KEY="..."

Create a glossary.txt for each client. One term per line, source = target. Use = with the same word on both sides for things that must never be translated:

Solvinc = Solvinc
warung = warung
gorengan = gorengan (fried snacks)
cuan = profit
Kak Rina = Rina

Step 2: Why structured output matters here

The naive approach is to paste the whole .srt into a chatbot and ask for a translation. It works for a short clip, then it breaks on a real video: the model merges two cues, skips one, renumbers them, or “fixes” the timestamps. Now every subtitle after minute 6 shows up a line late.

The fix is to never let the model touch timing at all. The script keeps timestamps in Python, sends only {id, text} pairs, asks for the same IDs back via a JSON schema, and refuses any reply where the IDs don’t match exactly.

Step 3: The script

Save this as translate_srt.py:

import argparse, json, os, sys, textwrap, time
from pathlib import Path

import srt
from google import genai
from google.genai import errors, types
from pydantic import BaseModel

MODEL = os.environ.get("SUB_MODEL", "gemini-3.5-flash-lite")
PRICE_IN, PRICE_OUT = 0.30, 2.50   # USD per 1M tokens, Gemini 3.5 Flash-Lite paid tier
BATCH = 40                         # cues per request
CONTEXT = 5                        # previous cues sent along for continuity
MAX_LINE, MAX_LINES, MAX_CPS = 42, 2, 20

class Cue(BaseModel):
    id: int
    text: str

client = genai.Client()            # reads GEMINI_API_KEY

def load_glossary(path):
    if not path:
        return ""
    lines = [l.strip() for l in Path(path).read_text(encoding="utf-8").splitlines()]
    return "\n".join(l for l in lines if l and "=" in l)

def build_prompt(src, tgt, glossary, context, batch):
    return f"""You are a professional subtitle translator from {src} to {tgt}.
Rules:
- Return exactly one item per input id, with the same id. Never merge, split or skip cues.
- Keep each cue's meaning inside that cue, even if a sentence continues in the next one.
- Write natural spoken {tgt}, short and easy to read. Shorten filler words if needed.
- Keep names, brands and numbers exactly. Follow the glossary strictly.
- Translate slang by meaning, not word for word. Do not add explanations or notes.
- No line breaks inside text.

Glossary (source = target):
{glossary or "(none)"}

Previous cues, for context only (do not return these):
{json.dumps(context, ensure_ascii=False)}

Translate these cues:
{json.dumps(batch, ensure_ascii=False)}"""

def translate_batch(prompt, ids):
    for attempt in range(3):
        try:
            resp = client.models.generate_content(
                model=MODEL,
                contents=prompt,
                config=types.GenerateContentConfig(
                    response_mime_type="application/json",
                    response_schema=list[Cue],
                    thinking_config=types.ThinkingConfig(thinking_level="minimal"),
                ),
            )
        except errors.APIError as e:
            print(f"  API error {e.code} (attempt {attempt + 1}), retrying", file=sys.stderr)
            time.sleep(5 * (attempt + 1))
            continue
        got = {c.id: c.text.strip() for c in (resp.parsed or [])}
        if set(got) == set(ids) and all(got.values()):
            return got, resp.usage_metadata
        print(f"  ids mismatch (attempt {attempt + 1}), retrying", file=sys.stderr)
    sys.exit(f"Batch starting at cue {ids[0] + 1} failed 3 times. Try a smaller BATCH.")

def wrap(text):
    lines = textwrap.wrap(text, MAX_LINE)
    return "\n".join(lines), len(lines)

def main():
    ap = argparse.ArgumentParser()
    ap.add_argument("srt_file")
    ap.add_argument("--src", default="Indonesian")
    ap.add_argument("--to", default="English")
    ap.add_argument("--code", default="en", help="suffix for output file, e.g. en, id, es")
    ap.add_argument("--glossary")
    args = ap.parse_args()

    path = Path(args.srt_file)
    subs = list(srt.parse(path.read_text(encoding="utf-8-sig")))
    glossary = load_glossary(args.glossary)
    done, tok_in, tok_out = {}, 0, 0

    for start in range(0, len(subs), BATCH):
        chunk = subs[start:start + BATCH]
        batch = [{"id": i, "text": s.content.replace("\n", " ")}
                 for i, s in enumerate(chunk, start)]
        context = [{"source": subs[i].content.replace("\n", " "), "translation": done[i]}
                   for i in range(max(0, start - CONTEXT), start)]
        prompt = build_prompt(args.src, args.to, glossary, context, batch)
        got, usage = translate_batch(prompt, [b["id"] for b in batch])
        done.update(got)
        tok_in += usage.prompt_token_count or 0
        tok_out += (usage.candidates_token_count or 0) + (usage.thoughts_token_count or 0)
        print(f"  cues {start + 1}-{start + len(chunk)} of {len(subs)} done")

    flags = []
    for i, s in enumerate(subs):
        text, n_lines = wrap(done[i])
        secs = max((s.end - s.start).total_seconds(), 0.001)
        cps = len(text.replace("\n", "")) / secs
        s.content = text
        if n_lines > MAX_LINES or cps > MAX_CPS:
            flags.append(f"#{s.index} {s.start} --> {s.end}  {cps:.0f} cps, "
                         f"{n_lines} lines\n    {text.replace(chr(10), ' / ')}")

    out = path.with_name(f"{path.stem}.{args.code}.srt")
    out.write_text(srt.compose(subs, reindex=False), encoding="utf-8")
    report = path.with_name(f"{path.stem}.{args.code}.review.txt")
    report.write_text("\n".join(flags) or "No flagged cues.", encoding="utf-8")

    cost = tok_in / 1e6 * PRICE_IN + tok_out / 1e6 * PRICE_OUT
    print(f"Wrote {out} ({len(subs)} cues), {len(flags)} flagged -> {report}")
    print(f"Tokens: {tok_in} in / {tok_out} out, about ${cost:.4f}")

if __name__ == "__main__":
    main()

What it does, in order:

  1. Parses the SRT with the srt library (utf-8-sig strips the byte-order mark Windows tools often add).
  2. Batches 40 cues at a time and sends the last 5 translated cues as context, so a name introduced in batch one is spelled the same in batch five.
  3. Validates the reply. If the returned IDs aren’t exactly the IDs sent, or the API returns an error such as a 429 rate limit, it retries up to three times, then stops loudly instead of writing a broken file.
  4. Re-wraps and checks readability. 42 characters per line and two lines per cue match Netflix’s English timed-text guidelines, a sensible default for YouTube too. Characters per second (CPS) above 20 means most viewers can’t finish reading before the line disappears.
  5. Writes two files: video.en.srt to deliver and video.en.review.txt listing the cues you need to look at.

Step 4: Run it

python translate_srt.py episode12.srt --src Indonesian --to English --code en --glossary glossary.txt

Example output for a 12-minute video (your token counts will differ):

  cues 1-40 of 187 done
  cues 41-80 of 187 done
  ...
Wrote episode12.en.srt (187 cues), 9 flagged -> episode12.en.review.txt
Tokens: 7412 in / 4180 out, about $0.0127

Several languages for the same video:

for pair in "English:en" "Malay:ms" "Spanish:es"; do
  python translate_srt.py episode12.srt --to "${pair%%:*}" --code "${pair##*:}" --glossary glossary.txt
done

Only sell languages that you, or a reviewer you pay, can actually read.

Step 5: Review before you deliver

This step is the product. Open the .srt next to the video in a free subtitle editor such as Subtitle Edit (Windows, runs on Linux via Mono) or Aegisub (Windows/macOS/Linux) and play it at normal speed.

Work through review.txt first. For each flagged cue, either shorten the text (“What I want to say is that you basically need to” → “You need to”) or, in the editor, stretch the cue’s end time into a gap. Then watch the whole thing once at 1× and check for:

  • Glossary misses, especially product names and people
  • Jokes and slang translated literally. “Cuan” means profit; it isn’t a sound effect.
  • Sentences split across cues that read oddly when shown one line at a time
  • Numbers and currencies. “Rp 25 ribu” should become “IDR 25,000” or “Rp25,000”, whichever the client prefers, used the same way throughout.

Rough planning figure: 1.5–2.5× the video length for a careful review, so 15–25 minutes for a 10-minute video. Clear talking-head content is quicker; podcasts with three people talking over each other are slower. Time your first five jobs and price from that, not from this estimate.

The creator uploads it in YouTube Studio: Subtitles → video → Add language → under Subtitles, Add → Upload file → With timing.

Turning it into income

Who buys. Creators with an existing audience who want a second one: Indonesian tech, cooking, travel and education channels adding English; English channels in finance, software tutorials or gaming adding Indonesian. Small course creators are another good fit, because course platforms accept uploaded caption files and the content doesn’t go out of date.

How to find them. Pick 20 channels in one niche with 10k–200k subscribers that post regularly and have no subtitles in your target language. Translate the first 3 minutes of one recent video for free and send them the file with a short note. A free sample on their own content beats any pitch.

How to price. Price per finished video minute, and look at what translators in your language pair charge on Fiverr and Upwork before you set a number. Don’t price at the bottom. Your costs per job look like this:

Cost per 10-minute videoAmount
Gemini 3.5 Flash-Lite API~$0.01
Your review time15–25 minutes
Platform fee (if via Fiverr)20% of the order

Example only, not market data: at Rp 7.500 per minute, a 10-minute video pays Rp 75.000 for about 20 minutes of work, roughly Rp 225.000 per hour of your time. One client posting two 12-minute videos a week is about 100 minutes a month, or Rp 750.000/month. Five steady clients is a meaningful side income. Getting those five clients usually takes weeks of outreach, not days.

Retainers beat one-offs. Offer a monthly package (“every upload translated within 48 hours”) so you aren’t re-selling every week. Keep each client’s glossary file; it makes their jobs faster and better over time, and it’s a reason for them to stay.

Reality check

You’re competing with free. YouTube auto-translates captions for viewers at no cost, and has been rolling out AI auto-dubbing to many channels. A creator who is happy with that isn’t your client. You’re selling accuracy, consistent terms and lines people can read in time. If you can’t clearly beat the auto-translation on the client’s own video, don’t take the job.

Don’t sell languages you can’t check. The script will happily output Japanese, and you won’t be able to tell when it’s wrong. Errors in a client’s subtitles are your errors.

Garbage in, garbage out. Translating uncorrected auto-captions multiplies their mistakes. Fix the source first, or charge extra for doing it.

Indonesian runs long. English → Indonesian usually produces more characters than the source, so expect more CPS flags in that direction. Shortening is part of the job, not a bug.

Batch size vs. reliability. Bigger batches give the model more context but fail the ID check more often. If you see repeated retries, drop BATCH to 20–25.

Free tier and privacy. On Google’s free tier, inputs may be used to improve Google’s products. Unreleased client videos and paid course content should go through a paid-tier key, where that doesn’t apply. Check the current terms on ai.google.dev before you take on NDA work.

Only translate content the client owns. Translating films or TV shows for subtitle download sites without permission is copyright infringement, and it pays nothing anyway. Work for the person who made the content.

Model names change. Google retires and reroutes Gemini models every few months; gemini-3.5-flash and gemini-3.7-flash were both rerouted in October 2026. If gemini-3.5-flash-lite goes away, set SUB_MODEL to the current Flash-Lite or Flash model (for example export SUB_MODEL=gemini-3.6-flash) and update PRICE_IN/PRICE_OUT so the cost line stays accurate. Don’t add temperature or thinking_budget back in: Google has deprecated both for 3.x models. The API ignores temperature for now and says it will reject it in future models, and thinking_budget can cause 400 errors.

Final checklist

  • Source .srt checked and corrected
  • Client glossary created and passed with --glossary
  • Paid-tier API key for client work
  • Every flagged cue in review.txt fixed
  • Full 1× watch-through in Subtitle Edit or Aegisub
  • Numbers, currency and name spellings consistent
  • Delivered as video.<lang>.srt with upload instructions

Start with your own videos or a friend’s channel. Once you’ve delivered three files you’d be proud to show, you have your portfolio and your first case study.