Plenty of Indonesian YouTubers would like English subtitles, and plenty of English-language creators would like to reach the Indonesian, Malaysian or Spanish-speaking audience. Most never do it. Hiring a human translator for every upload is expensive, and YouTube’s free auto-translated captions are run through machine translation from auto-generated captions, so they often turn names, slang and product terms into nonsense.
That gap is a small, real service you can sell: clean, reviewed subtitle files in another language, delivered as an .srt the creator uploads in YouTube Studio. AI does the first draft. You do the part clients actually pay for: checking it, fixing terms, and making sure every line can be read in time.
In this guide you’ll build a Python tool that:
- reads an
.srtfile and sends the cues to Gemini 3.5 Flash-Lite in batches - gets back structured JSON with exactly one translation per cue, so timing never drifts
- uses a glossary so names and brand terms stay consistent
- re-wraps lines to a maximum of 42 characters and flags cues that are too fast to read
- prints the token cost of every run
AI cost: about $0.01 per 10 minutes of video. Build time: about 45 minutes.
What you need
| Item | Cost | Notes |
|---|---|---|
| Python 3.10+ | $0 | Windows, macOS or Linux |
| Gemini API key | $0 to start | aistudio.google.com → “Get API key” |
google-genai, pydantic, srt | $0 | pip install |
A source .srt file | $0 | From YouTube Studio, Whisper, or the creator |
| Fluency in the target language | Not optional | You review every file before it ships |
Pricing as of October 2026 (check ai.google.dev/pricing before you rely on it, because Google changes models and prices often):
- Gemini 3.5 Flash-Lite (
gemini-3.5-flash-lite, paid tier): $0.30 per 1M input tokens, $2.50 per 1M output tokens. This is the default in the script. - Gemini 3.6 Flash (
gemini-3.6-flash, paid tier): $0.75 per 1M input, $3.75 per 1M output until Dec 31, 2026, then $1.50 / $7.50. Worth trying only if Flash-Lite struggles with a language pair. - Free tier: rate-limited, fine for testing. Don’t use it for client work (see Reality check).
A note on Gemini 2.5: older tutorials use gemini-2.5-flash. Since September 2026 Google only serves the 2.5 models to projects that already used them, so a new API key can’t call them. Use a 3.x model.
Where the source subtitles come from
- The creator’s own captions. In YouTube Studio: Subtitles → pick the video → next to the language, open the ⋮ menu → Download →
.srt. Auto-generated captions can usually be downloaded the same way, but fix obvious mistakes before you translate them. Errors in the source end up in every language. - Whisper. If there are no captions, transcribe the audio first. Our guide on batch captioning with faster-whisper produces
.srtfiles for free on your own machine, and the Groq Whisper guide does it for about $0.04 per hour of audio.
How the cost works out
People speak roughly 130–160 words per minute in a talking-head video, so 10 minutes is about 1,300–1,600 words. With JSON overhead, the prompt, the glossary and a few lines of context per batch, a 10-minute file comes to roughly 6,000 input tokens and 3,500 output tokens.
On Gemini 3.5 Flash-Lite:
- Input: 6,000 × $0.30/1M = $0.0018
- Output: 3,500 × $2.50/1M = $0.0088
- Total: about $0.01 per 10 minutes, or about $0.06 per hour of video
Gemini 3.x models “think” before answering, and thinking tokens are billed as output. Flash-Lite can’t switch thinking off completely, but the script sets thinking_level to minimal (the lowest setting). Raise it and the cost can go up several times with no clear gain for this job. The script prints real token counts, so you don’t have to trust these estimates.
Step 1: Set up the project
mkdir subtrans && cd subtrans
python3 -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install google-genai pydantic srt
export GEMINI_API_KEY="paste-your-key-here" # Windows PowerShell: $env:GEMINI_API_KEY="..."
Create a glossary.txt for each client. One term per line, source = target. Use = with the same word on both sides for things that must never be translated:
Solvinc = Solvinc
warung = warung
gorengan = gorengan (fried snacks)
cuan = profit
Kak Rina = Rina
Step 2: Why structured output matters here
The naive approach is to paste the whole .srt into a chatbot and ask for a translation. It works for a short clip, then it breaks on a real video: the model merges two cues, skips one, renumbers them, or “fixes” the timestamps. Now every subtitle after minute 6 shows up a line late.
The fix is to never let the model touch timing at all. The script keeps timestamps in Python, sends only {id, text} pairs, asks for the same IDs back via a JSON schema, and refuses any reply where the IDs don’t match exactly.
Step 3: The script
Save this as translate_srt.py:
import argparse, json, os, sys, textwrap, time
from pathlib import Path
import srt
from google import genai
from google.genai import errors, types
from pydantic import BaseModel
MODEL = os.environ.get("SUB_MODEL", "gemini-3.5-flash-lite")
PRICE_IN, PRICE_OUT = 0.30, 2.50 # USD per 1M tokens, Gemini 3.5 Flash-Lite paid tier
BATCH = 40 # cues per request
CONTEXT = 5 # previous cues sent along for continuity
MAX_LINE, MAX_LINES, MAX_CPS = 42, 2, 20
class Cue(BaseModel):
id: int
text: str
client = genai.Client() # reads GEMINI_API_KEY
def load_glossary(path):
if not path:
return ""
lines = [l.strip() for l in Path(path).read_text(encoding="utf-8").splitlines()]
return "\n".join(l for l in lines if l and "=" in l)
def build_prompt(src, tgt, glossary, context, batch):
return f"""You are a professional subtitle translator from {src} to {tgt}.
Rules:
- Return exactly one item per input id, with the same id. Never merge, split or skip cues.
- Keep each cue's meaning inside that cue, even if a sentence continues in the next one.
- Write natural spoken {tgt}, short and easy to read. Shorten filler words if needed.
- Keep names, brands and numbers exactly. Follow the glossary strictly.
- Translate slang by meaning, not word for word. Do not add explanations or notes.
- No line breaks inside text.
Glossary (source = target):
{glossary or "(none)"}
Previous cues, for context only (do not return these):
{json.dumps(context, ensure_ascii=False)}
Translate these cues:
{json.dumps(batch, ensure_ascii=False)}"""
def translate_batch(prompt, ids):
for attempt in range(3):
try:
resp = client.models.generate_content(
model=MODEL,
contents=prompt,
config=types.GenerateContentConfig(
response_mime_type="application/json",
response_schema=list[Cue],
thinking_config=types.ThinkingConfig(thinking_level="minimal"),
),
)
except errors.APIError as e:
print(f" API error {e.code} (attempt {attempt + 1}), retrying", file=sys.stderr)
time.sleep(5 * (attempt + 1))
continue
got = {c.id: c.text.strip() for c in (resp.parsed or [])}
if set(got) == set(ids) and all(got.values()):
return got, resp.usage_metadata
print(f" ids mismatch (attempt {attempt + 1}), retrying", file=sys.stderr)
sys.exit(f"Batch starting at cue {ids[0] + 1} failed 3 times. Try a smaller BATCH.")
def wrap(text):
lines = textwrap.wrap(text, MAX_LINE)
return "\n".join(lines), len(lines)
def main():
ap = argparse.ArgumentParser()
ap.add_argument("srt_file")
ap.add_argument("--src", default="Indonesian")
ap.add_argument("--to", default="English")
ap.add_argument("--code", default="en", help="suffix for output file, e.g. en, id, es")
ap.add_argument("--glossary")
args = ap.parse_args()
path = Path(args.srt_file)
subs = list(srt.parse(path.read_text(encoding="utf-8-sig")))
glossary = load_glossary(args.glossary)
done, tok_in, tok_out = {}, 0, 0
for start in range(0, len(subs), BATCH):
chunk = subs[start:start + BATCH]
batch = [{"id": i, "text": s.content.replace("\n", " ")}
for i, s in enumerate(chunk, start)]
context = [{"source": subs[i].content.replace("\n", " "), "translation": done[i]}
for i in range(max(0, start - CONTEXT), start)]
prompt = build_prompt(args.src, args.to, glossary, context, batch)
got, usage = translate_batch(prompt, [b["id"] for b in batch])
done.update(got)
tok_in += usage.prompt_token_count or 0
tok_out += (usage.candidates_token_count or 0) + (usage.thoughts_token_count or 0)
print(f" cues {start + 1}-{start + len(chunk)} of {len(subs)} done")
flags = []
for i, s in enumerate(subs):
text, n_lines = wrap(done[i])
secs = max((s.end - s.start).total_seconds(), 0.001)
cps = len(text.replace("\n", "")) / secs
s.content = text
if n_lines > MAX_LINES or cps > MAX_CPS:
flags.append(f"#{s.index} {s.start} --> {s.end} {cps:.0f} cps, "
f"{n_lines} lines\n {text.replace(chr(10), ' / ')}")
out = path.with_name(f"{path.stem}.{args.code}.srt")
out.write_text(srt.compose(subs, reindex=False), encoding="utf-8")
report = path.with_name(f"{path.stem}.{args.code}.review.txt")
report.write_text("\n".join(flags) or "No flagged cues.", encoding="utf-8")
cost = tok_in / 1e6 * PRICE_IN + tok_out / 1e6 * PRICE_OUT
print(f"Wrote {out} ({len(subs)} cues), {len(flags)} flagged -> {report}")
print(f"Tokens: {tok_in} in / {tok_out} out, about ${cost:.4f}")
if __name__ == "__main__":
main()
What it does, in order:
- Parses the SRT with the
srtlibrary (utf-8-sigstrips the byte-order mark Windows tools often add). - Batches 40 cues at a time and sends the last 5 translated cues as context, so a name introduced in batch one is spelled the same in batch five.
- Validates the reply. If the returned IDs aren’t exactly the IDs sent, or the API returns an error such as a 429 rate limit, it retries up to three times, then stops loudly instead of writing a broken file.
- Re-wraps and checks readability. 42 characters per line and two lines per cue match Netflix’s English timed-text guidelines, a sensible default for YouTube too. Characters per second (CPS) above 20 means most viewers can’t finish reading before the line disappears.
- Writes two files:
video.en.srtto deliver andvideo.en.review.txtlisting the cues you need to look at.
Step 4: Run it
python translate_srt.py episode12.srt --src Indonesian --to English --code en --glossary glossary.txt
Example output for a 12-minute video (your token counts will differ):
cues 1-40 of 187 done
cues 41-80 of 187 done
...
Wrote episode12.en.srt (187 cues), 9 flagged -> episode12.en.review.txt
Tokens: 7412 in / 4180 out, about $0.0127
Several languages for the same video:
for pair in "English:en" "Malay:ms" "Spanish:es"; do
python translate_srt.py episode12.srt --to "${pair%%:*}" --code "${pair##*:}" --glossary glossary.txt
done
Only sell languages that you, or a reviewer you pay, can actually read.
Step 5: Review before you deliver
This step is the product. Open the .srt next to the video in a free subtitle editor such as Subtitle Edit (Windows, runs on Linux via Mono) or Aegisub (Windows/macOS/Linux) and play it at normal speed.
Work through review.txt first. For each flagged cue, either shorten the text (“What I want to say is that you basically need to” → “You need to”) or, in the editor, stretch the cue’s end time into a gap. Then watch the whole thing once at 1× and check for:
- Glossary misses, especially product names and people
- Jokes and slang translated literally. “Cuan” means profit; it isn’t a sound effect.
- Sentences split across cues that read oddly when shown one line at a time
- Numbers and currencies. “Rp 25 ribu” should become “IDR 25,000” or “Rp25,000”, whichever the client prefers, used the same way throughout.
Rough planning figure: 1.5–2.5× the video length for a careful review, so 15–25 minutes for a 10-minute video. Clear talking-head content is quicker; podcasts with three people talking over each other are slower. Time your first five jobs and price from that, not from this estimate.
The creator uploads it in YouTube Studio: Subtitles → video → Add language → under Subtitles, Add → Upload file → With timing.
Turning it into income
Who buys. Creators with an existing audience who want a second one: Indonesian tech, cooking, travel and education channels adding English; English channels in finance, software tutorials or gaming adding Indonesian. Small course creators are another good fit, because course platforms accept uploaded caption files and the content doesn’t go out of date.
How to find them. Pick 20 channels in one niche with 10k–200k subscribers that post regularly and have no subtitles in your target language. Translate the first 3 minutes of one recent video for free and send them the file with a short note. A free sample on their own content beats any pitch.
How to price. Price per finished video minute, and look at what translators in your language pair charge on Fiverr and Upwork before you set a number. Don’t price at the bottom. Your costs per job look like this:
| Cost per 10-minute video | Amount |
|---|---|
| Gemini 3.5 Flash-Lite API | ~$0.01 |
| Your review time | 15–25 minutes |
| Platform fee (if via Fiverr) | 20% of the order |
Example only, not market data: at Rp 7.500 per minute, a 10-minute video pays Rp 75.000 for about 20 minutes of work, roughly Rp 225.000 per hour of your time. One client posting two 12-minute videos a week is about 100 minutes a month, or Rp 750.000/month. Five steady clients is a meaningful side income. Getting those five clients usually takes weeks of outreach, not days.
Retainers beat one-offs. Offer a monthly package (“every upload translated within 48 hours”) so you aren’t re-selling every week. Keep each client’s glossary file; it makes their jobs faster and better over time, and it’s a reason for them to stay.
Reality check
You’re competing with free. YouTube auto-translates captions for viewers at no cost, and has been rolling out AI auto-dubbing to many channels. A creator who is happy with that isn’t your client. You’re selling accuracy, consistent terms and lines people can read in time. If you can’t clearly beat the auto-translation on the client’s own video, don’t take the job.
Don’t sell languages you can’t check. The script will happily output Japanese, and you won’t be able to tell when it’s wrong. Errors in a client’s subtitles are your errors.
Garbage in, garbage out. Translating uncorrected auto-captions multiplies their mistakes. Fix the source first, or charge extra for doing it.
Indonesian runs long. English → Indonesian usually produces more characters than the source, so expect more CPS flags in that direction. Shortening is part of the job, not a bug.
Batch size vs. reliability. Bigger batches give the model more context but fail the ID check more often. If you see repeated retries, drop BATCH to 20–25.
Free tier and privacy. On Google’s free tier, inputs may be used to improve Google’s products. Unreleased client videos and paid course content should go through a paid-tier key, where that doesn’t apply. Check the current terms on ai.google.dev before you take on NDA work.
Only translate content the client owns. Translating films or TV shows for subtitle download sites without permission is copyright infringement, and it pays nothing anyway. Work for the person who made the content.
Model names change. Google retires and reroutes Gemini models every few months; gemini-3.5-flash and gemini-3.7-flash were both rerouted in October 2026. If gemini-3.5-flash-lite goes away, set SUB_MODEL to the current Flash-Lite or Flash model (for example export SUB_MODEL=gemini-3.6-flash) and update PRICE_IN/PRICE_OUT so the cost line stays accurate. Don’t add temperature or thinking_budget back in: Google has deprecated both for 3.x models. The API ignores temperature for now and says it will reject it in future models, and thinking_budget can cause 400 errors.
Final checklist
- Source
.srtchecked and corrected - Client glossary created and passed with
--glossary - Paid-tier API key for client work
- Every flagged cue in
review.txtfixed - Full 1× watch-through in Subtitle Edit or Aegisub
- Numbers, currency and name spellings consistent
- Delivered as
video.<lang>.srtwith upload instructions
Start with your own videos or a friend’s channel. Once you’ve delivered three files you’d be proud to show, you have your portfolio and your first case study.