You know the problem. A client sends a four-minute WhatsApp voice note with the order details in it. A supplier call runs 40 minutes and nobody writes anything down. A week later someone asks “what did we agree on the price?” and you scrub through audio on your phone trying to find the bit you need.
Speech-to-text is one of the few AI tools that is cheap, accurate and dull, which is what you want here. OpenAI’s Whisper models are open source, they handle Indonesian and English (including people switching between the two mid-sentence) well enough to be useful, and hosted versions now cost almost nothing.
In this guide you’ll build a small Python tool that:
- takes any audio file: meeting recordings, WhatsApp
.opusvoice notes, phone call recordings, Zoom.m4afiles - shrinks it and splits it into chunks with
ffmpeg - transcribes it with Whisper large-v3-turbo on Groq, with timestamps
- asks an LLM for a summary, decisions and action items
- writes one Markdown file per recording, plus an
.srtsubtitle file
Running cost: about $0.04 per hour of audio to transcribe, plus well under a cent for the summary. Build time is 30–45 minutes.
What you need
| Item | Cost | Notes |
|---|---|---|
| Python 3.10+ | $0 | Windows, macOS or Linux |
ffmpeg | $0 | Converts and splits audio |
| Groq API key | $0 to start | console.groq.com → API Keys |
groq Python SDK | $0 | pip install groq |
| Audio files | $0 | Phone recorder, Zoom/Meet recordings, exported WhatsApp voice notes |
Pricing as of October 2026. Check groq.com/pricing and openai.com/api/pricing before you commit, because both change model line-ups often.
| Service | Model | Price per hour of audio |
|---|---|---|
| Groq | whisper-large-v3-turbo | $0.04 |
| Groq | whisper-large-v3 | $0.111 |
| OpenAI | gpt-4o-mini-transcribe | $0.18 ($0.003/min) |
| OpenAI | gpt-transcribe | $0.27 ($0.0045/min) |
| OpenAI | Whisper / gpt-4o-transcribe | $0.36 ($0.006/min) |
| Your own computer | faster-whisper (open source) | $0, but slow without a GPU |
For the summary step this guide uses openai/gpt-oss-120b on Groq at $0.15 per 1M input tokens and $0.60 per 1M output tokens. (Older tutorials use llama-3.3-70b-versatile. Groq shut that model down for free and developer accounts on 16 August 2026, so code that uses it now fails.) You only need one API key for the whole thing.
Groq has a free tier with rate limits that is plenty for testing and light personal use. At the time of writing, the Whisper limits are 7,200 audio seconds per hour and 28,800 per day (2 hours and 8 hours of audio). The catch is the summary model: on the free tier gpt-oss-120b allows only 8,000 tokens per minute, so a summary request for anything much longer than about 15 minutes of speech will be rejected as too large. Short voice notes and short calls work on the free tier. For hour-long meetings, add a payment method to move to the Developer tier. Your current limits are at console.groq.com/settings/limits. Groq also bills each transcription request as at least 10 seconds, which matters if you feed it hundreds of two-second voice notes.
How it compares to subscriptions: Otter.ai Pro is listed at $16.99/month billed monthly ($8.33/month billed annually) for 1,200 minutes. At Groq turbo rates, 20 hours of audio costs about $0.80. What Otter gives you for the money is a polished app, live transcription and calendar bots that join your meetings. If you want those, pay for them. If you just want transcripts and summaries of files you already have, a script does it for pennies.
Step 1: Install ffmpeg and set up the project
# Ubuntu/Debian
sudo apt install -y ffmpeg
# macOS
brew install ffmpeg
# Windows
winget install Gyan.FFmpeg
Then:
mkdir voicenotes && cd voicenotes
python3 -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install groq
mkdir inbox notes
export GROQ_API_KEY="paste-your-key-here" # Windows PowerShell: $env:GROQ_API_KEY="..."
Check that ffmpeg works with ffmpeg -version.
Step 2: Why we shrink the audio first
Groq’s free tier accepts files up to 25 MB per request (paid accounts get up to 100 MB). A one-hour Zoom recording at normal quality is often 50–60 MB. Whisper processes audio at 16 kHz mono internally anyway, so sending a high-quality stereo file just wastes upload time and hits the size cap sooner.
The script converts everything to 16 kHz, mono, 32 kbps MP3. That works out to roughly 14.4 MB per hour (32,000 bits × 3,600 seconds ÷ 8). It also splits recordings into 20-minute chunks (about 4.8 MB each), so long meetings never hit the limit and one failed request doesn’t lose the whole file.
You can test the conversion by hand:
ffmpeg -i inbox/meeting.m4a -ac 1 -ar 16000 -b:a 32k \
-f segment -segment_time 1200 -reset_timestamps 1 /tmp/chunk_%03d.mp3
Step 3: The script
Save this as transcribe.py:
import json, os, shutil, subprocess, sys, tempfile
from pathlib import Path
from groq import Groq
STT_MODEL = os.environ.get("STT_MODEL", "whisper-large-v3-turbo")
LLM_MODEL = os.environ.get("LLM_MODEL", "openai/gpt-oss-120b")
LANGUAGE = os.environ.get("STT_LANGUAGE", "id") # "id", "en", or "" to auto-detect
CHUNK_SECONDS = 1200
INBOX, NOTES = Path("inbox"), Path("notes")
AUDIO_EXT = {".mp3", ".m4a", ".wav", ".ogg", ".opus", ".mp4", ".webm", ".aac", ".flac"}
# Names, products and jargon Whisper would otherwise misspell. Keep it short.
VOCAB = os.environ.get("STT_VOCAB", "Solvinc, QRIS, Tokopedia, Shopee, invoice, PPN")
client = Groq()
def split_audio(src: Path, workdir: Path) -> list[Path]:
"""Convert to 16 kHz mono 32 kbps MP3 and split into chunks."""
pattern = workdir / "chunk_%03d.mp3"
subprocess.run(
["ffmpeg", "-loglevel", "error", "-y", "-i", str(src),
"-vn", "-ac", "1", "-ar", "16000", "-b:a", "32k",
"-f", "segment", "-segment_time", str(CHUNK_SECONDS),
"-reset_timestamps", "1", str(pattern)],
check=True,
)
return sorted(workdir.glob("chunk_*.mp3"))
def transcribe_chunk(path: Path) -> list[dict]:
with open(path, "rb") as f:
kwargs = dict(
file=(path.name, f.read()),
model=STT_MODEL,
response_format="verbose_json",
temperature=0.0,
prompt=VOCAB,
)
if LANGUAGE:
kwargs["language"] = LANGUAGE
result = client.audio.transcriptions.create(**kwargs)
data = result if isinstance(result, dict) else result.model_dump()
return data.get("segments") or [{"start": 0, "end": 0, "text": data["text"]}]
def ts(seconds: float, srt: bool = False) -> str:
h, rem = divmod(int(seconds), 3600)
m, s = divmod(rem, 60)
if srt:
ms = int((seconds - int(seconds)) * 1000)
return f"{h:02}:{m:02}:{s:02},{ms:03}"
return f"{h:02}:{m:02}:{s:02}"
def transcribe(src: Path) -> list[dict]:
segments = []
with tempfile.TemporaryDirectory() as tmp:
for i, chunk in enumerate(split_audio(src, Path(tmp))):
offset = i * CHUNK_SECONDS
print(f" chunk {i + 1}: {chunk.name}")
for seg in transcribe_chunk(chunk):
segments.append({
"start": seg["start"] + offset,
"end": seg["end"] + offset,
"text": seg["text"].strip(),
})
return segments
SUMMARY_PROMPT = """Below is a timestamped transcript of a business conversation or voice note.
It may mix Indonesian and English. Write in the main language of the transcript.
Return Markdown with exactly these sections:
## Summary
3-6 bullet points.
## Decisions
Bullets. Only things clearly agreed. Write "None" if there are none.
## Action items
Bullets as "- [ ] who: what (by when, if said) [timestamp]". Write "None" if there are none.
## Numbers mentioned
Every price, quantity, date and deadline, each with its [timestamp].
Rules: do not invent names, numbers or deadlines. If something is unclear, say "unclear"
and give the timestamp so a human can check the audio."""
def summarise(transcript: str) -> str:
resp = client.chat.completions.create(
model=LLM_MODEL,
temperature=0.2,
reasoning_effort="low", # gpt-oss option; remove it if you switch to a model without it
messages=[
{"role": "system", "content": SUMMARY_PROMPT},
{"role": "user", "content": transcript},
],
)
return resp.choices[0].message.content
def process(src: Path) -> None:
print(f"→ {src.name}")
segments = transcribe(src)
if not segments:
print(" no speech found, skipping")
return
transcript = "\n".join(f"[{ts(s['start'])}] {s['text']}" for s in segments)
try:
summary = summarise(transcript)
except Exception as e: # don't lose a paid-for transcript over a summary error
summary = f"_Summary failed: {e}_"
print(f" summary failed, saving transcript only: {e}")
stem = NOTES / src.stem
stem.with_suffix(".md").write_text(
f"# {src.name}\n\n{summary}\n\n---\n\n## Full transcript\n\n{transcript}\n",
encoding="utf-8",
)
srt = "\n".join(
f"{i}\n{ts(s['start'], True)} --> {ts(s['end'], True)}\n{s['text']}\n"
for i, s in enumerate(segments, 1)
)
stem.with_suffix(".srt").write_text(srt, encoding="utf-8")
stem.with_suffix(".json").write_text(json.dumps(segments, ensure_ascii=False), encoding="utf-8")
done = INBOX / "done"
done.mkdir(exist_ok=True)
shutil.move(str(src), done / src.name)
minutes = segments[-1]["end"] / 60
print(f" {minutes:.1f} min → {stem.with_suffix('.md')}")
if __name__ == "__main__":
NOTES.mkdir(exist_ok=True)
files = [Path(p) for p in sys.argv[1:]] or sorted(
p for p in INBOX.iterdir() if p.suffix.lower() in AUDIO_EXT
)
for f in files:
try:
process(f)
except Exception as e: # keep going; the file stays in inbox for a retry
print(f" FAILED: {e}")
Run it:
cp ~/Downloads/supplier-call.m4a inbox/
python transcribe.py
You get three files in notes/:
supplier-call.md: summary, decisions, action items, numbers with timestamps, then the full transcriptsupplier-call.srt: subtitles you can load into VLC, CapCut or YouTubesupplier-call.json: raw segments, in case you want to build search later
The timestamps on every number are the important part. When the summary says “price agreed: Rp 42.000/kg [00:17:32]”, you can jump to 17:32 and listen for ten seconds instead of trusting the model.
Step 4: Getting WhatsApp voice notes out of your phone
WhatsApp stores voice notes as .opus files, which ffmpeg reads directly.
- One voice note: long-press it → Share (Android) or Forward → Share (iPhone) → save to Google Drive, Files, or send it to yourself on Telegram, then download it on your computer.
- A whole chat: open the chat → ⋮ → More → Export chat → Include media. You get a
.zipwith all the.opusfiles and a.txtof the messages. Unzip the audio intoinbox/. - Android only: voice notes also sit in
Android/media/com.whatsapp/WhatsApp/Media/WhatsApp Voice Notes/, which you can sync with Syncthing or copy over USB.
Short voice notes are where the 10-second minimum billing shows up. A batch of 200 notes of five seconds each is billed as 2,000 seconds (about 33 minutes), or roughly 2 cents. It’s still cheap, just not as cheap as the raw maths suggests.
Step 5: Make it automatic (optional)
On Linux or macOS, run it every 15 minutes with cron (crontab -e):
*/15 * * * * cd /home/you/voicenotes && GROQ_API_KEY=xxx .venv/bin/python transcribe.py >> run.log 2>&1
Point Google Drive for desktop, Dropbox or Syncthing at the inbox/ folder and the whole flow becomes: share audio from phone to a Drive folder, notes appear on your computer a few minutes later. On Windows, use Task Scheduler with the same command.
Want it fully offline? Use faster-whisper
If the audio is sensitive (HR conversations, medical, legal) and you don’t want it leaving your machine, swap the API call for faster-whisper, an open-source reimplementation that runs Whisper locally:
pip install faster-whisper
from faster_whisper import WhisperModel
model = WhisperModel("large-v3-turbo", device="cpu", compute_type="int8") # "cuda" + "float16" with an NVIDIA GPU
segments, info = model.transcribe("inbox/meeting.m4a", language="id", vad_filter=True)
for s in segments:
print(f"[{s.start:.0f}s] {s.text}")
It costs nothing per minute, but speed depends heavily on your hardware. With an NVIDIA GPU it runs many times faster than real time. On a laptop CPU, large-v3-turbo can take roughly as long as the recording itself, or longer. The small model is much faster but noticeably worse on Indonesian. Time one 10-minute file before you decide. For the summary step offline, you’d need a local LLM through Ollama, which is a separate project.
Honest cost breakdown
A small business owner recording client and supplier calls, plus a lot of voice notes:
| Usage per month | Transcription (Groq turbo) | Summaries (gpt-oss-120b) | Total |
|---|---|---|---|
| 10 hours of calls | $0.40 | ~$0.05 | ~$0.45 |
| 300 voice notes, avg 30 s | ~$0.10 | ~$0.15 | ~$0.25 |
| 40 hours (heavy, e.g. a podcaster) | $1.60 | ~$0.20 | ~$1.80 |
The summary estimate assumes roughly 15–20k input tokens for a one-hour transcript plus a few thousand output and reasoning tokens. Many short voice notes cost more to summarise than their length suggests, because each one sends the full prompt and gets its own reply. Indonesian uses more tokens per word than English, so budget at the high end. Even if you switch both steps to OpenAI, the 10-hour row comes to about $2–4.
Time saved: writing up a 45-minute call properly takes most people 20–30 minutes, and most people skip it. With the script it takes the 2 minutes you spend reading and correcting the summary. At four calls a week, that’s 5–7 hours a month back, and you have a written record of what was agreed.
Can you make money with this?
Yes, but be realistic about what people will pay for. Raw transcripts are close to free now: CapCut, YouTube and Google Recorder (on Pixel phones) all do automatic captions at no cost. Nobody will pay you to press a button they already have.
What people do pay for is the cleanup and packaging on top:
- Edited subtitles for YouTubers and podcasters, especially Indonesian–English bilingual content where auto-captions do badly. You run the script, then spend 20–40 minutes per hour of audio fixing names, slang and line breaks. Upwork and Fiverr list plenty of captioning and transcription gigs; check current rates in your target market before you set prices.
- Meeting minutes for small organisations: RT/RW meetings, schools, cooperatives, churches and mosque committees that need written minutes (notulen) in a set format. Change
SUMMARY_PROMPTto match their template. - Setting this up for a client: the same “inbox folder → notes” flow, installed on an office PC or a small VPS, as a one-off job.
The AI cost is a rounding error in all of these. You’re being paid for accuracy and a format the client can use as-is, and that comes from the human review step.
Reality check
Whisper hallucinates on silence and noise. Long pauses, hold music or background noise can produce invented sentences, repeated lines, or things like “Terima kasih telah menonton” that nobody said. Trim long silences, and be suspicious of a sentence that repeats several times in a row. The local vad_filter=True option above helps a lot.
Speaker labels are not included. Whisper does not tell you who said what. For a two-person call the LLM can often guess from context, but it will sometimes guess wrong. If you need reliable speaker labels, you have to add a diarization tool such as pyannote, or use a paid service that includes it (AssemblyAI and Deepgram both offer it, at a higher per-minute price).
Accuracy drops with bad audio. A phone on the table in a noisy warung, speakerphone calls, or six people talking over each other will produce a messy transcript. A cheap clip-on mic or simply holding the phone closer makes a bigger difference than switching models.
Regional languages are weaker. Indonesian and English work well. Javanese, Sundanese and heavy slang are hit and miss. Test with your real recordings, not a clean demo clip.
Never trust numbers without checking. “Lima belas” vs “lima puluh” is exactly the kind of mistake that costs money. That’s why the prompt asks for a timestamp next to every number. Check prices and deadlines against the audio before you act on them.
Get consent before recording. Tell people you’re recording calls, and treat the transcripts like any other customer data. Indonesia’s Personal Data Protection Law (UU PDP, No. 27/2022) applies to business records like these. Also read Groq’s data policy before sending sensitive audio. If that’s a concern, use the faster-whisper route.
Free tier limits will bite during a backlog. If you dump three months of recordings in at once, you’ll hit the hourly audio limit (about 2 hours of audio per hour on the free tier) and get 429 errors. The script leaves failed files in inbox/ so the next run picks them up, or you can add a payment method to move to Groq’s higher-limit developer tier.
Model names change. Groq has retired models before. If you get a “model not found” error, check console.groq.com/docs/models and update STT_MODEL or LLM_MODEL. No code changes needed.
Wrapping up
This is one of the most boring AI projects you can build, and probably one of the most useful. For less than a dollar a month you get a written, searchable record of every call and voice note, with the numbers marked so you can check them. Start with the last week of supplier calls and voice notes, read what comes out, and adjust the vocabulary list and summary prompt to fit how your business talks.