Every small business has the same pile somewhere: thermal-paper receipts from the supplier, Indomaret runs for office supplies, fuel slips, a handwritten nota from the bengkel. At the end of the month somebody types them into a spreadsheet, or nobody does and the books are a guess.
Typing a receipt with line items into a sheet takes about one to three minutes once you include finding it, reading the faded print, and fixing typos. Time yourself on ten receipts, because that number decides whether this project is worth it for you. At 300 receipts a month and two minutes each, that’s 10 hours of data entry.
In this guide you’ll build a small Python tool that:
- watches an
inbox/folder of receipt photos - sends each photo to Google’s Gemini API and gets back structured JSON (merchant, date, total, tax, line items, category)
- checks the maths so wrong reads get flagged, not trusted
- appends the results to a CSV, and optionally to Google Sheets
- moves processed photos to an archive so nothing is counted twice
The AI cost is roughly $0.002 per receipt, so under $1 for 300 receipts. Build time is about an hour.
What you need
| Item | Cost | Notes |
|---|---|---|
| Python 3.10+ | $0 | Laptop, mini PC, or any VPS |
| Gemini API key | $0 to start | From aistudio.google.com, “Get API key” |
google-genai, pydantic, pillow, pillow-heif | $0 | pip install (pillow-heif lets it read iPhone HEIC photos) |
gspread (optional) | $0 | Only if you want Google Sheets output |
| Phone camera | $0 | Any phone from the last 5 years is fine |
| A way to get photos onto the computer | $0 | Google Drive for desktop, Syncthing, or a USB cable |
Pricing as of October 2026 (check ai.google.dev/pricing before you commit, because Google changes model names and prices often):
- Gemini 2.5 Flash (paid tier): $0.30 per 1M input tokens (text/image), $2.50 per 1M output tokens
- Gemini 2.5 Flash-Lite (paid tier): $0.10 per 1M input, $0.40 per 1M output
- Free tier: both models have a rate-limited free tier, which is plenty for testing. Read the privacy note in the Reality check before you put real financial documents through it.
Google’s docs now feature newer Gemini 3.x models, but gemini-2.5-flash and gemini-2.5-flash-lite are still listed as stable with no shutdown date announced, and they remain among the cheapest options for this job. If you switch to a 3.x model, change the model string and the thinking setting: 3.x models are controlled with thinking_level rather than thinking_budget (see Reality check).
How the token maths works
Gemini charges for images by tokens. Under Gemini 2.x rules, an image that is 384 px or smaller on both sides costs 258 tokens. Bigger images are split into 768×768 tiles and each tile costs 258 tokens. A receipt photo resized to 1536 px on the long side comes out at roughly 4–8 tiles, so about 1,000–2,000 input tokens. Very long, narrow supermarket receipts can reach 10–12 tiles (~3,000 tokens), which adds less than a tenth of a cent. Add about 300 tokens of prompt and about 400 tokens of JSON output.
Per receipt on Gemini 2.5 Flash:
- Input: ~2,000 tokens × $0.30/1M = $0.0006
- Output: ~400 tokens × $2.50/1M = $0.0010
- Total: about $0.0016, call it $0.002
One catch: Gemini 2.5 Flash “thinks” by default, and thinking tokens are billed as output. For reading a receipt you don’t need it, so the script sets thinking_budget=0. Leave it on and your per-receipt cost can go up several times.
Step 1: Set up the project
mkdir receipts && cd receipts
python3 -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install google-genai pydantic pillow pillow-heif gspread
mkdir inbox processed review
export GEMINI_API_KEY="paste-your-key-here" # Windows PowerShell: $env:GEMINI_API_KEY="..."
The folder layout:
receipts/
├── inbox/ # drop new photos here
├── processed/ # photos that were read and passed the checks
├── review/ # photos that need a human to look
├── receipts.csv # one row per receipt
├── items.csv # one row per line item
└── scan.py
On Windows, $env: only lasts for the current window. To make the key permanent, also run setx GEMINI_API_KEY "..." and open a new terminal.
Step 2: Define what you want back
The most important trick is structured output. Instead of asking for “the receipt details” and parsing free text, you give Gemini a schema. The API then returns JSON that matches it. With the google-genai SDK you can pass a Pydantic model directly.
from pydantic import BaseModel
class LineItem(BaseModel):
description: str
qty: float
unit_price: float
total: float
class Receipt(BaseModel):
merchant: str
date: str # YYYY-MM-DD
currency: str # ISO code, e.g. IDR, USD
subtotal: float # 0 if not printed
tax: float # 0 if not printed (PPN)
other_charges: float # service charge, delivery, rounding; negative for discounts
total: float
payment_method: str # cash, card, qris, transfer, unknown
category: str
items: list[LineItem]
legible: bool # model's own judgement: is the photo readable?
Step 3: The script
Save this as scan.py:
import csv, hashlib, os, shutil, sys
from io import BytesIO
from pathlib import Path
from google import genai
from google.genai import types
from PIL import Image, ImageOps
from pydantic import BaseModel
try: # iPhone photos are HEIC by default
from pillow_heif import register_heif_opener
register_heif_opener()
except ImportError:
pass
MODEL = os.environ.get("RECEIPT_MODEL", "gemini-2.5-flash")
CATEGORIES = ["inventory", "supplies", "fuel_transport", "food_meals",
"utilities", "rent", "services", "other"]
INBOX, DONE, REVIEW = Path("inbox"), Path("processed"), Path("review")
SEEN_FILE = Path("seen_hashes.txt")
class LineItem(BaseModel):
description: str
qty: float
unit_price: float
total: float
class Receipt(BaseModel):
merchant: str
date: str
currency: str
subtotal: float
tax: float
other_charges: float
total: float
payment_method: str
category: str
items: list[LineItem]
legible: bool
PROMPT = f"""You are extracting data from a photo of a purchase receipt or invoice.
Rules:
- Copy numbers exactly as printed. Return plain numbers: Rp 25.000 -> 25000, $4.50 -> 4.5.
In Indonesian receipts '.' is usually a thousands separator.
- date as YYYY-MM-DD. If the year is missing or unreadable, use "unknown".
- If subtotal or tax is not printed, return 0.
- other_charges = service charge, delivery, rounding (pembulatan); discounts are negative.
- category must be one of: {", ".join(CATEGORIES)}.
- payment_method: cash, card, qris, transfer, or unknown.
- Never guess a number you cannot read. If key fields are unreadable, set legible=false.
"""
client = genai.Client() # reads GEMINI_API_KEY
def load_image(path: Path) -> bytes:
img = ImageOps.exif_transpose(Image.open(path)).convert("RGB")
img.thumbnail((1536, 1536)) # caps tokens and upload size
buf = BytesIO()
img.save(buf, format="JPEG", quality=85)
return buf.getvalue()
def extract(path: Path) -> Receipt:
resp = client.models.generate_content(
model=MODEL,
contents=[types.Part.from_bytes(data=load_image(path), mime_type="image/jpeg"),
PROMPT],
config=types.GenerateContentConfig(
response_mime_type="application/json",
response_schema=Receipt,
temperature=0,
thinking_config=types.ThinkingConfig(thinking_budget=0),
),
)
return resp.parsed
def problems(r: Receipt) -> list[str]:
issues = []
if not r.legible:
issues.append("model says not legible")
if r.date == "unknown":
issues.append("no date")
if r.category not in CATEGORIES:
issues.append(f"bad category {r.category}")
tol = max(1.0, r.total * 0.01) # 1% or 1 unit, absorbs rounding
items_sum = sum(i.total for i in r.items)
exclusive_ok = abs(items_sum + r.tax + r.other_charges - r.total) <= tol
inclusive_ok = abs(items_sum - r.total) <= tol # tax already inside item prices
if r.items and not (exclusive_ok or inclusive_ok):
issues.append(f"items {items_sum:.0f} + tax/charges != total {r.total:.0f}")
for i in r.items:
if abs(i.qty * i.unit_price - i.total) > max(1.0, i.total * 0.01):
issues.append(f"line '{i.description[:20]}' qty x price != total")
return issues
def append_csv(path: str, row: dict):
new = not Path(path).exists()
with open(path, "a", newline="", encoding="utf-8") as f:
w = csv.DictWriter(f, fieldnames=row.keys())
if new:
w.writeheader()
w.writerow(row)
def main():
seen = set(SEEN_FILE.read_text().split()) if SEEN_FILE.exists() else set()
files = sorted(p for p in INBOX.iterdir()
if p.suffix.lower() in {".jpg", ".jpeg", ".png", ".webp", ".heic"})
for path in files:
digest = hashlib.sha256(path.read_bytes()).hexdigest()[:16]
if digest in seen:
print(f"skip duplicate {path.name}")
shutil.move(path, REVIEW / f"DUPLICATE_{path.name}")
continue
try:
r = extract(path)
except Exception as e: # network, quota, bad file
print(f"ERROR {path.name}: {e}", file=sys.stderr)
continue # leave it in inbox, retry next run
if r is None: # response didn't match the schema
print(f"ERROR {path.name}: unparseable response", file=sys.stderr)
continue
issues = problems(r)
status = "review" if issues else "ok"
append_csv("receipts.csv", {
"id": digest, "file": path.name, "status": status,
"issues": "; ".join(issues), "merchant": r.merchant, "date": r.date,
"currency": r.currency, "subtotal": r.subtotal, "tax": r.tax,
"other_charges": r.other_charges, "total": r.total,
"payment_method": r.payment_method, "category": r.category,
})
for i in r.items:
append_csv("items.csv", {"receipt_id": digest, **i.model_dump()})
shutil.move(path, (REVIEW if issues else DONE) / f"{digest}_{path.name}")
seen.add(digest)
SEEN_FILE.write_text("\n".join(seen)) # save after each file, survives crashes
print(f"{status:6} {r.date} {r.merchant[:25]:25} {r.total:>12,.0f} {r.currency}")
if __name__ == "__main__":
main()
Drop 5–10 receipt photos in inbox/ and run:
python scan.py
Output looks like this:
ok 2026-10-02 INDOMARET JL RAYA BOGOR 87,500 IDR
ok 2026-10-03 TB SUMBER JAYA 1,240,000 IDR
review 2026-10-03 SPBU 34-16712 200,000 IDR
Anything marked review lands in the review/ folder with the reason in the issues column. You open the photo, fix the row, done. That’s where most of your remaining minutes go, and it’s the reason this tool is safe to use.
Step 4: Why the validation step matters
The model doesn’t “know” your receipt is wrong. It returns confident JSON either way. The problems() function catches the mistakes you’ll actually see:
- Thousands-separator errors:
25.000read as25.0. The line total stops matching qty × price, so it gets flagged. - Missed lines on long supermarket receipts: items no longer add up to the total.
- Faded thermal paper: the model sets
legible=false, or the sums break. - Tax-inclusive vs. tax-exclusive receipts: the check accepts either, because Indonesian retail receipts often print prices that already include PPN.
Always trust the printed total over the line items. For bookkeeping, the total is what matters. The line items are a bonus for inventory and price tracking.
Step 5 (optional): Push to Google Sheets
If the business owner lives in Google Sheets:
- In Google Cloud Console, create a project, enable the Google Sheets API, create a service account and download its JSON key as
sa.json. - Share your target spreadsheet with the service account’s email (it looks like
name@project.iam.gserviceaccount.com) as an Editor. - Create a tab called
receiptswith the same header row asreceipts.csv.
Save this as sync_sheet.py:
import csv, gspread
gc = gspread.service_account(filename="sa.json")
ws = gc.open_by_key("YOUR_SHEET_ID").worksheet("receipts")
existing_ids = set(ws.col_values(1)[1:]) # column A = id
with open("receipts.csv", encoding="utf-8") as f:
rows = list(csv.reader(f))[1:] # skip header
new_rows = [r for r in rows if r[0] not in existing_ids]
for r in new_rows:
r[0] = "'" + r[0] # keep ids as text; an all-digit id would otherwise become a number
if new_rows:
ws.append_rows(new_rows, value_input_option="USER_ENTERED")
print(f"pushed {len(new_rows)} rows")
Because it matches on the id column, you can run it as often as you like without creating duplicates.
Step 6: Make it hands-off
Getting photos in: the easiest route on Windows or macOS is Google Drive for desktop (there is no official Linux client; use Syncthing or rclone there). Make a Drive folder called receipts-inbox, point INBOX at the synced local path, and anyone on the team can upload from the Drive app on their phone. Syncthing (free, open source) does the same thing without the cloud.
Running it on a schedule (Linux/macOS, every evening at 21:00):
crontab -e
# add:
0 21 * * * cd /home/you/receipts && GEMINI_API_KEY=xxx .venv/bin/python scan.py >> scan.log 2>&1 && .venv/bin/python sync_sheet.py >> scan.log 2>&1
On Windows, use Task Scheduler with .venv\Scripts\python.exe scan.py as the action and the receipts folder as “Start in”.
Honest cost breakdown
Scenario: a small shop with 300 receipts/month.
| Item | Monthly |
|---|---|
| Gemini 2.5 Flash, ~300 × $0.0016 | ~$0.50 |
| Re-runs, failed calls, testing | ~$0.20 |
| Google Sheets, Drive, Python | $0 |
| Computer it runs on | $0 if existing; ~$5 if you rent a small VPS |
| Total | under $1 (or ~$6 with a VPS) |
Switching to Flash-Lite cuts the AI line to about a third. Test it on your own receipts first. On blurry or handwritten notas the cheaper model gets more of them wrong, and every extra trip to review/ costs you more in time than the cents you saved.
Time: if manual entry was 2 minutes per receipt (10 hours/month) and about 10–15% of receipts still need review at around 1 minute each, you’re down to roughly 30–45 minutes a month, plus snapping the photos. These are illustrative numbers. Measure your own for one month before you promise anyone anything.
Turning it into money
This works as a paid service because the buyer can see the result: a clean spreadsheet instead of a shoebox. None of the figures below are guaranteed. They’re a way to work out a price, not proven income.
1. Setup for local businesses. Look for anyone with lots of small paper receipts and no bookkeeping software: workshops, building-supply shops, caterers, small restaurants, contractors. You install the folder, the script and the Sheet on their computer, on their own Gemini key so the API bill and the data stay theirs, and teach one staff member the photo routine. That’s 2–4 hours of your time.
2. Monthly upkeep. You check the review/ folder, fix mistakes, and send a monthly expense summary by category. That’s a bookkeeping-assistant job done in a fraction of the hours.
How to price it: start from the time you save the client, not from your API cost. If they currently pay staff for 10 hours of data entry a month, your monthly fee has to come in clearly below that and still be worth your 1–2 hours. Work it out with local wages. Don’t copy a dollar figure from a blog, including this one.
3. Accountants as a channel. Small accounting practices (konsultan pajak, freelance bookkeepers) get receipts from clients as messy photos every month. One accountant with 20 small clients is a better lead than 20 separate shops. Demo it on their own receipts.
Realistic expectations: getting a first paying client usually takes weeks of showing the demo in person, and it may not happen at all. “I’ll turn your receipts into a spreadsheet, here’s last month’s pile done” sells. “AI-powered OCR pipeline” doesn’t.
Reality check: what does NOT work
Skipping human review. Most clear receipts come out right. Not all of them. A clean printed Alfamart receipt is easy. A crumpled, sun-faded fuel slip or a handwritten nota with a scribbled total is not. The validation step plus a human looking at review/ is the product. Without it, you’re just making wrong numbers faster.
Putting client financial data through the free tier. Under Google’s Gemini API terms, content sent on the unpaid tier can be used to improve Google’s products and may be seen by human reviewers. The paid tier doesn’t do this. For your own coffee receipts, the free tier is fine. For a client’s supplier invoices, turn on billing. It costs under a dollar a month.
Bad photos. Most “the AI is bad” complaints are really photo problems: shadows, a curled receipt, flash glare. Lay it flat, use daylight, fill the frame. Photograph thermal receipts the same day, because thermal print fades, faster in a hot car or a sunny drawer.
Treating this as accounting. This gives you clean expense data. It doesn’t make the data tax-compliant, decide what’s deductible, or replace an accountant for tax filings. Keep the original photos (the processed/ folder) as evidence and back them up.
Leaving thinking on and not resizing. A full-size 12-megapixel photo with thinking enabled costs several times more per receipt for no accuracy gain on this task. It’s still cheap, but across several clients’ bills it adds up.
Assuming duplicates are caught. Two staff members snapping the same nota happens all the time. The SHA-256 hash only catches identical files, so a second photo of the same receipt gets through. Once a month, sort the sheet by date + merchant + total and look for twins.
Hard-coding the model name forever. Google retires model versions on a schedule (check ai.google.dev/gemini-api/docs/deprecations). When gemini-2.5-flash gets a shutdown date, set the RECEIPT_MODEL environment variable to its replacement. For Gemini 3.x models, also swap thinking_budget=0 for the lowest thinking_level that model supports (for example thinking_level="minimal" or "low"). Then run 10 test receipts and compare against old results before you switch a client over.
Final checklist
- Script runs on 10 of your own receipts, results checked by hand
-
thinking_budget=0and image resize in place - Billing enabled before any client data goes through
-
review/folder checked weekly by a named person - Original photos backed up off the machine
- Cron or Task Scheduler job running,
scan.logchecked after the first week - Monthly duplicate check on date + merchant + total
Run it on your own receipts for a month. If it saves you hours and the review pile stays small, you’ve got a tool for yourself and a demo you can show a business owner.