Every small business has the same pile somewhere: thermal-paper receipts from the supplier, Indomaret runs for office supplies, fuel slips, a handwritten nota from the bengkel. At the end of the month somebody types them into a spreadsheet, or nobody does and the books are a guess.

Typing a receipt with line items into a sheet takes about one to three minutes once you include finding it, reading the faded print, and fixing typos. Time yourself on ten receipts, because that number decides whether this project is worth it for you. At 300 receipts a month and two minutes each, that’s 10 hours of data entry.

In this guide you’ll build a small Python tool that:

  • watches an inbox/ folder of receipt photos
  • sends each photo to Google’s Gemini API and gets back structured JSON (merchant, date, total, tax, line items, category)
  • checks the maths so wrong reads get flagged, not trusted
  • appends the results to a CSV, and optionally to Google Sheets
  • moves processed photos to an archive so nothing is counted twice

The AI cost is roughly $0.002 per receipt, so under $1 for 300 receipts. Build time is about an hour.

What you need

ItemCostNotes
Python 3.10+$0Laptop, mini PC, or any VPS
Gemini API key$0 to startFrom aistudio.google.com, “Get API key”
google-genai, pydantic, pillow, pillow-heif$0pip install (pillow-heif lets it read iPhone HEIC photos)
gspread (optional)$0Only if you want Google Sheets output
Phone camera$0Any phone from the last 5 years is fine
A way to get photos onto the computer$0Google Drive for desktop, Syncthing, or a USB cable

Pricing as of October 2026 (check ai.google.dev/pricing before you commit, because Google changes model names and prices often):

  • Gemini 2.5 Flash (paid tier): $0.30 per 1M input tokens (text/image), $2.50 per 1M output tokens
  • Gemini 2.5 Flash-Lite (paid tier): $0.10 per 1M input, $0.40 per 1M output
  • Free tier: both models have a rate-limited free tier, which is plenty for testing. Read the privacy note in the Reality check before you put real financial documents through it.

Google’s docs now feature newer Gemini 3.x models, but gemini-2.5-flash and gemini-2.5-flash-lite are still listed as stable with no shutdown date announced, and they remain among the cheapest options for this job. If you switch to a 3.x model, change the model string and the thinking setting: 3.x models are controlled with thinking_level rather than thinking_budget (see Reality check).

How the token maths works

Gemini charges for images by tokens. Under Gemini 2.x rules, an image that is 384 px or smaller on both sides costs 258 tokens. Bigger images are split into 768×768 tiles and each tile costs 258 tokens. A receipt photo resized to 1536 px on the long side comes out at roughly 4–8 tiles, so about 1,000–2,000 input tokens. Very long, narrow supermarket receipts can reach 10–12 tiles (~3,000 tokens), which adds less than a tenth of a cent. Add about 300 tokens of prompt and about 400 tokens of JSON output.

Per receipt on Gemini 2.5 Flash:

  • Input: ~2,000 tokens × $0.30/1M = $0.0006
  • Output: ~400 tokens × $2.50/1M = $0.0010
  • Total: about $0.0016, call it $0.002

One catch: Gemini 2.5 Flash “thinks” by default, and thinking tokens are billed as output. For reading a receipt you don’t need it, so the script sets thinking_budget=0. Leave it on and your per-receipt cost can go up several times.

Step 1: Set up the project

mkdir receipts && cd receipts
python3 -m venv .venv && source .venv/bin/activate   # Windows: .venv\Scripts\activate
pip install google-genai pydantic pillow pillow-heif gspread

mkdir inbox processed review
export GEMINI_API_KEY="paste-your-key-here"           # Windows PowerShell: $env:GEMINI_API_KEY="..."

The folder layout:

receipts/
├── inbox/        # drop new photos here
├── processed/    # photos that were read and passed the checks
├── review/       # photos that need a human to look
├── receipts.csv  # one row per receipt
├── items.csv     # one row per line item
└── scan.py

On Windows, $env: only lasts for the current window. To make the key permanent, also run setx GEMINI_API_KEY "..." and open a new terminal.

Step 2: Define what you want back

The most important trick is structured output. Instead of asking for “the receipt details” and parsing free text, you give Gemini a schema. The API then returns JSON that matches it. With the google-genai SDK you can pass a Pydantic model directly.

from pydantic import BaseModel

class LineItem(BaseModel):
    description: str
    qty: float
    unit_price: float
    total: float

class Receipt(BaseModel):
    merchant: str
    date: str            # YYYY-MM-DD
    currency: str        # ISO code, e.g. IDR, USD
    subtotal: float      # 0 if not printed
    tax: float           # 0 if not printed (PPN)
    other_charges: float # service charge, delivery, rounding; negative for discounts
    total: float
    payment_method: str  # cash, card, qris, transfer, unknown
    category: str
    items: list[LineItem]
    legible: bool        # model's own judgement: is the photo readable?

Step 3: The script

Save this as scan.py:

import csv, hashlib, os, shutil, sys
from io import BytesIO
from pathlib import Path

from google import genai
from google.genai import types
from PIL import Image, ImageOps
from pydantic import BaseModel

try:                                     # iPhone photos are HEIC by default
    from pillow_heif import register_heif_opener
    register_heif_opener()
except ImportError:
    pass

MODEL = os.environ.get("RECEIPT_MODEL", "gemini-2.5-flash")
CATEGORIES = ["inventory", "supplies", "fuel_transport", "food_meals",
              "utilities", "rent", "services", "other"]
INBOX, DONE, REVIEW = Path("inbox"), Path("processed"), Path("review")
SEEN_FILE = Path("seen_hashes.txt")

class LineItem(BaseModel):
    description: str
    qty: float
    unit_price: float
    total: float

class Receipt(BaseModel):
    merchant: str
    date: str
    currency: str
    subtotal: float
    tax: float
    other_charges: float
    total: float
    payment_method: str
    category: str
    items: list[LineItem]
    legible: bool

PROMPT = f"""You are extracting data from a photo of a purchase receipt or invoice.
Rules:
- Copy numbers exactly as printed. Return plain numbers: Rp 25.000 -> 25000, $4.50 -> 4.5.
  In Indonesian receipts '.' is usually a thousands separator.
- date as YYYY-MM-DD. If the year is missing or unreadable, use "unknown".
- If subtotal or tax is not printed, return 0.
- other_charges = service charge, delivery, rounding (pembulatan); discounts are negative.
- category must be one of: {", ".join(CATEGORIES)}.
- payment_method: cash, card, qris, transfer, or unknown.
- Never guess a number you cannot read. If key fields are unreadable, set legible=false.
"""

client = genai.Client()  # reads GEMINI_API_KEY

def load_image(path: Path) -> bytes:
    img = ImageOps.exif_transpose(Image.open(path)).convert("RGB")
    img.thumbnail((1536, 1536))           # caps tokens and upload size
    buf = BytesIO()
    img.save(buf, format="JPEG", quality=85)
    return buf.getvalue()

def extract(path: Path) -> Receipt:
    resp = client.models.generate_content(
        model=MODEL,
        contents=[types.Part.from_bytes(data=load_image(path), mime_type="image/jpeg"),
                  PROMPT],
        config=types.GenerateContentConfig(
            response_mime_type="application/json",
            response_schema=Receipt,
            temperature=0,
            thinking_config=types.ThinkingConfig(thinking_budget=0),
        ),
    )
    return resp.parsed

def problems(r: Receipt) -> list[str]:
    issues = []
    if not r.legible:
        issues.append("model says not legible")
    if r.date == "unknown":
        issues.append("no date")
    if r.category not in CATEGORIES:
        issues.append(f"bad category {r.category}")
    tol = max(1.0, r.total * 0.01)        # 1% or 1 unit, absorbs rounding
    items_sum = sum(i.total for i in r.items)
    exclusive_ok = abs(items_sum + r.tax + r.other_charges - r.total) <= tol
    inclusive_ok = abs(items_sum - r.total) <= tol   # tax already inside item prices
    if r.items and not (exclusive_ok or inclusive_ok):
        issues.append(f"items {items_sum:.0f} + tax/charges != total {r.total:.0f}")
    for i in r.items:
        if abs(i.qty * i.unit_price - i.total) > max(1.0, i.total * 0.01):
            issues.append(f"line '{i.description[:20]}' qty x price != total")
    return issues

def append_csv(path: str, row: dict):
    new = not Path(path).exists()
    with open(path, "a", newline="", encoding="utf-8") as f:
        w = csv.DictWriter(f, fieldnames=row.keys())
        if new:
            w.writeheader()
        w.writerow(row)

def main():
    seen = set(SEEN_FILE.read_text().split()) if SEEN_FILE.exists() else set()
    files = sorted(p for p in INBOX.iterdir()
                   if p.suffix.lower() in {".jpg", ".jpeg", ".png", ".webp", ".heic"})
    for path in files:
        digest = hashlib.sha256(path.read_bytes()).hexdigest()[:16]
        if digest in seen:
            print(f"skip duplicate {path.name}")
            shutil.move(path, REVIEW / f"DUPLICATE_{path.name}")
            continue
        try:
            r = extract(path)
        except Exception as e:                  # network, quota, bad file
            print(f"ERROR {path.name}: {e}", file=sys.stderr)
            continue                            # leave it in inbox, retry next run
        if r is None:                           # response didn't match the schema
            print(f"ERROR {path.name}: unparseable response", file=sys.stderr)
            continue
        issues = problems(r)
        status = "review" if issues else "ok"
        append_csv("receipts.csv", {
            "id": digest, "file": path.name, "status": status,
            "issues": "; ".join(issues), "merchant": r.merchant, "date": r.date,
            "currency": r.currency, "subtotal": r.subtotal, "tax": r.tax,
            "other_charges": r.other_charges, "total": r.total,
            "payment_method": r.payment_method, "category": r.category,
        })
        for i in r.items:
            append_csv("items.csv", {"receipt_id": digest, **i.model_dump()})
        shutil.move(path, (REVIEW if issues else DONE) / f"{digest}_{path.name}")
        seen.add(digest)
        SEEN_FILE.write_text("\n".join(seen))   # save after each file, survives crashes
        print(f"{status:6} {r.date} {r.merchant[:25]:25} {r.total:>12,.0f} {r.currency}")

if __name__ == "__main__":
    main()

Drop 5–10 receipt photos in inbox/ and run:

python scan.py

Output looks like this:

ok     2026-10-02 INDOMARET JL RAYA BOGOR         87,500 IDR
ok     2026-10-03 TB SUMBER JAYA               1,240,000 IDR
review 2026-10-03 SPBU 34-16712                  200,000 IDR

Anything marked review lands in the review/ folder with the reason in the issues column. You open the photo, fix the row, done. That’s where most of your remaining minutes go, and it’s the reason this tool is safe to use.

Step 4: Why the validation step matters

The model doesn’t “know” your receipt is wrong. It returns confident JSON either way. The problems() function catches the mistakes you’ll actually see:

  • Thousands-separator errors: 25.000 read as 25.0. The line total stops matching qty × price, so it gets flagged.
  • Missed lines on long supermarket receipts: items no longer add up to the total.
  • Faded thermal paper: the model sets legible=false, or the sums break.
  • Tax-inclusive vs. tax-exclusive receipts: the check accepts either, because Indonesian retail receipts often print prices that already include PPN.

Always trust the printed total over the line items. For bookkeeping, the total is what matters. The line items are a bonus for inventory and price tracking.

Step 5 (optional): Push to Google Sheets

If the business owner lives in Google Sheets:

  1. In Google Cloud Console, create a project, enable the Google Sheets API, create a service account and download its JSON key as sa.json.
  2. Share your target spreadsheet with the service account’s email (it looks like name@project.iam.gserviceaccount.com) as an Editor.
  3. Create a tab called receipts with the same header row as receipts.csv.

Save this as sync_sheet.py:

import csv, gspread

gc = gspread.service_account(filename="sa.json")
ws = gc.open_by_key("YOUR_SHEET_ID").worksheet("receipts")

existing_ids = set(ws.col_values(1)[1:])            # column A = id
with open("receipts.csv", encoding="utf-8") as f:
    rows = list(csv.reader(f))[1:]                   # skip header
new_rows = [r for r in rows if r[0] not in existing_ids]
for r in new_rows:
    r[0] = "'" + r[0]    # keep ids as text; an all-digit id would otherwise become a number
if new_rows:
    ws.append_rows(new_rows, value_input_option="USER_ENTERED")
print(f"pushed {len(new_rows)} rows")

Because it matches on the id column, you can run it as often as you like without creating duplicates.

Step 6: Make it hands-off

Getting photos in: the easiest route on Windows or macOS is Google Drive for desktop (there is no official Linux client; use Syncthing or rclone there). Make a Drive folder called receipts-inbox, point INBOX at the synced local path, and anyone on the team can upload from the Drive app on their phone. Syncthing (free, open source) does the same thing without the cloud.

Running it on a schedule (Linux/macOS, every evening at 21:00):

crontab -e
# add:
0 21 * * * cd /home/you/receipts && GEMINI_API_KEY=xxx .venv/bin/python scan.py >> scan.log 2>&1 && .venv/bin/python sync_sheet.py >> scan.log 2>&1

On Windows, use Task Scheduler with .venv\Scripts\python.exe scan.py as the action and the receipts folder as “Start in”.

Honest cost breakdown

Scenario: a small shop with 300 receipts/month.

ItemMonthly
Gemini 2.5 Flash, ~300 × $0.0016~$0.50
Re-runs, failed calls, testing~$0.20
Google Sheets, Drive, Python$0
Computer it runs on$0 if existing; ~$5 if you rent a small VPS
Totalunder $1 (or ~$6 with a VPS)

Switching to Flash-Lite cuts the AI line to about a third. Test it on your own receipts first. On blurry or handwritten notas the cheaper model gets more of them wrong, and every extra trip to review/ costs you more in time than the cents you saved.

Time: if manual entry was 2 minutes per receipt (10 hours/month) and about 10–15% of receipts still need review at around 1 minute each, you’re down to roughly 30–45 minutes a month, plus snapping the photos. These are illustrative numbers. Measure your own for one month before you promise anyone anything.

Turning it into money

This works as a paid service because the buyer can see the result: a clean spreadsheet instead of a shoebox. None of the figures below are guaranteed. They’re a way to work out a price, not proven income.

1. Setup for local businesses. Look for anyone with lots of small paper receipts and no bookkeeping software: workshops, building-supply shops, caterers, small restaurants, contractors. You install the folder, the script and the Sheet on their computer, on their own Gemini key so the API bill and the data stay theirs, and teach one staff member the photo routine. That’s 2–4 hours of your time.

2. Monthly upkeep. You check the review/ folder, fix mistakes, and send a monthly expense summary by category. That’s a bookkeeping-assistant job done in a fraction of the hours.

How to price it: start from the time you save the client, not from your API cost. If they currently pay staff for 10 hours of data entry a month, your monthly fee has to come in clearly below that and still be worth your 1–2 hours. Work it out with local wages. Don’t copy a dollar figure from a blog, including this one.

3. Accountants as a channel. Small accounting practices (konsultan pajak, freelance bookkeepers) get receipts from clients as messy photos every month. One accountant with 20 small clients is a better lead than 20 separate shops. Demo it on their own receipts.

Realistic expectations: getting a first paying client usually takes weeks of showing the demo in person, and it may not happen at all. “I’ll turn your receipts into a spreadsheet, here’s last month’s pile done” sells. “AI-powered OCR pipeline” doesn’t.

Reality check: what does NOT work

Skipping human review. Most clear receipts come out right. Not all of them. A clean printed Alfamart receipt is easy. A crumpled, sun-faded fuel slip or a handwritten nota with a scribbled total is not. The validation step plus a human looking at review/ is the product. Without it, you’re just making wrong numbers faster.

Putting client financial data through the free tier. Under Google’s Gemini API terms, content sent on the unpaid tier can be used to improve Google’s products and may be seen by human reviewers. The paid tier doesn’t do this. For your own coffee receipts, the free tier is fine. For a client’s supplier invoices, turn on billing. It costs under a dollar a month.

Bad photos. Most “the AI is bad” complaints are really photo problems: shadows, a curled receipt, flash glare. Lay it flat, use daylight, fill the frame. Photograph thermal receipts the same day, because thermal print fades, faster in a hot car or a sunny drawer.

Treating this as accounting. This gives you clean expense data. It doesn’t make the data tax-compliant, decide what’s deductible, or replace an accountant for tax filings. Keep the original photos (the processed/ folder) as evidence and back them up.

Leaving thinking on and not resizing. A full-size 12-megapixel photo with thinking enabled costs several times more per receipt for no accuracy gain on this task. It’s still cheap, but across several clients’ bills it adds up.

Assuming duplicates are caught. Two staff members snapping the same nota happens all the time. The SHA-256 hash only catches identical files, so a second photo of the same receipt gets through. Once a month, sort the sheet by date + merchant + total and look for twins.

Hard-coding the model name forever. Google retires model versions on a schedule (check ai.google.dev/gemini-api/docs/deprecations). When gemini-2.5-flash gets a shutdown date, set the RECEIPT_MODEL environment variable to its replacement. For Gemini 3.x models, also swap thinking_budget=0 for the lowest thinking_level that model supports (for example thinking_level="minimal" or "low"). Then run 10 test receipts and compare against old results before you switch a client over.

Final checklist

  • Script runs on 10 of your own receipts, results checked by hand
  • thinking_budget=0 and image resize in place
  • Billing enabled before any client data goes through
  • review/ folder checked weekly by a named person
  • Original photos backed up off the machine
  • Cron or Task Scheduler job running, scan.log checked after the first week
  • Monthly duplicate check on date + merchant + total

Run it on your own receipts for a month. If it saves you hours and the review pile stays small, you’ve got a tool for yourself and a demo you can show a business owner.