Every online seller has the same backlog. The photos are done and the stock is on the shelf, but 80 products are still sitting unlisted because nobody has written the titles and descriptions. Or they are listed with one line of text copied from the supplier’s WhatsApp message.

Writing a decent listing by hand takes about 10 to 20 minutes per product once you count the title, a few bullet points, a spec list and some search keywords. Time yourself on five products, because that number decides whether this project is worth it for you. At 15 minutes each, 100 products is 25 hours of writing.

Pasting products one by one into ChatGPT is faster, but it brings a new problem. LLMs fill gaps. Ask for a description of a canvas tote bag and you may get “waterproof”, “premium 16 oz canvas” or “1-year warranty” when you never said any of that. On a marketplace, an invented claim means returns and bad reviews. In Indonesia it can also break consumer protection law (UU No. 8/1999 bans misleading product information).

In this guide you’ll build a small Python tool that:

  • reads a CSV where each row holds the facts about one product
  • sends each row to Google’s Gemini API and gets back structured JSON: title, bullets, description, search keywords
  • checks the output against your facts: title length, numbers that aren’t in your data, and risky claims like “original”, “BPOM”, “garansi” or “waterproof”
  • sends the model its own mistakes once to fix, then marks anything still wrong as review
  • writes everything to a CSV you can paste into a marketplace mass-upload template

The AI cost is roughly $0.002 per product, so about $1 for 500 products. Build time is about an hour.

What you need

ItemCostNotes
Python 3.10+$0Any laptop, mini PC or VPS
Gemini API key$0 to startaistudio.google.com → “Get API key”
google-genai, pydantic$0pip install
A spreadsheet of product factsYour timeGoogle Sheets or Excel, exported as CSV
Seller Center access$0For the mass-upload template, if you sell on a marketplace

Pricing as of October 2026 (check ai.google.dev/pricing before you commit, because Google changes model names and prices often):

  • Gemini 3.5 Flash-Lite (gemini-3.5-flash-lite, paid tier): $0.30 per 1M input tokens, $2.50 per 1M output tokens. Thinking tokens are billed as output.
  • Gemini 3.1 Flash-Lite (gemini-3.1-flash-lite, paid tier): $0.25 per 1M input, $1.50 per 1M output. Google lists May 2027 as its earliest shutdown date.
  • Free tier: both models have one. It’s rate-limited and fine for testing. Read the privacy note in the Reality check before you put a client’s catalogue through it.
  • Gemini 2.5 Flash / Flash-Lite: older tutorials use these. Google now gives access to the 2.5 models only to accounts that already used them, so a new API key probably won’t work with them.

How the cost works out

One product sends about 500 tokens of instructions plus 100 to 200 tokens of facts. The reply (title, 3 to 5 bullets, a description of up to 1,500 characters and keywords) is about 500 to 700 tokens. In Indonesian it’s at the upper end, because Indonesian text splits into more tokens than English does.

Per product on Gemini 3.5 Flash-Lite:

  • Input: ~700 tokens × $0.30/1M = $0.0002
  • Output: ~700 tokens × $2.50/1M = $0.0018
  • Total: about $0.002. Minimal thinking adds a few output tokens, and if 20% of products need a retry, add about 20%.

Two ways to make it cheaper:

  1. Keep thinking at the minimum. Gemini 3.x models “think” before answering and bill those tokens as output. Writing a product description doesn’t need reasoning, so the script sets thinking_level="minimal". The 3.x models can’t switch thinking off completely, so the script adds the thinking tokens to its cost estimate.
  2. Try gemini-3.1-flash-lite (set LISTING_MODEL). Output costs 40% less. Run 10 products through both models and compare the writing before you switch.

Step 1: Set up the project

mkdir listings && cd listings
python3 -m venv .venv && source .venv/bin/activate   # Windows: .venv\Scripts\activate
pip install google-genai pydantic
export GEMINI_API_KEY="your-key-here"                # Windows: set GEMINI_API_KEY=...

Step 2: Build the facts sheet (this is the real work)

The AI can only be as accurate as this sheet. Make one row per product and one column per fact. Leave a cell empty when you don’t know. Never write “-” or “N/A”, because the model may try to “fill” them.

Save it as products.csv:

sku,product_type,brand,material,size,colors,weight,features,care,notes
TOTE-01,tote bag,Kain Kita,canvas 12 oz,40 x 35 x 10 cm,"black, cream",350 g,"inner zip pocket; magnetic snap closure",hand wash cold,handmade in Bandung
MUG-03,ceramic mug,,stoneware,350 ml,"sage green",,"microwave safe; dishwasher safe",,glaze colour varies slightly per piece
SOAP-07,bar soap,Rumah Sabun,"olive oil, coconut oil, shea butter",100 g,,,"cold process; unscented",keep dry between uses,

Tips that save you editing later:

  • Put measured numbers in the sheet: weight on a kitchen scale, size with a tape measure. These are exactly the details buyers ask about in chat.
  • Use features for things you have checked. “Microwave safe” goes here only if you actually tested it, or the maker states it.
  • Use notes for honesty details (“colour varies per piece”). They prevent returns.
  • Leave price and stock out. Marketplaces have separate fields for them, and a price written into the description goes stale.

Step 3: The script

Save this as write_listings.py:

import csv
import os
import re
import sys
import time

from google import genai
from google.genai import types
from pydantic import BaseModel

MODEL = os.getenv("LISTING_MODEL", "gemini-3.5-flash-lite")
LANG = os.getenv("LISTING_LANG", "Indonesian")
TITLE_MAX = int(os.getenv("TITLE_MAX", "100"))
DESC_MIN, DESC_MAX = 300, 1500
PRICE_IN, PRICE_OUT = 0.30, 2.50  # USD per 1M tokens, Gemini 3.5 Flash-Lite paid tier

# Claims that cause returns or legal trouble if they are not true.
# They are only allowed when the same word appears in your facts.
RISKY = ["original", "ori", "asli", "bpom", "sni", "halal", "garansi", "warranty",
         "waterproof", "anti air", "tahan air", "100%", "terbaik", "best", "no.1",
         "#1", "dijamin", "guaranteed", "menyembuhkan", "obat", "cure", "premium"]


class Listing(BaseModel):
    title: str
    bullets: list[str]
    description: str
    search_keywords: list[str]


SYSTEM = f"""You write marketplace product listings for a small online shop.
Rules:
- Use ONLY the facts provided. Never invent materials, sizes, quantities,
  certifications, warranties or claims. If a fact is missing, leave it out.
- Write in {LANG}. Plain and specific. No hype words, no emojis, no ALL CAPS.
- title: max {TITLE_MAX} characters. Product type first, then brand (if given),
  then the key attributes buyers filter by (material, size, variant).
  No symbols like ! $ ? and do not repeat a word.
- bullets: 3 to 5 short lines, each one concrete fact or use.
- description: {DESC_MIN}-{DESC_MAX} characters. Two short paragraphs, then a
  specifications list using only the given facts, then care instructions if given.
  Include the notes field if present, it is there for honesty.
- search_keywords: 5 to 10 lowercase phrases a buyer would actually type."""


def numbers(text: str) -> set[str]:
    return {n.replace(",", ".") for n in re.findall(r"\d+(?:[.,]\d+)?", text)}


def check(listing: Listing, facts: str) -> list[str]:
    problems = []
    if len(listing.title) > TITLE_MAX:
        problems.append(f"title is {len(listing.title)} characters, max is {TITLE_MAX}")
    if not DESC_MIN <= len(listing.description) <= DESC_MAX:
        problems.append(f"description is {len(listing.description)} characters, "
                        f"must be {DESC_MIN}-{DESC_MAX}")
    if not 3 <= len(listing.bullets) <= 5:
        problems.append(f"{len(listing.bullets)} bullets, need 3 to 5")

    text = " ".join([listing.title, *listing.bullets, listing.description])
    invented = numbers(text) - numbers(facts)
    if invented:
        problems.append("numbers not in the facts: " + ", ".join(sorted(invented)))

    low_text, low_facts = text.lower(), facts.lower()
    for word in RISKY:
        pattern = rf"(?<!\w){re.escape(word)}(?!\w)"
        if re.search(pattern, low_text) and not re.search(pattern, low_facts):
            problems.append(f"claim '{word}' is not in the facts")
    return problems


def generate(client, facts: str, feedback: list[str] | None = None):
    prompt = f"Product facts:\n{facts}"
    if feedback:
        prompt += "\n\nYour previous draft had these problems. Fix them:\n- " + "\n- ".join(feedback)
    for attempt in range(4):
        try:
            resp = client.models.generate_content(
                model=MODEL,
                contents=prompt,
                config=types.GenerateContentConfig(
                    system_instruction=SYSTEM,
                    response_mime_type="application/json",
                    response_schema=Listing,
                    thinking_config=types.ThinkingConfig(thinking_level="minimal"),
                ),
            )
            if resp.parsed is None:
                raise ValueError("no valid JSON in the response")
            usage = resp.usage_metadata
            out_tokens = (usage.candidates_token_count or 0) + (usage.thoughts_token_count or 0)
            cost = ((usage.prompt_token_count or 0) * PRICE_IN + out_tokens * PRICE_OUT) / 1e6
            return resp.parsed, cost
        except Exception as e:  # rate limits on the free tier, network blips, bad JSON
            wait = 15 * (attempt + 1)
            print(f"  error: {e}. Retrying in {wait}s", file=sys.stderr)
            time.sleep(wait)
    raise RuntimeError("Gemini failed 4 times in a row")


def main(src="products.csv", dst="listings.csv"):
    client = genai.Client()  # reads GEMINI_API_KEY
    done = set()
    if os.path.exists(dst):
        with open(dst, newline="", encoding="utf-8") as f:
            done = {row["sku"] for row in csv.DictReader(f)}

    fields = ["sku", "status", "problems", "title", "bullets", "description", "search_keywords"]
    new_file = not os.path.exists(dst)
    total = 0.0
    with open(src, newline="", encoding="utf-8-sig") as f_in, \
         open(dst, "a", newline="", encoding="utf-8") as f_out:
        writer = csv.DictWriter(f_out, fieldnames=fields)
        if new_file:
            writer.writeheader()
        for row in csv.DictReader(f_in):
            sku = row.get("sku", "").strip()
            if not sku or sku in done:
                continue
            facts = "\n".join(f"{k}: {v.strip()}" for k, v in row.items()
                              if k != "sku" and v and v.strip())
            listing, cost = generate(client, facts)
            problems = check(listing, facts)
            if problems:  # one repair round with the exact problems listed
                listing, cost2 = generate(client, facts, problems)
                cost += cost2
                problems = check(listing, facts)
            total += cost
            writer.writerow({
                "sku": sku,
                "status": "review" if problems else "ok",
                "problems": "; ".join(problems),
                "title": listing.title,
                "bullets": "\n".join(f"• {b}" for b in listing.bullets),
                "description": listing.description,
                "search_keywords": ", ".join(listing.search_keywords),
            })
            f_out.flush()
            print(f"{sku}: {'REVIEW' if problems else 'ok'}  (${cost:.4f})")
    print(f"Done. Estimated cost this run: ${total:.4f}")


if __name__ == "__main__":
    main(*sys.argv[1:])

What the important parts do:

  • response_schema=Listing makes Gemini return JSON that matches the Pydantic model, and resp.parsed gives you a Listing object. You don’t need to strip markdown fences or fix broken JSON.
  • check() is the part that keeps you safe. Any number in the output that isn’t in your facts gets flagged. That catches “16 oz” when you wrote “12 oz”, or “2 year warranty” when you wrote nothing. Risky words are only allowed if the same word is in your facts.
  • One repair round. The model gets back its exact problems (“title is 118 characters, max is 100”) and tries again. In practice this fixes most length problems. Anything still wrong is marked review rather than shipped.
  • Resume. SKUs already in listings.csv are skipped, so if the run stops at product 312 you just run it again.

Step 4: Run it

Start with five products, not five hundred:

head -n 6 products.csv > sample.csv
python write_listings.py sample.csv sample_out.csv

The output looks something like this:

TOTE-01: ok  ($0.0019)
MUG-03: REVIEW  ($0.0041)
SOAP-07: ok  ($0.0017)
Done. Estimated cost this run: $0.0077

Open sample_out.csv in Google Sheets and read every listing. If MUG-03 was flagged with numbers not in the facts: 2, look at it. Sometimes the model invented something (“set of 2”). Sometimes it’s a false alarm, like “2 paragraphs” or a number written as a word in your facts. The checker is deliberately strict. A false alarm costs you 30 seconds, while a missed invented claim can cost you a return.

When the samples look right, run the full sheet:

python write_listings.py products.csv listings.csv

For English listings in Amazon or Etsy style, write to a separate output file. Otherwise the resume logic sees the SKUs already in listings.csv and skips them all:

LISTING_LANG=English TITLE_MAX=200 python write_listings.py products.csv listings_en.csv

Step 5: Tune the title length for your marketplace

Title limits vary by platform and sometimes by category, and they change. Check the limit shown in the product form in your Seller Center and set TITLE_MAX to match. Amazon’s title rules, in force since January 21, 2025, cap most categories at 200 characters, and some apparel categories at less. They also ban symbols like ! $ ? _ { } ^ unless they’re part of the brand name, and allow any word (apart from articles, prepositions and conjunctions) at most twice. That’s why the prompt bans those symbols and repeated words.

A longer title is not automatically better. Marketplace search results cut titles off on mobile, so the first 40 to 60 characters do most of the work. That’s why the prompt puts the product type and key attributes first.

Step 6: Get it into the marketplace

Shopee and Tokopedia both offer mass upload in Seller Center: you download an Excel template, fill it in and upload it. The column names and required fields change from time to time, so don’t hard-code them. Copy the title and description columns from listings.csv into the template, and fill price, stock, weight and category in the template itself.

If you don’t have many products, copy-pasting from the sheet into the product form also works. Even then, you save the writing time, which was the slow part.

Cost breakdown

ItemOne-timePer 100 products
Python, libraries$0–
Gemini 3.5 Flash-Lite (with ~20% retries)–~$0.25
Building the facts sheet2–5 min per product3–8 hours
Reviewing output1–2 min per product2–3 hours
Script setup~1 hour–

The API bill is a rounding error. Your time goes into collecting facts and reviewing the output, not writing. Compared with 25 hours of hand-writing, a realistic total for 100 products is 5 to 11 hours, and most of that is measuring and checking things you should have recorded anyway.

Turning it into money

Your own shop. The direct gain is getting stock listed weeks sooner. An unlisted product earns nothing. Bilingual listings also become cheap: if you sell cross-border, run the same sheet with LISTING_LANG=English into a second output file.

Offering it as a service. Many small sellers have the same backlog and don’t want to touch Python. A “listing cleanup” package is a realistic side service:

  • You collect product facts from the seller, through a Google Form or a video call where they hold each product up.
  • You run the script, review every listing and fix the flagged ones.
  • You deliver a filled mass-upload template.

Before you set a price, check what freelancers on Fastwork, Sribu or Fiverr charge for “product description” or “deskripsi produk”. Price per SKU, and charge more when you also collect the facts and do the upload, because that’s where the hours are. Be upfront that AI drafts the text and you check it. Clients find out anyway, and honesty is a better selling point than pretending.

What you’re really selling is the facts sheet, the review and the upload, not the AI text. Anyone can paste into ChatGPT. Few people will measure 80 mugs and catch the model calling stoneware “porcelain”.

Reality check: what does NOT work

Better text won’t fix a bad listing. Marketplace search ranking weighs sales, reviews, conversion, price and photos heavily. A clean description helps buyers decide and cuts down chat questions, but it won’t push a product with poor photos and no reviews up the rankings. Don’t promise clients ranking gains.

Keyword stuffing. Asking the model for a 200-character title packed with every synonym (“Tas Tote Bag Wanita Pria Kanvas Canvas Murah Kuliah Kerja…”) makes listings look like spam. Shopee, Tokopedia and Amazon all have rules against misleading or irrelevant keywords. Keep search_keywords as a separate list you use carefully, not something you dump into the title.

Trusting the model on facts it wasn’t given. The checker catches invented numbers and a list of risky words, but it can’t catch everything. “Stoneware” turning into “porcelain” or “cotton” into “linen” has no number in it. That’s why every listing still gets a human read. For cosmetics, food and supplements, check every claim by hand. In Indonesia, wording about BPOM registration, halal status or health benefits has legal weight.

Same text across sellers. If ten resellers of the same supplier run the supplier’s spec sheet through the same prompt, they get near-identical listings. Add facts only you have: your own measurements, how it looks in real light, what customers have asked in chat.

Leaving the free tier on for client work. Under Google’s Gemini API terms, content sent on the free tier can be used to improve Google’s products and train its models, and human reviewers may read it. Paid-tier prompts and responses are not used to improve products. They are only logged for a limited time to detect abuse. (In the EEA, Switzerland and the UK, the paid-tier terms apply to the free tier too.) A product catalogue usually isn’t sensitive, but if a client’s unreleased product line goes through it, switch on billing first. At these volumes it costs cents.

Skipping the five-product test. Running 500 products before reading any output is how you end up with 500 listings that all open with the same stiff sentence. Read the samples, adjust the prompt rules (tone, structure, things to always mention), then scale up.

Model churn. Google retires and renames models regularly. In 2026 it restricted the 2.5 models to existing users. If the script suddenly errors on the model name, check the deprecations page in the docs and set LISTING_MODEL. The thinking levels differ by model: some don’t accept "minimal", and their lowest level is "low". Also, Google now recommends its newer Interactions API for new projects. The generate_content call used here is labelled legacy but is still fully supported.

Final checklist

  • Timed yourself writing 5 listings by hand, so you know what you’re saving
  • Facts sheet with measured numbers, empty cells for unknowns
  • TITLE_MAX set to your marketplace’s current limit
  • 5-product test run, every listing read
  • Prompt adjusted, then full run
  • Every review row fixed by hand, ok rows spot-checked
  • Billing enabled before processing client catalogues

The AI part of this project is cheap and quick. What makes it work is the facts sheet and the checker, because they stop the model from making things up. Do those two properly and a weekend of measuring and reviewing replaces weeks of writing.