How to scrape Amazon: Python, the law and official APIs
Scrape public Amazon product and price data with Python: what the terms and courts say, the official API routes, polite rate limits and where proxies fit.

Quick summary · TL;DR
- There are three routes to Amazon data. The official APIs if you qualify (Creators API for Associates, SP-API for sellers), your own scraper on public logged-out pages, or a bought dataset.
- Amazon's terms prohibit scraping. The Conditions of Use exclude any collection of product listings or prices and any use of robots or data mining tools, and robots.txt bars 99 named agents, Scrapy included, from the whole site.
- Courts separate access from contract. US rulings since 2021 treat logged-out collection of public pages as mostly outside computer-crime law, but contract, database right and the GDPR still apply.
- The code should stop, not evade. A 503, a robot-check page or a sign-in redirect means log it, back off and end the run after a short streak.
- Most small jobs need no proxy. A few hundred ASINs a day at a slow pace can run from one IP; residential exits fit other countries' marketplaces, and volume past one polite exit belongs on the official APIs or a bought dataset.
The direct answer to how to scrape Amazon comes in three routes. Use Amazon’s own APIs if you qualify: the Creators API for Associates, SP-API for sellers and vendors. Otherwise run your own scraper on public, logged-out product and search pages, paced slowly. Or buy a dataset or a price-history service. This guide covers the second route, done within the rules.
Most Amazon scraping tutorials open with pip install and never mention that Amazon’s terms forbid the thing they are teaching. The rest mention it in one line and move on to header rotation. Neither helps a data engineer who has to explain the job to a legal reviewer before it runs.
How to scrape Amazon: three routes
Route 1: the official APIs. If you are an Amazon Associate with recent sales, the Creators API returns product data under Amazon’s own license. If you sell or supply on Amazon, SP-API returns pricing and catalog data for the marketplaces you trade in. Both are sanctioned, metered and stable.
Route 2: your own scraper. A script that fetches public product pages (/dp/{ASIN}) and search pages (/s?k=) while logged out, slowly, and parses the HTML. Every tutorial teaches this route, this guide covers it, and Amazon’s terms prohibit it.
Route 3: bought data. Price-history services and dataset vendors already collect Amazon prices at scale. For a one-off market study, buying a dataset is often less work than building and running a scraper, and it moves the collection question to someone else’s contract.
Which route to pick. If you are an Associate with 10 qualifying sales in the last 30 days, use the Creators API. If you sell on Amazon and need buy box data for your own marketplaces, use SP-API. If you need a one-off study, price a bought dataset before building anything. If you still need your own scraper, run the script from your own IP for a week and read the status column. Add residential exits only when the job needs other countries’ marketplaces; volume past one polite exit belongs on route 1 or route 3.
Amazon’s Conditions of Use on scraping
Amazon’s Conditions of Use (last updated August 14, 2026, read September 29, 2026) grant “a limited, non-exclusive, non-transferable, non-sublicensable license to access and make personal and non-commercial use” of the site. That license excludes “any collection and use of any product listings, descriptions, or prices” and “any use of data mining, robots, or similar data gathering and extraction tools”. Amazon’s terms prohibit scraping.
Amazon’s robots.txt, fetched September 29, 2026, reads like this:
- The
User-agent: *group has 118 Disallow lines and no Crawl-delay. - Product pages (
/dp/{ASIN}) and search pages (/s) are not disallowed for that group. Sub-paths such as/dp/shipping/and/dp/rate-this-item/are. - The all-sellers pages under
/gp/offer-listing/are disallowed, with two narrow Allow exceptions. So are/gp/cart,/ap/signin,/dp/product-availability/,/ss/twister/ajaxand/gp/product/product-availability. - 99 named agents (ClaudeBot is listed twice) get
Disallow: /, the whole site. The list includes GPTBot, ClaudeBot, PerplexityBot, Google-Extended, Bytespider andScrapy.
The Scrapy line hits Python users first. Scrapy’s settings docs say scrapy startproject writes ROBOTSTXT_OBEY = True into every new project, and the default user agent is Scrapy/VERSION (+https://scrapy.org). A stock Scrapy project that respects the file will fetch nothing from amazon.com. Renaming the agent to get past that line defeats the file, so don’t.
The practical rule this guide follows sits between the two documents: public, logged-out pages only. No account data, no cart, no checkout, no sign-in, no reviews behind the login wall, no personal data. That matches proxymint’s own acceptable-use terms, which tell customers to collect public data only and not to get behind logins. A proxy acceptable use policy sets those lines for every job on the network, not only Amazon.
Is scraping Amazon legal?
At a high level, and not legal advice: “is Amazon scraping legal” is three questions with different answers.
Computer-access law. In the US, the Supreme Court’s Van Buren v. United States (June 3, 2021) narrowed the Computer Fraud and Abuse Act to a “gates-up-or-down” test: are you in a part of the system you have no right to enter? The Ninth Circuit then held in hiQ Labs v. LinkedIn (April 18, 2022) that scraping pages open to the public is likely not access “without authorization” under the CFAA.
Contract. The same hiQ case did not end well for hiQ. In November 2022 the district court found it had breached LinkedIn’s user agreement, and on December 8, 2022 the court entered a consent judgment with a permanent injunction against further scraping. Public is not the same as permitted. In Meta’s 2024 case against a data-collection vendor (N.D. Cal., January 23, 2024), the court held that Meta’s terms bound the vendor only while it was logged in, which is why logged-out collection matters so much. Amazon v. Perplexity was about an AI shopping agent acting inside shoppers’ logged-in accounts, the area this guide stays out of. A district court granted Amazon a preliminary injunction in March 2026, and the Ninth Circuit vacated it on August 4, 2026, holding that the user, not the agent’s maker, accessed Amazon’s computers.
Data law. The EU has no CFAA equivalent, but it has two other levers. The Database Directive’s sui generis right protects a maker’s investment: in CV-Online Latvia v Melons (C-762/19, June 3, 2021) the CJEU held that extraction infringes when it harms that investment. Where no database right applies, Ryanair v PR Aviation (C-30/14, January 15, 2015) says the site’s own terms can restrict use. On top of both sits the GDPR, which applies to personal data even when it is public. The Dutch data protection authority said so in its May 2024 scraping guidance, and the EDPB’s Guidelines 03/2026 on web scraping in the context of generative AI (adopted July 7, 2026) say the same, with a public consultation open until October 30, 2026.
For an Amazon price job, product titles, prices, ratings and ASINs are not personal data. Reviewer names and profile links are, so do not collect them.
Amazon’s terms forbid scraping. Courts have mostly treated logged-out collection of public pages as outside computer-crime law. Contract, database and data-protection law still apply. Talk to a lawyer before this runs commercially.
The official routes first
Before writing a scraper, check whether Amazon will give you the data.
Creators API. This is the successor to Product Advertising API 5. PA-API 5 is deprecated, and calls to it return 403 with a message to migrate (checked September 29, 2026). Any tutorial from 2020 that sends you to PA-API is now wrong. The Creators API needs enrolment in the Associates program for the target marketplace and at least 10 qualifying sales in the past 30 days. It authenticates with OAuth 2.0 client credentials from Login with Amazon, and offers four operations: SearchItems, GetItems, GetVariations and GetBrowseNodes.
SP-API. For sellers and vendors, the Selling Partner API covers Product Pricing, Catalog Items and pricing notifications. getItemOffersBatch returns the buy box price, the lowest offers and offer counts, on a published usage plan of 0.1 requests per second with a burst of 1. Amazon announced annual and usage fees for SP-API developers in late 2025, then dropped them in 2026.
When the API cannot serve the job. You have no Associates sales yet. You are not a seller. You need search rank positions as a shopper sees them, or prices in a marketplace you do not sell in. That is where the rest of this guide starts.
What a public product page shows
A logged-out product page carries more than most jobs need: title, price, currency, list price, availability text, rating, rating count, ASIN, brand, bullet points, images, breadcrumb category, best-seller rank, and the few featured reviews shown on the page. Take review text only, never reviewer names.
What you cannot get this way: the full review history, which has sat behind sign-in since late 2024 (on September 29, 2026, a logged-out request for /product-reviews/{ASIN}/ on amazon.com still redirected to /ap/signin); anything under /gp/offer-listing/; the cart; and delivery estimates tied to an account.
Search pages (/s?k=) give ASINs, titles, prices, the sponsored flag and position on the page.
A working Python example
This is how to scrape Amazon product pages in one complete script: Python 3.12, using httpx (0.26 or later, for the proxy argument) and selectolax. It reads a list of ASINs for one marketplace, checks robots.txt once per run, fetches each canonical product URL at a fixed pace with jitter, and writes one row per fetch with a timestamp, the marketplace and the HTTP status. Every selector was present on a live, logged-out amazon.com product page on September 29, 2026. Amazon runs page variants, so every field falls back to empty instead of crashing.
"""Polite Amazon price collector: public, logged-out product pages only.
Stops on 503/429, robot-check pages and sign-in redirects. Never solves them."""
import csv, os, random, sys, time
from datetime import datetime, timezone
from decimal import Decimal, InvalidOperation
from urllib.robotparser import RobotFileParser
import httpx
from selectolax.parser import HTMLParser
MARKET = "www.amazon.com" # www.amazon.de, www.amazon.co.uk, ...
LANG = "en-US,en;q=0.8" # match the marketplace
UA = os.environ["SCRAPER_UA"] # e.g. "acme-prices/1.0 (+https://acme.example/bot)"
PACE = 5.0 # seconds between requests on this exit
MAX_STRIKES = 3 # blocks in a row before the run stops
OUT = "amazon_prices.csv"
FIELDS = ["fetched_at", "market", "asin", "status", "note",
"title", "price", "rating", "reviews", "availability", "image"]
def robots_gate(client):
r = client.get(f"https://{MARKET}/robots.txt")
if r.status_code != 200:
sys.exit(f"robots.txt returned {r.status_code}: not starting")
rp = RobotFileParser()
rp.parse(r.text.splitlines())
return rp
def blocked(r):
if r.status_code in (429, 503):
return "throttled"
if "/ap/signin" in str(r.url):
return "sign-in"
if "/errors/validateCaptcha" in r.text:
return "robot-check"
return None
def pick(tree, sel, attr=None):
node = tree.css_first(sel)
if node is None:
return None
return (node.attributes.get(attr) if attr else node.text(strip=True)) or None
def to_decimal(raw):
if not raw:
return None
s = "".join(c for c in raw if c.isdigit() or c in ",.")
s = s.replace(".", "").replace(",", ".") if "," in s[-3:] else s.replace(",", "")
try:
return Decimal(s)
except InvalidOperation:
return None
def parse(tree):
price = pick(tree, "span.a-price span.a-offscreen") or pick(tree, ".a-offscreen")
return [pick(tree, "#productTitle"), to_decimal(price),
pick(tree, "#acrPopover", "title"), pick(tree, "#acrCustomerReviewText"),
pick(tree, "#availability"), pick(tree, "#landingImage", "src")]
def main(asins):
proxy = os.environ.get("PROXY_URL") # http://USERNAME:PASSWORD@HOST:PORT, both percent-encoded
headers = {"User-Agent": UA, "Accept-Language": LANG}
new_file = not os.path.exists(OUT)
with httpx.Client(headers=headers, proxy=proxy, timeout=30,
follow_redirects=True) as client, open(OUT, "a", newline="") as f:
out = csv.writer(f)
if new_file:
out.writerow(FIELDS)
rp = robots_gate(client)
strikes, wait = 0, PACE
for asin in asins:
url = f"https://{MARKET}/dp/{asin}"
now = datetime.now(timezone.utc).isoformat(timespec="seconds")
if not rp.can_fetch(UA, url):
out.writerow([now, MARKET, asin, None, "robots-disallow"] + [None] * 6)
continue
r = client.get(url)
tree = HTMLParser(r.text)
reason = blocked(r) or (None if tree.css_first("#productTitle") else "no-title")
if reason:
strikes += 1
out.writerow([now, MARKET, asin, r.status_code, reason] + [None] * 6)
if strikes >= MAX_STRIKES:
sys.exit(f"stopped after {strikes} blocks in a row ({reason}) at {asin}")
wait *= 2 # back off; the blocked ASIN is logged, not retried
else:
strikes, wait = 0, PACE
out.writerow([now, MARKET, asin, r.status_code, "ok"] + parse(tree))
time.sleep(wait + random.uniform(0, PACE))
if __name__ == "__main__":
main([line.strip() for line in open(sys.argv[1]) if line.strip()])
Run it with a text file of ASINs, one per line: python amazon_prices.py asins.txt. The output looks like this (values illustrative):
fetched_at,market,asin,status,note,title,price,rating,reviews,availability,image
2026-09-29T06:10:04+00:00,www.amazon.com,B0XXXXXXX1,200,ok,Example Kettle 1.7L,34.99,4.5 out of 5 stars,"1,204 ratings",In Stock,https://m.media-amazon.com/...
2026-09-29T06:10:13+00:00,www.amazon.com,B0XXXXXXX2,503,throttled,,,,,,
Every blocked or empty page becomes a logged row, so the CSV doubles as an audit trail and a per-ASIN price series. The currency follows the marketplace; store it next to the price if you run more than one.
The proxy is one optional line. With PROXY_URL unset, the script runs from your own address. Percent-encode the username and password before building PROXY_URL, for example with urllib.parse.quote(PASSWORD, safe=""), so an @, : or % in a password cannot break the URL. The URL uses http://; that scheme carries HTTPS pages end to end through a CONNECT tunnel, and socks5h:// works the same way (httpx 0.28 or later, with the httpx[socks] extra).
Send an honest user agent: a project name and a URL where Amazon can reach you, the format Scrapy’s docs recommend. Every ranking tutorial sends a desktop browser string to pass as a shopper. If Amazon turns an honest agent away, that is its answer, and the job moves to route 1 or route 3. The script never rotates the header.
For search pages, the same client, gate and pacing apply. A minimal variant:
from urllib.parse import quote_plus
def search(client, rp, query):
url = f"https://{MARKET}/s?k={quote_plus(query)}"
if not rp.can_fetch(UA, url):
return []
r = client.get(url)
if blocked(r):
return []
cards = [c for c in HTMLParser(r.text).css("div[data-asin]") if c.attributes.get("data-asin")]
return [(pos, c.attributes["data-asin"], pick(c, "h2 span"),
to_decimal(pick(c, ".a-price .a-offscreen")), "Sponsored" in c.text())
for pos, c in enumerate(cards, 1)]
For page two onwards, add &page=2 and so on to the same URL, one page per paced request, and stop at the page depth the job needs. A rank tracker for 20 keywords that reads the first two pages makes 40 requests a run.
urllib.robotparser handles the basics but not every rule in RFC 9309, such as * wildcards and $ anchors inside paths. Amazon’s file uses both (Disallow: /slp/*/b$). For a job that runs every day, a stricter gate is worth the extra twenty lines.
Pacing and politeness, in numbers
Amazon publishes no crawl rate for its storefront, and robots.txt carries no Crawl-delay. So the scraper has to set its own pace and be able to say why.
- Start slow. One request every few seconds per exit, with jitter.
- One request at a time per exit. No parallel connections on the same address.
- Cap the day. Fetch only what the job needs. A daily price check on 300 ASINs is 300 requests, not 3,000.
- Cache. A page fetched this morning does not need fetching again at noon unless the price is the point.
- Go off-peak. Run when the marketplace’s own time zone is asleep.
- Treat 503 as “slow down”. The script doubles its wait after each block and quits after three.
Amazon’s sanctioned pricing API, getItemOffersBatch, runs on a usage plan of 0.1 requests per second, one call every ten seconds. It is an API quota, not a website rule, but it shows how Amazon meters licensed access.
A 2016 write-up by Hartley Brody describes crawling Amazon with 200 threads through 500 proxies. That is the shape of job this guide advises against.
Where proxies fit, and when not
A few hundred ASINs a day at a slow pace can run from one IP with no proxy at all.
Proxies fit one case: the job needs other countries’ marketplaces as a local shopper sees them. Volume past what one polite exit handles belongs on route 1 or route 3, not on more exits.
Residential. The default fit for marketplace price and stock monitoring. Rotating residential proxies are billed per GB, with rotation set on the order: a new IP every request, a timer from 5 to 60 minutes, or sticky while the device stays online. How residential sits next to the other tiers for scraping is covered in best proxies for web scraping.
Datacenter. Datacenter proxies are a listed misfit for marketplaces with anti-bot scoring, and this guide does not test around that. Why is covered in datacenter vs residential proxies.
Static ISP and mobile. Both are built for other jobs. A static ISP order is a fixed set of addresses, so a broad ASIN list puts all its load on the same few IPs. Mobile is built for carrier-origin targets, and price monitoring does not check for one.
What Amazon sees from the exit
A request reaches Amazon with three things the scraper controls and one it does not. It controls the marketplace domain, the Accept-Language header and the pace. It does not control how Amazon classifies the exit’s address: the network that owns it (a home ISP or a hosting company), its country, and whatever that address did before. A proxy changes the address class and the country. It does not change what the script sends, so a scraper that is too fast from home is too fast through a residential exit as well.
A sticky session keeps one address for a worker for 5 to 60 minutes, or while the device stays online, so the cookies Amazon sets and the address they came from stay together for the whole run. Per-request rotation gives every page a new address with the same cookie jar, which is harder to reason about when you read the status column afterwards.
Scrape Amazon for free first
You can scrape Amazon for free at small scale: Python, the script above and your own connection. A proxy only enters when the job needs other countries. Run the free version for a week first, and look at the status column.
Price depends on request origin
Price, availability, currency and the delivery promise all change with the marketplace domain and the delivery location. Use the domain for the country (amazon.de for Germany, amazon.co.uk for the UK) and an exit in the same country, so the page matches what a local shopper sees. Residential location is set on the order by country, region and city, from places with devices online at the time. It does not go down to postal code, and Amazon’s delivery-location setting is a site feature that no proxy sets.
What it costs
GB needed = pages x MB per page / 1,024, rounded up to a whole GB. Count failed and retried requests too: they move bytes like any other.
This guide does not print a median MB per Amazon page. There is no reliable public figure, and page weight varies by variant, so a guessed page size would make every total below it wrong. Measure yours: len(r.content) in the script gives the decoded HTML size, while a per-GB proxy meters the bytes inside the HTTPS tunnel: the compressed response plus TLS overhead, in both directions (httpx asks for gzip by default). The biggest cost lever is the fetch mode. An HTML-only fetch like the script above skips the images and scripts a full browser render downloads.
Once you have your own MB per page, size the package from it; residential is usage-based per GB in one-off packages, and the rate falls as the package grows, with volume bands on the pricing page.
Troubleshooting inside the rules
- Empty fields. Amazon served a page variant. Update the selector for that field; the rest of the row still saves.
- Wrong currency. The marketplace domain, the exit country or
Accept-Languagedo not match. Align all three. - 503 streaks. The pace is too fast. Lengthen
PACE, cut the daily cap, and check nothing else is sharing the exit. - Robot-check page. Stop, reduce the rate, try again another day. Do not solve it. Why proxies trigger captchas explains what the page is reacting to.
- Sign-in redirect. That data is behind a login and out of scope. Drop the URL from the list.
- Prices differ from your browser. Your browser has a delivery location and maybe a signed-in account. The script has neither. Compare with a logged-out private window set to the same country.
Frequently asked questions
Amazon's Conditions of Use prohibit robots, data mining and the collection of product listings and prices, so scraping breaks Amazon's terms. US courts have mostly treated logged-out collection of public pages as outside the Computer Fraud and Abuse Act, but contract claims, EU database right and the GDPR still apply. This is not legal advice, and a commercial project should be reviewed by a lawyer before it runs.
Yes. The Creators API replaced Product Advertising API 5, which is deprecated and returns 403, and it needs an active Associates account with at least 10 qualifying sales in the past 30 days. Sellers and vendors can use the Selling Partner API, whose Product Pricing operations return buy box prices, lowest offers and offer counts. Neither API gives search rank positions as shoppers see them or prices in markets a seller does not trade in.
Web scraping is not illegal as such, but three kinds of law can apply to it: computer-access law, the site's contract terms, and data law such as copyright, database right and data protection. Collecting public pages without logging in is the lowest-risk pattern, while getting behind a login or collecting personal data raises the risk sharply. A lawyer should review any commercial use.
Fetch the public product page at the canonical /dp/ URL for each ASIN while logged out, and read the price from the span.a-price span.a-offscreen element, with a fallback selector because Amazon runs page variants. Normalise the price to a decimal with its currency and store it with a timestamp so a price series can be rebuilt. Use the marketplace domain and an exit in the same country, because prices change with location.
The same laws apply to AI crawlers as to any scraper, and site owners are drawing their own lines: Amazon's robots.txt bars 99 named agents from the whole site, including GPTBot, ClaudeBot and PerplexityBot. In the EU, the EDPB's 2026 draft guidelines say the GDPR applies whenever scraping collects personal data. Amazon's case against Perplexity concerned an AI agent acting inside shoppers' logged-in accounts, a different situation from reading public pages.
Only the few featured reviews shown on a public product page are available without signing in; the full review history has been behind sign-in since late 2024. Collecting reviews past the login breaks the logged-out rule that keeps a scraper on the safer side of the law. Reviewer names and profile links are personal data, so keep the review text and drop the names.
Pace the scraper slowly, send one request at a time per exit, cache pages already fetched and cap the daily volume to what the job needs. Treat a 503, a robot-check page or a sign-in redirect as a signal to slow down, and stop the run after a short streak instead of retrying. No method guarantees zero blocks, and a policy-first scraper never works around one.