Back to Blog

How to Scrape Amazon Reviews From Public Product Pages

Daniel Zhao

Sep 20, 2026 · Guides · 21 min read

TL;DR

Use Python, Playwright, and Beautiful Soup to extract reviews present on an Amazon.com product page you are authorized to collect from. Validate the page state, parse and deduplicate records, and compare the reopened CSV and JSON with the original records. Stop on redirects, sign-in, challenges, denied access, or rate limits. This example cannot retrieve a product’s complete review history, and a proxy does not grant access rights.

You can scrape Amazon reviews that are present in the HTML of a product page you are authorized to collect from, using Python, Playwright, and Beautiful Soup. A signed-out session may expose only a subset of reviews, require sign-in, or return an access challenge. Public visibility alone does not establish permission to collect or republish the content.

This guide is for Python developers who can use a terminal and inspect HTML. It covers one Amazon.com product page per run, CSV and JSON export, and explicit failure conditions. It does not inject account cookies, solve CAPTCHAs, or promise complete review history.

What Has Been Validated?

Version 1.1.0 completed one live, signed-out Amazon.com run on September 18, 2026. The target was https://www.amazon.com/dp/B07FZ8S74R, an Echo Dot product page. It used a direct connection, received HTTP 200, found eight review containers, exported eight records, and compared every field after reopening both files. All eight dates were normalized; two records had no displayed helpful-vote count and retained null.

Validation Scope and result
Live Amazon.com run, version 1.1.0 One ASIN, one signed-out context, no proxy, eight reviews exported and verified
Offline regression suite Nine tests passed; includes real Chrome with synthetic local responses
Challenges, sign-in, and HTTP errors Tested with synthetic cases; not encountered in this single live run
Page-2 check Implemented but not exercised in the current live run
Other marketplaces, live proxy authentication, and repeated-run reliability Not verified

The automated tests cover URL restrictions, navigation guards, page-state classification, parsing and deduplication, CSV/JSON integrity, nullable fields, HTTP errors, bounded scrolling, and the browser workflow. Synthetic fixtures validate those behaviors; the separate live run establishes only the observed result for this product and context.

Echo Dot product page captured during the September 18 version 1.1.0 signed-out run

The screenshots show product inspection, review fields, export checks, and access errors across separate test sessions. The validation table above describes the version 1.1.0 run; screenshots of other sessions illustrate the workflow rather than additional results from that run.

The live run and regression tests were recorded separately. The table above summarizes their scope; the full records are retained internally for editorial verification. No downloadable test files are required to follow this tutorial.

HTTP 200 and a recognizable ASIN do not prove that review records are available. A successful run must pass both the page-state checks and the export validation described below. One successful run does not establish a success rate, future access, or complete review coverage.

Fields Extracted by the Amazon Review Scraper

The script below stores the following fields:

Field Behavior
asin Required; extracted from a recognized Amazon product URL
review_id Used for deduplication when present
rating Parsed as a number; otherwise null
title Preserved as displayed; otherwise null
body Required for an exported review record
date_text Original Amazon date text
review_date ISO date for recognized en-US dates; otherwise null
verified_purchase true when the badge is present; otherwise null
helpful_votes Parsed count when displayed; otherwise null
source_url Final product-page URL after navigation

The exported review records omit reviewer display names because they are unnecessary for this example. Raw HTML and screenshots can still contain identities and other personal information; keep those diagnostic files private and redact copies before sharing. Missing helpful-vote values remain null, because “not displayed” is not the same as zero.

Prerequisites and Tested Versions

You need permission for the intended collection and use, Python 3.12, a supported browser, and a text editor. Install the packages below, then copy the complete script in Step 4 into a local file named amazon_review_scraper.py. This tutorial does not require any downloadable attachments.

Component Current validation environment
OS Windows build 19045
Python 3.12.10
Playwright for Python 1.63.0
Beautiful Soup 4.15.0
lxml 6.1.3
Browser Google Chrome 153.0.8010.47, headless
Script 1.1.0
Proxy authentication Not tested with a live proxy account
macOS/Linux Installation commands provided; not tested in this revision

The pinned dependencies were installed successfully for this revision. The browser context used en-US, the America/New_York timezone, and a 1440 by 900 viewport. It used one page at a time, no account cookies, and no automatic retries. Navigation is bounded by a 90-second page timeout and a 30-second intercepted document-fetch timeout; review waiting is bounded by the loop shown below. A browser locale or delivery-location label does not establish the network’s exit region, which was not independently measured.

Create an isolated environment and install the pinned packages.

Windows PowerShell

python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install beautifulsoup4==4.15.0 lxml==6.1.3 playwright==1.63.0
python -m playwright install chromium

macOS or Linux

python3 -m venv .venv
source .venv/bin/activate
python -m pip install beautifulsoup4==4.15.0 lxml==6.1.3 playwright==1.63.0
python -m playwright install chromium

On Windows, the script uses Chrome from C:\Program Files\Google\Chrome\Application\chrome.exe when present. Otherwise it uses Playwright Chromium. The Chromium installation step is necessary when that fallback is used.

Step 1: Validate the Amazon URL and ASIN

This example only supports HTTPS URLs on amazon.com and www.amazon.com. The script rejects embedded credentials, custom ports, and lookalike domains such as amazon.com.example.org. It requires a ten-character ASIN in a supported path and normalizes the input to /dp/{ASIN}.

The full validator is included in the complete script in Step 4. Browser document requests are checked against the same allowlist. Document HTTP redirects are stopped rather than followed, including redirects within Amazon.com; an otherwise legitimate redirected product URL may therefore fail in this example. Use the intended, authorized canonical product URL instead of automatically following a redirect chain. This is a narrow tutorial guard, not a general-purpose server-side URL security boundary.

Amazon product page with the signed-out state and target product marked

Step 2: Inspect the Available Review Markup

Open the product page, scroll to the review section, right-click one review, and choose Inspect. Start with the outer review container, then query fields inside that container. Scoping selectors to one review prevents a title from one record being combined with the date from another.

One review card from the current live run with identity, title, and body text redacted

The current parser uses semantic data-hook attributes and keeps a small fallback list:

Data Primary selector Fallback
Review container div[data-hook="review"] div[id^="customer_review-"]
Rating [data-hook="review-star-rating"] [data-hook="review-star-rating-view-point"]
Title [data-hook="reviewTitle"] [data-hook="review-title"]
Body [data-hook="reviewText"] [data-hook="reviewRichContentContainer"], then [data-hook="review-body"]
Date [data-hook="review-date"] Return null
Verified purchase [data-hook="avp-badge"] Return null
Helpful votes [data-hook="helpful-vote-statement"] Return null

Avoid generated class names such as _Y3Itd_...; they are much more likely to change than semantic hooks.

In DevTools, expand one div[data-hook="review"] node and verify each field against that same card. In the saved HTML from the current run, the outer container was a div; reviewTitle appeared on an h5, reviewText on a div, review-date on a span, and review-star-rating on an i. These observations support the selectors for this captured layout, not every Amazon page. If no matching container exists, do not treat unrelated text as a review.

Amazon review card with its container boundary marked and reviewer names covered

Chrome DevTools beside a saved Amazon product page and its review section

Step 3: Check Access, Then Wait for Review Containers

The page may lazy-load content, so parsing immediately after DOMContentLoaded is unreliable. The script first checks HTTP status and the initial page state, waits for the product title where available, then scrolls in finite viewport-sized steps. It stops scrolling as soon as a review container appears.

The implementation checks both supported review-container selectors. It scrolls at most 60 times, waits 250 milliseconds between scrolls, and stops early when it reaches the bottom or finds a container. It then waits up to 15 seconds for a container to attach; a timeout returns False. The full function appears in Step 4.

The loop is bounded. If no review appears, the script classifies the page and writes diagnostic HTML and a screenshot instead of retrying indefinitely.

Step 4: Parse and Deduplicate the Reviews

Beautiful Soup parses the final rendered HTML. Each record is deduplicated by review_id; if no ID exists, the parser uses a SHA-256 hash of the title, body, and displayed date.

Containers with missing bodies and duplicate record keys are skipped. Optional data stays nullable. The code retains both date_text and review_date because the original wording is useful for auditing, while an ISO date is easier to analyze. Date normalization is limited to English month-name dates in the shown en-US format; run under an English-compatible Python time locale.

Complete Python script

Copy the entire code block below into a plain-text file named amazon_review_scraper.py in your working folder. Save it as UTF-8. It includes all imports, configuration, parsing, validation, and error handling; no separate code attachment is needed. Keep the terminal in that folder when running the commands in Step 5.

from __future__ import annotations

import argparse
import csv
import hashlib
import json
import os
import platform
import re
import sys
from dataclasses import asdict, dataclass
from datetime import datetime, timezone
from importlib.metadata import version
from pathlib import Path
from typing import Any
from urllib.parse import urlparse

from bs4 import BeautifulSoup
from playwright.sync_api import (
    Browser,
    Error as PlaywrightError,
    Page,
    TimeoutError as PlaywrightTimeoutError,
    sync_playwright,
)


SCRIPT_VERSION = "1.1.0"
ASIN_RE = re.compile(
    r"/(?:dp|gp/product|product-reviews)/([A-Z0-9]{10})(?:[/?]|$)", re.I
)
ALLOWED_HOSTS = {"amazon.com", "www.amazon.com"}
REVIEW_SELECTORS = {
    "container": ['div[data-hook="review"]', 'div[id^="customer_review-"]'],
    "rating": [
        '[data-hook="review-star-rating"]',
        '[data-hook="review-star-rating-view-point"]',
    ],
    "title": ['[data-hook="reviewTitle"]', '[data-hook="review-title"]'],
    "body": [
        '[data-hook="reviewText"]',
        '[data-hook="reviewRichContentContainer"]',
        '[data-hook="review-body"]',
    ],
    "date": ['[data-hook="review-date"]'],
    "verified": ['[data-hook="avp-badge"]'],
    "helpful": ['[data-hook="helpful-vote-statement"]'],
}


@dataclass
class Review:
    asin: str
    review_id: str | None
    rating: float | None
    title: str | None
    body: str
    date_text: str | None
    review_date: str | None
    verified_purchase: bool | None
    helpful_votes: int | None
    source_url: str


def parse_args() -> argparse.Namespace:
    parser = argparse.ArgumentParser(
        description="Export reviews visible on a public Amazon product page."
    )
    parser.add_argument(
        "url",
        nargs="?",
        default="https://www.amazon.com/dp/B07FZ8S74R",
        help="Public Amazon product URL containing a 10-character ASIN.",
    )
    parser.add_argument(
        "--output-dir",
        type=Path,
        default=Path(__file__).resolve().parent / "output",
    )
    parser.add_argument(
        "--check-pagination",
        action="store_true",
        help="Check one signed-out page-2 URL and report the observed access state.",
    )
    parser.add_argument(
        "--capture-screenshots",
        action="store_true",
        help="Save unedited browser screenshots for verification.",
    )
    parser.add_argument("--headed", action="store_true")
    return parser.parse_args()


def validate_destination(url: str) -> None:
    parsed = urlparse(url)
    if (parsed.scheme != "https" or parsed.hostname not in ALLOWED_HOSTS
            or parsed.username is not None or parsed.password is not None
            or parsed.port not in {None, 443}):
        raise ValueError("Use an HTTPS amazon.com or www.amazon.com URL without credentials or a custom port.")


def validate_target(url: str) -> tuple[str, str]:
    validate_destination(url)
    parsed = urlparse(url)
    match = ASIN_RE.search(parsed.path)
    if not match:
        raise ValueError("No 10-character ASIN was found in the URL path.")
    asin = match.group(1).upper()
    return asin, f"https://{parsed.hostname}/dp/{asin}"


def guard_navigation(route: Any) -> None:
    """Allow approved documents only; stop HTTP redirects; allow CDN resources."""
    if route.request.is_navigation_request():
        try:
            validate_destination(route.request.url)
        except ValueError:
            route.abort("blockedbyclient")
            return
    # Never follow a document redirect. Playwright routing must not be assumed
    # to intercept every hop in a redirect chain.
    if route.request.is_navigation_request():
        response = route.fetch(max_redirects=0, timeout=30_000)
        if 300 <= response.status < 400 and response.headers.get("location"):
            route.fulfill(status=403, content_type="text/html", body="<title>Redirect stopped</title><p>This tutorial does not follow document redirects.</p>")
        else:
            route.fulfill(response=response)
    else:
        route.continue_()


def clean_text(value: str | None) -> str | None:
    if value is None:
        return None
    cleaned = " ".join(value.split())
    return cleaned or None


def first_text(node: Any, selectors: list[str]) -> str | None:
    for selector in selectors:
        match = node.select_one(selector)
        if match:
            value = clean_text(match.get_text(" ", strip=True))
            if value:
                return value
    return None


def parse_rating(text: str | None) -> float | None:
    match = re.search(r"([0-5](?:[.,]\d)?)", text or "")
    return float(match.group(1).replace(",", ".")) if match else None


def parse_helpful_votes(text: str | None) -> int | None:
    if not text:
        return None
    if re.search(r"\bone\s+person\b", text, re.I):
        return 1
    match = re.search(r"([\d,]+)", text)
    return int(match.group(1).replace(",", "")) if match else None


def parse_review_date(text: str | None) -> str | None:
    """Normalize the en-US date while preserving Amazon's display text."""
    match = re.search(r"\bon\s+([A-Za-z]+\s+\d{1,2},\s+\d{4})$", text or "")
    if not match:
        return None
    try:
        return datetime.strptime(match.group(1), "%B %d, %Y").date().isoformat()
    except ValueError:
        return None


def review_id_from(container: Any) -> str | None:
    html_id = clean_text(container.get("id"))
    if html_id and html_id.startswith("customer_review-"):
        return html_id.removeprefix("customer_review-")
    review_link = container.select_one('a[href*="customer-reviews/srp/-/"]')
    if review_link:
        match = re.search(
            r"/customer-reviews/srp/-/([A-Z0-9]+)", review_link.get("href", "")
        )
        if match:
            return match.group(1)
    return None


def parse_reviews(html: str, asin: str, source_url: str) -> tuple[list[Review], int]:
    soup = BeautifulSoup(html, "lxml")
    containers = []
    for selector in REVIEW_SELECTORS["container"]:
        containers = soup.select(selector)
        if containers:
            break

    reviews: list[Review] = []
    seen: set[str] = set()
    for container in containers:
        body = first_text(container, REVIEW_SELECTORS["body"])
        if not body:
            continue

        review_id = review_id_from(container)
        title = first_text(container, REVIEW_SELECTORS["title"])
        date_text = first_text(container, REVIEW_SELECTORS["date"])
        dedupe_key = review_id or hashlib.sha256(
            f"{title}|{body}|{date_text}".encode("utf-8")
        ).hexdigest()
        if dedupe_key in seen:
            continue
        seen.add(dedupe_key)

        reviews.append(
            Review(
                asin=asin,
                review_id=review_id,
                rating=parse_rating(first_text(container, REVIEW_SELECTORS["rating"])),
                title=title,
                body=body,
                date_text=date_text,
                review_date=parse_review_date(date_text),
                verified_purchase=(
                    True
                    if any(container.select_one(s) for s in REVIEW_SELECTORS["verified"])
                    else None
                ),
                helpful_votes=parse_helpful_votes(
                    first_text(container, REVIEW_SELECTORS["helpful"])
                ),
                source_url=source_url,
            )
        )
    return reviews, len(containers)


def classify_page(page: Page, html: str) -> str:
    validate_destination(page.url)
    soup = BeautifulSoup(html, "lxml")
    for node in soup.select("script, style, template, noscript"):
        node.decompose()
    # Inspect page text and challenge controls, not arbitrary script keywords.
    sample = f"{page.title()}\n{soup.get_text(' ', strip=True)}".lower()
    if re.search(
        r"click the button below.{0,200}continue shopping", sample, re.DOTALL
    ):
        return "continue_shopping_challenge"
    if "robot check" in sample or "enter the characters you see below" in sample:
        return "robot_check"
    if soup.select_one('form[action*="validateCaptcha"], form[action*="validatecaptcha"], input#captchacharacters'):
        return "captcha"
    if "/ap/signin" in page.url or (
        "sign in or create account" in sample and soup.select_one('input[name="email"]')
    ):
        return "sign_in"
    if "page not found" in sample or "dogs of amazon" in sample:
        return "not_found"
    if re.search(r"sign in.{0,200}to see customer reviews", sample, re.DOTALL):
        return "review_sign_in_required"
    if re.search(r"/dp/[A-Z0-9]{10}", page.url, re.I):
        return "product_page"
    if re.search(r"/product-reviews/[A-Z0-9]{10}", page.url, re.I):
        return "review_page"
    return "unexpected_page"


def check_http_status(response: Any, report: dict[str, Any]) -> None:
    status = response.status if response else None
    report["http_status"] = status
    if status == 429:
        report["page_type"] = "rate_limited"
        report["retry_after"] = response.headers.get("retry-after")
        raise RuntimeError("HTTP 429: stopped without retrying. Respect Retry-After before any later authorized run.")
    if status is None or status >= 400:
        report["page_type"] = {403: "access_denied", 407: "proxy_authentication_required", 404: "not_found"}.get(status, "http_error")
        raise RuntimeError(f"HTTP {status}: stopped before review extraction.")


def launch_browser(playwright: Any, headed: bool) -> Browser:
    options: dict[str, Any] = {"headless": not headed}
    chrome = Path(r"C:\Program Files\Google\Chrome\Application\chrome.exe")
    if chrome.exists():
        options["executable_path"] = str(chrome)

    proxy_server = os.getenv("ROLA_PROXY_SERVER")
    if proxy_server:
        options["proxy"] = {
            "server": proxy_server,
            "username": os.getenv("ROLA_PROXY_USERNAME", ""),
            "password": os.getenv("ROLA_PROXY_PASSWORD", ""),
        }
    return playwright.chromium.launch(**options)


def scroll_until_reviews(page: Page, max_steps: int = 60) -> bool:
    review_locator = page.locator(', '.join(REVIEW_SELECTORS["container"]))
    for _ in range(max_steps):
        if review_locator.count() > 0:
            return True
        state = page.evaluate(
            """
            () => {
              const step = Math.max(600, Math.floor(window.innerHeight * 0.75));
              window.scrollBy(0, step);
              return {
                y: window.scrollY,
                max: Math.max(0, document.documentElement.scrollHeight - window.innerHeight)
              };
            }
            """
        )
        page.wait_for_timeout(250)
        if state["y"] >= state["max"]:
            break

    try:
        review_locator.first.wait_for(state="attached", timeout=15_000)
        return True
    except PlaywrightTimeoutError:
        return False


def save_browser_evidence(page: Page, output_dir: Path, name: str) -> None:
    evidence_dir = output_dir / "browser_evidence"
    evidence_dir.mkdir(parents=True, exist_ok=True)
    (evidence_dir / f"{name}.html").write_text(page.content(), encoding="utf-8")
    page.screenshot(path=str(evidence_dir / f"{name}.png"), full_page=False)


def record_key(row: dict[str, Any]) -> str:
    return row["review_id"] or hashlib.sha256(
        f"{row['title']}|{row['body']}|{row['date_text']}".encode("utf-8")
    ).hexdigest()


def csv_record(row: dict[str, Any]) -> dict[str, str]:
    return {key: "" if value is None else str(value) for key, value in row.items()}


def verify_exports(rows: list[dict[str, Any]], output_dir: Path) -> dict[str, Any]:
    with (output_dir / "amazon_reviews.csv").open("r", newline="", encoding="utf-8-sig") as handle:
        reader = csv.DictReader(handle)
        if reader.fieldnames != list(Review.__dataclass_fields__):
            raise RuntimeError("Round-trip validation failed: CSV schema differs.")
        csv_rows = list(reader)
    json_rows = json.loads((output_dir / "amazon_reviews.json").read_text(encoding="utf-8"))
    if not rows or len(csv_rows) != len(rows) or len(json_rows) != len(rows):
        raise RuntimeError("Round-trip validation failed: empty input or row counts differ.")
    expected_csv = [csv_record(row) for row in rows]
    # JSON comparison preserves types (True and 1 must not compare as equivalent).
    if json.dumps(json_rows, sort_keys=True) != json.dumps(rows, sort_keys=True) or csv_rows != expected_csv:
        raise RuntimeError("Round-trip validation failed: exported fields differ.")
    for label, loaded in (("JSON", json_rows), ("CSV", csv_rows)):
        keys = [record_key(row) for row in loaded]
        if len(keys) != len(set(keys)) or any(not row["body"].strip() for row in loaded):
            raise RuntimeError(f"Round-trip validation failed: {label} duplicate keys or empty body.")
    return {
        "csv_rows_reopened": len(csv_rows),
        "json_rows_reopened": len(json_rows),
        "unique_record_keys": len({record_key(row) for row in json_rows}),
        "all_fields_compared": True,
        "round_trip_validated": True,
    }


def export_and_verify(reviews: list[Review], output_dir: Path) -> dict[str, Any]:
    rows = [asdict(review) for review in reviews]
    json_path = output_dir / "amazon_reviews.json"
    csv_path = output_dir / "amazon_reviews.csv"
    json_path.write_text(json.dumps(rows, indent=2, ensure_ascii=False), encoding="utf-8")

    fieldnames = list(Review.__dataclass_fields__)
    with csv_path.open("w", newline="", encoding="utf-8-sig") as handle:
        writer = csv.DictWriter(handle, fieldnames=fieldnames)
        writer.writeheader()
        writer.writerows(rows)

    return verify_exports(rows, output_dir)


def check_public_page_two(
    page: Page,
    asin: str,
    origin: str,
    report: dict[str, Any],
    output_dir: Path,
) -> None:
    url = f"{origin}/product-reviews/{asin}/?pageNumber=2"
    response = None
    state_report: dict[str, Any] = {}
    try:
        response = page.goto(url, wait_until="commit", timeout=30_000)
        validate_destination(page.url)
        check_http_status(response, state_report)
        page.locator("body").wait_for(state="attached", timeout=15_000)
        page.wait_for_timeout(3_000)
        html = page.content()
        page_type = classify_page(page, html)
        reviews, containers = (parse_reviews(html, asin, page.url)
                               if page_type in {"product_page", "review_page"} else ([], 0))
        report.update(
            {
                "public_pagination_tested": True,
                "pagination_http_status": response.status if response else None,
                "pagination_final_url": page.url,
                "pagination_page_type": page_type,
                "pagination_review_containers": containers,
                "pagination_reviews_parsed": len(reviews),
                "public_pagination_available": bool(reviews),
            }
        )
        save_browser_evidence(page, output_dir, "pagination_page_2")
    except (PlaywrightError, RuntimeError, ValueError) as exc:
        report.update(
            {
                "public_pagination_tested": True,
                "public_pagination_available": False,
                "pagination_error": type(exc).__name__,
                "pagination_page_type": state_report.get("page_type", "navigation_error"),
                "pagination_http_status": state_report.get("http_status"),
                "pagination_retry_after": state_report.get("retry_after"),
            }
        )
        try:
            save_browser_evidence(page, output_dir, "pagination_timeout")
        except PlaywrightError:
            pass


def run() -> int:
    args = parse_args()
    asin, normalized_url = validate_target(args.url)
    # A separate run directory prevents old successful exports being mistaken
    # for the output of a later failed run.
    args.output_dir = args.output_dir / datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%S%fZ")
    args.output_dir.mkdir(parents=True, exist_ok=False)
    parsed_target = urlparse(normalized_url)
    origin = f"{parsed_target.scheme}://{parsed_target.hostname}"
    report: dict[str, Any] = {
        "script_version": SCRIPT_VERSION,
        "tested_at_utc": datetime.now(timezone.utc).isoformat(timespec="seconds"),
        "python_version": sys.version.split()[0],
        "os": platform.platform(),
        "dependencies": {name: version(name) for name in ("playwright", "beautifulsoup4", "lxml")},
        "script_sha256": hashlib.sha256(Path(__file__).read_bytes()).hexdigest(),
        "automatic_retries": 0,
        "success": False,
        "target_url": normalized_url,
        "asin": asin,
        "access_state": "signed_out_fresh_context",
        "proxy_configured": bool(os.getenv("ROLA_PROXY_SERVER")),
    }

    print(f"Script version: {SCRIPT_VERSION}")
    print(f"Target ASIN: {asin}")
    print(f"Opening: {normalized_url}")

    with sync_playwright() as playwright:
        browser = None
        page = None
        response = None
        try:
            browser = launch_browser(playwright, args.headed)
            report["browser_version"] = browser.version
            context = browser.new_context(
                locale="en-US", timezone_id="America/New_York",
                viewport={"width": 1440, "height": 900}, service_workers="block",
            )
            context.route("**/*", guard_navigation)
            page = context.new_page()
            response = page.goto(
                normalized_url, wait_until="domcontentloaded", timeout=90_000
            )
            validate_destination(page.url)
            check_http_status(response, report)
            page.locator("body").wait_for(state="attached", timeout=15_000)
            page_type = classify_page(page, page.content())
            report["page_type"] = page_type
            if page_type != "product_page":
                raise RuntimeError(f"Stopped before scrolling: page_type={page_type}.")
            try:
                page.locator("#productTitle:visible").first.wait_for(
                    state="visible", timeout=15_000
                )
            except PlaywrightTimeoutError:
                pass

            if args.capture_screenshots:
                page.screenshot(
                    path=str(args.output_dir / "product_page_viewport.png"),
                    full_page=False,
                )

            review_loaded = scroll_until_reviews(page)
            if review_loaded:
                first_review = page.locator(', '.join(REVIEW_SELECTORS["container"])).first
                first_review.scroll_into_view_if_needed()
                page.wait_for_timeout(1_000)
                if args.capture_screenshots:
                    page.screenshot(
                        path=str(args.output_dir / "review_page_viewport.png"),
                        full_page=False,
                    )
                    review_area = page.locator("#customer-reviews_feature_div")
                    if review_area.count() and review_area.first.is_visible():
                        review_area.first.screenshot(
                            path=str(args.output_dir / "review_section.png")
                        )

            html = page.content()
            page_type = classify_page(page, html)
            reviews, container_count = (parse_reviews(html, asin, page.url)
                                        if page_type == "product_page" else ([], 0))
            report.update(
                {
                    "http_status": response.status if response else None,
                    "page_title": page.title(),
                    "final_url": page.url,
                    "page_type": page_type,
                    "review_locator_loaded": review_loaded,
                    "review_containers_found": container_count,
                    "reviews_parsed": len(reviews),
                    "containers_not_exported": container_count - len(reviews),
                    "public_pagination_tested": False,
                    "public_pagination_available": None,
                }
            )

            if page_type != "product_page" or not reviews:
                save_browser_evidence(page, args.output_dir, "failed_product_page")
                raise RuntimeError(
                    f"No usable public reviews were found; page_type={page_type}. "
                    "Inspect browser_evidence before retrying."
                )

            report["export_validation"] = export_and_verify(reviews, args.output_dir)
            snapshot = args.output_dir / "source.html"
            snapshot.write_text(html, encoding="utf-8")
            report["source_html_sha256"] = hashlib.sha256(snapshot.read_bytes()).hexdigest()
            report["export_sha256"] = {name: hashlib.sha256((args.output_dir / name).read_bytes()).hexdigest()
                                       for name in ("amazon_reviews.csv", "amazon_reviews.json")}
            report["success"] = True
            report["field_quality"] = {
                "missing_review_id": sum(r.review_id is None for r in reviews),
                "missing_rating": sum(r.rating is None for r in reviews),
                "missing_title": sum(r.title is None for r in reviews),
                "missing_body": sum(not r.body for r in reviews),
                "missing_date_text": sum(r.date_text is None for r in reviews),
                "missing_review_date": sum(r.review_date is None for r in reviews),
                "missing_verified_purchase": sum(
                    r.verified_purchase is None for r in reviews
                ),
                "missing_helpful_votes": sum(r.helpful_votes is None for r in reviews),
            }

            if args.check_pagination:
                pagination_page = context.new_page()
                try:
                    check_public_page_two(
                        pagination_page, asin, origin, report, args.output_dir
                    )
                finally:
                    pagination_page.close()
        except (PlaywrightError, RuntimeError, ValueError, OSError) as exc:
            report["success"] = False
            report["error_type"] = type(exc).__name__
            # Do not persist raw exception strings, which can include proxy details.
            report.setdefault("page_type", "navigation_or_browser_error")
            if page is not None:
                try:
                    save_browser_evidence(page, args.output_dir, "failed_run")
                except (PlaywrightError, OSError):
                    report["evidence_capture_failed"] = True
            raise
        finally:
            (args.output_dir / "verification_report.json").write_text(
                json.dumps(report, indent=2, ensure_ascii=False), encoding="utf-8"
            )
            if browser is not None:
                browser.close()
            print(f"Run record: {args.output_dir.resolve()}")

    print(f"HTTP status: {report.get('http_status')}")
    print(f"Page type: {report.get('page_type')}")
    print(f"Review containers found: {report.get('review_containers_found')}")
    print(f"Reviews parsed: {report.get('reviews_parsed')}")
    print(
        "CSV round-trip validated: "
        f"{report.get('export_validation', {}).get('round_trip_validated')}"
    )
    if args.check_pagination:
        print(f"Pagination page type: {report.get('pagination_page_type')}")
        print(
            "Public pagination available: "
            f"{report.get('public_pagination_available')}"
        )
    print(f"Saved: {args.output_dir.resolve()}")
    return 0


if __name__ == "__main__":
    try:
        raise SystemExit(run())
    except (ValueError, RuntimeError, PlaywrightError, OSError) as exc:
        message = str(exc) if isinstance(exc, (ValueError, RuntimeError)) else "Browser, network, or file operation failed. See the run report; check proxy authentication, DNS, connection, and timeout settings."
        print(f"ERROR: {message}", file=sys.stderr)
        raise SystemExit(2)

Step 5: Run the Scraper and Reopen the Output

Run the script from PowerShell:

python amazon_review_scraper.py `
  "https://www.amazon.com/dp/B07FZ8S74R" `
  --capture-screenshots `
  --output-dir output

For macOS or Linux, use the same command on one line after activating the environment:

python amazon_review_scraper.py "https://www.amazon.com/dp/B07FZ8S74R" --capture-screenshots --output-dir output

To make one bounded page-2 check, add --check-pagination. The check stays on the supported Amazon.com host, waits no more than 30 seconds for navigation to commit, and records the page type. It is diagnostic and does not append page-2 records to the export. It does not continue when Amazon returns sign-in, a challenge, or a missing page.

Each run creates a new timestamped subdirectory under output. On a successful run, that subdirectory contains:

  • amazon_reviews.csv
  • amazon_reviews.json
  • verification_report.json with environment, script hash, counts, and result
  • source.html with the captured input and its hash in the report
  • optional raw browser screenshots

After writing CSV and JSON, the script reopens both files and compares every field with the original records. Each reopened file is also checked for unique record keys and nonempty bodies. JSON retains nulls, numbers, and booleans. CSV represents null as an empty cell and other values as strings, including True for a present verified-purchase badge. The comparison accounts for that CSV representation.

Success requires success: true and export_validation.round_trip_validated: true in the current run’s report, with a nonzero review count and all_fields_compared: true. A file’s existence or an HTTP 200 response is insufficient. The report also records exported-file hashes and the number of containers not exported.

PowerShell showing reopened CSV rows and export validation counts

Rendered current-run verification report showing eight records and successful CSV and JSON field comparisons

On failure, inspect the new run’s report and diagnostic files. Do not use CSV or JSON from an earlier directory as the latest result. Treat raw HTML and screenshots as private evidence: they may contain names, review text, session data, or third-party media even though the exported review schema omits names. Redact before sharing and apply your project’s retention limits.

Step 6: Detect Challenge, Sign-In, and Missing Pages

Amazon Continue Shopping challenge displayed instead of the requested product page

Amazon missing-page response at a product review page-2 URL

An Amazon review scraper should identify content states before it trusts any selector. The script checks the final URL, page title, and HTML for these outcomes:

  • product_page
  • continue_shopping_challenge
  • robot_check
  • captcha
  • review_sign_in_required
  • sign_in
  • not_found
  • unexpected_page

The script evaluates page text and challenge controls after removing scripts and style elements. A JavaScript variable containing captcha alone does not classify the page as a challenge. These are heuristic checks, not a guarantee of detecting every future Amazon layout. It rechecks the content before exporting.

HTTP 403, 407, 429, and other error responses stop the run. For 429, the report preserves Retry-After. There is no automatic retry: wait at least the indicated interval, whether expressed as seconds or an HTTP date, and confirm that a later attempt is permitted. Without that header, stop and review the access policy instead of selecting an arbitrary aggressive retry interval. Browser, network, and filesystem failures also exit with an error and write a report where initialization has progressed far enough.

Optional Proxy Routing for an Authorized Test

The current live run and local tests used no proxy. No live proxy account was used for this revision, so the configuration below is not presented as verified proxy-authentication evidence.

For an authorized regional test, a web scraping proxy can keep the network route consistent with the marketplace being evaluated. A residential proxy may be appropriate when the project genuinely needs a residential exit region, but it does not grant access to login-only reviews and does not make an otherwise prohibited collection acceptable.

The script reads proxy settings from environment variables so credentials are not committed to source control:

$Env:ROLA_PROXY_SERVER="http://HOST:PORT"
$Env:ROLA_PROXY_USERNAME="YOUR_USERNAME"
$Env:ROLA_PROXY_PASSWORD="YOUR_PASSWORD"
python amazon_review_scraper.py "https://www.amazon.com/dp/B07FZ8S74R"

Replace HOST, PORT, YOUR_USERNAME, and YOUR_PASSWORD with the values issued for your account. Refer to Python proxy integration for account setup, and check your provider dashboard for the assigned endpoint. This example does not claim that any particular endpoint or authentication method was tested. Never include real credentials, account cookies, or session tokens in screenshots.

Common Amazon Review Scraping Errors

Symptom Likely cause How to verify Correct response
HTTP 200 but zero reviews Reviews are absent, lazy content did not load, or sign-in is required Inspect the final HTML and screenshot Scroll progressively once; then stop and record the state
Continue Shopping page Amazon returned an access challenge Match the visible message and final screenshot Stop; do not automate challenge completion
Sign-in message in the review section Reviews are gated for that context Search HTML for “Sign in to see customer reviews” Do not inject personal cookies
Page Not Found Invalid, unavailable, or unsupported target URL Check the final title, URL, and screenshot Correct the ASIN or mark it unavailable
Containers found but no records exported Bodies are absent or body selectors no longer match Inspect one review container and its body hook Update selectors only after checking the current markup
Duplicate output rows The same card loaded more than once Compare IDs or content hashes Deduplicate before export
Missing helpful-vote count Amazon did not display one Inspect the individual review Keep null; do not convert it to zero
CSV exists but validation fails A field, schema, or record changed Compare the current run’s reopened files against original records Treat that run as failed
HTTP 403 or redirect stopped Access denied, or a document redirect hit this tutorial’s stop rule Check the status and saved page Stop; confirm the canonical URL and permission
HTTP 429 Rate limit Read retry_after in the report Stop; respect the indicated interval before any later permitted run
HTTP 407 or proxy authentication error Proxy credentials or authentication mode Check the account-issued endpoint and authentication method Correct configuration; do not publish credentials
DNS, refused connection, or timeout Network, endpoint, browser, or target availability Check connectivity and the error type separately from page content Resolve the specific fault; do not assume poor IP quality

Use HTTP status, content state, and network diagnostics together. A status code alone cannot distinguish all product pages, sign-in screens, challenges, and error pages.

Scope, Limitations, and Responsible Use

This example collects only review records displayed on a public product page in the tested browser context. When scraping Amazon reviews, treat every session as a new observation: this script does not promise complete coverage, stable selectors, or consistent results across regions, products, and Amazon experiments.

Before scraping Amazon customer reviews:

  • Confirm that you are authorized to collect and use the data.
  • Review Amazon’s current Conditions of Use and crawler directives for the intended path and crawler. The Conditions of Use URL returned HTTP 403 during this revision, so its current terms have not been verified here; this guide makes no claim that Amazon permits the proposed collection.
  • Treat robots.txt as a technical directive, not a license or legal opinion.
  • Do not copy account cookies, purchase history, or login-only content into this workflow.
  • Avoid collecting reviewer identities when they are unnecessary.
  • Do not republish long review text or customer media without the required rights.
  • Use finite timeouts, low request volume, and explicit stop conditions. This script performs no automatic retries.
  • Prefer a licensed source or written data agreement when completeness or redistribution rights matter.

This article provides technical guidance, not legal advice.

Conclusion

The reliable way to scrape reviews from Amazon is not to assume success. Validate the ASIN, render a fresh browser context, scroll in bounded steps, parse only real review containers, deduplicate records, reopen the exports, and stop on sign-in, challenges, or missing pages.

The current live run exported and verified eight records, while the regression suite checks failure handling and data integrity with synthetic fixtures. Neither result establishes a live Amazon success rate. If your authorized project needs stable regional routing, Rola IP can be evaluated as the network layer, while access rights and page availability must still be handled separately.

Frequently asked questions