18 AI Prompts for Web Scraping

The best Web Scraping prompts in the Coding library. Tested on ChatGPT, Claude, Gemini and every major model.

Browse the prompts

## Role

You are an expert web scraping engineer specializing in ethical data extraction, Python development, and compliance with website policies and robots.txt standards.

## Task

Generate a complete, production-ready Python web scraping script that extracts structured data from the specified URLs while adhering to ethical scraping practices: robots.txt compliance, rate limiting, graceful error handling, user-agent rotation, and proper attribution.

## Context

{{scraping-requirements}}

The script must:
- Check and respect robots.txt before scraping
- Implement polite request delays and user-agent rotation
- Identify appropriate HTML selectors with fallback logic for structure changes
- Detect and handle pagination automatically
- Include comprehensive error handling and logging
- Output clean, timestamped data with source attribution
- Be modular, well-documented, and maintainable

Generate Web Scraping Script

Generates production-ready Python web scraping scripts that respect robots.txt, implement rate limiting, and extract structured data ethically. Runs on ChatGPT, Claude, and Cursor.

95
You are an expert web scraper specialising in job-board data extraction and normalisation.

Write a complete, production-ready Python scraper that collects job postings from {{target-url}} and outputs them into a structured CSV tracker with the following columns:

- role_title
- company
- location (normalised to "City, Country" or "Remote")
- remote_status (one of: Full Remote | Hybrid | On-site | Not Specified)
- salary_range (normalised to "$min–$max currency/period" or "Not Listed")
- posting_date (ISO 8601 format YYYY-MM-DD, inferred if relative like "2 days ago")
- application_link (absolute URL)
- listing_id (unique stable identifier derived from the posting to enable deduplication)
- status (Active | Closed)
- last_seen (ISO 8601 timestamp of last successful scrape)

## Functional requirements

1. **Deduplication logic**: use listing_id to detect reposts. If a job with the same li

Job Listing Scraper and Tracker

Generate a robust web scraper in Python that extracts and tracks job postings in a structured CSV format from specified job boards.

18

Horse Racing Data Scraper

Automate horse racing data collection with this AI prompt, streamlining information gathering from multiple websites including your RacingTV account.

17

Scraper Breakage Detection Harness

Generate a test harness for web scrapers to detect silent breakage due to site markup changes. Produces code output.

17

Paginated Scraper With Resume Support

Generate a resilient paginated scraper with automatic pagination detection and recovery features using Python.

17
You are an expert Python developer specializing in ethical web scraping and data engineering.

Write a complete, production-ready web scraper that collects publicly listed business contact information from a specified directory site and outputs a structured lead sheet. The scraper must:

**Input**
- Target URL: {{directory-url}}
- Output format: {{output-format}}

**Data to Extract**
For each business listing, collect:
- Business name
- Category / industry
- Physical address
- Phone number
- Email address (only if publicly displayed on the listing page)
- Website URL

**Compliance Requirements**
The script MUST:
1. Check and respect the site's robots.txt before scraping
2. Skip any content behind authentication / login walls
3. Only collect publicly published business contact details (no personal data, no scraping of employee profiles or internal pages)
4. Include reasonable rate limitin

Business Directory Contact Extractor

Generate a Python script to extract business contact information from a public directory, ensuring compliance with web scraping ethics and data protection laws.

15
You are an expert Python developer specializing in ethical web scraping and data engineering.

Build a scheduled price-tracking scraper that monitors competitor products over time and detects significant price changes.

# Requirements

## Input
{{scraper-config}}

## Core Features

**Data Collection**
- Extract product identifier, current price, and timestamp on each run
- Append new records to a persistent CSV or SQLite historical store (never overwrite)
- Handle missing prices, unavailable products, and parsing errors gracefully

**Change Detection**
- Compare current prices against the most recent prior record for each product
- Flag any price change that exceeds the threshold specified in the config
- Calculate both absolute change (currency units) and percentage change

**Reporting**
- Generate a simple summary after each run showing:
  - Products with flagged changes, sorted by mag

Competitor Price Monitoring Scraper

Generate a Python-based web scraper for tracking and reporting competitor product price changes over time, with ethical compliance.

14

Customer Contact Information Scraper

Generate customer contact databases with this AI prompt, extracting primary contact names, job titles, email addresses, and phone numbers efficiently.

13

Playwright Browser Automation Scraper

Generate a production-ready scraping script for JavaScript-rendered sites using Playwright. Produces structured JavaScript output with handling for common obstacles like overlays and lazy-loads.

12
You are a tactical intelligence analyst specializing in rapid-response synthesis from raw communication feeds. Your mission is to transform unstructured chat logs, message threads, and social media exports into actionable intelligence briefs that drive immediate decision-making.

# Input Material

{{raw-communication-data}}

# Analysis Framework

Apply structured intelligence tradecraft to the input:

**Event Extraction**: Identify the 5 Wsβ€”Who (actors, organizations, roles), What (specific events, claims, decisions), When (timestamps, deadlines, time-sensitive windows), Where (locations, platforms, channels), and Why (stated motivations, inferred intent).

**Impact Assessment**: Evaluate immediate consequences and second-order effects. Consider operational risks, reputational exposure, resource implications, and stakeholder impact. Distinguish between confirmed impacts and probabilistic

TGscrape

Generate tactical intelligence briefs by analyzing raw communication data from various platforms. Produces structured outputs that prioritize decision-making.

11
You are an expert data engineer specializing in web scraping and ETL pipelines for sentiment analysis.

Write a production-ready web scraper that collects public product reviews from {{target-site}} and outputs a clean, analysis-ready dataset. The scraper must extract:

- Review text (full body)
- Star rating
- Review date
- Verified purchase flag (boolean)
- Helpful vote count

**Privacy & compliance requirements:**
- Exclude all reviewer names, profile links, and personal identifiers
- Respect robots.txt and implement polite rate-limiting (minimum 1-second delay between requests)
- Handle authentication/session management if required by the site

**Pagination & reliability:**
- Page through all available reviews using the site's pagination mechanism (query parameters, infinite scroll, or "load more" buttons)
- Implement retry logic with exponential backoff for failed requests
- Log pro

Product Review Scraper for Sentiment Analysis

Generate a web scraper script that collects and cleans product reviews for sentiment analysis.

11

Article Content Extractor for Research

Generate a Python script that scrapes and extracts structured article content from news or blog pages using readability heuristics.

10

Web Table to Spreadsheet Converter

Generate spreadsheet-ready data from HTML table markup, handling complex structures and data cleaning.

10

HTML to Structured JSON Parser

Generate a robust HTML parser and validation function to convert messy HTML into structured JSON matching a user-supplied schema.

10
You are a web scraping compliance and reliability auditor with expertise in ethical data collection, legal requirements, and production-grade scraper architecture.

Review the scraper code below for compliance, reliability, and best practices. Examine:

1. **robots.txt handling** – Does the code fetch and respect robots.txt? Are disallowed paths honored? Is the crawl-delay directive observed?
2. **Request rate & concurrency** – Are requests throttled appropriately? Does concurrency risk overwhelming the target server?
3. **User-Agent honesty** – Is the User-Agent string present, accurate, and identifying? Does it include contact information?
4. **Caching strategy** – Does the scraper cache responses to avoid redundant requests? Are conditional requests (ETag, Last-Modified) used where applicable?
5. **Retry & backoff behavior** – Are transient errors retried? Is exponential backoff imple

Polite Scraper Compliance Checklist

Generate a compliance and reliability review for a web scraper, including guidance for ethical and legal practices.

8
You are an expert Python web scraping engineer writing production-grade, ethical scrapers.

Write a complete, runnable Python script that extracts structured product data from an e-commerce category page and exports it to CSV. The script must:

**Compliance & Ethics**
- Check the site's robots.txt and print a warning if scraping the path may be disallowed
- Include a comment reminding the user to review the site's Terms of Service before running
- Use a realistic User-Agent header (desktop browser)
- Implement polite rate limiting (2–3 second delay between requests if paginating)

**Data Extraction**
Extract these fields for every product on the page:
- Product name
- Price (numeric value)
- Currency
- Rating (numeric, e.g. 4.5)
- Review count (integer)
- Availability status (in stock / out of stock / pre-order, etc.)
- Product URL (absolute)

**Robustness**
- Handle missing or malformed

Ecommerce Product Listing Scraper

Generate a Python web scraper script that extracts product data from e-commerce category pages into a CSV file, following ethical standards.

7

Lunch atop a Skyscraper

Generate a vintage black-and-white photograph of humanoid robotic power armor suits perched on a steel beam above a city skyline.

6

Event Scraper and Analyzer

Generate targeted event opportunities with this AI prompt, finding conferences, conventions, investor meetings, exhibitor pricing, success metrics, and strategic recommendations.

5

What are AI prompts for Web Scraping?

AI prompts for Web Scraping are engineered instructions that already work. These are not one-line questions. Each one fixes the role, the context, the task and the output format before you type a word, so you get a usable result on the first run instead of the fourth.

They cover the work Web Scraping actually get asked for: research and briefs, copy and content, analysis and reporting, planning, outreach and the admin that eats the day. Open a card to see the full prompt and the output it returns.

Popular on this page right now: "Generate Web Scraping Script", "Job Listing Scraper and Tracker", "Horse Racing Data Scraper".

18 on this page, every one scoped to Web Scraping. Free to read, free to copy.

Why these prompts work for Web Scraping

A weak prompt costs you the hour you were trying to save: you rewrite it three times, get something generic, then finish the job by hand. An engineered prompt front-loads that thinking once.

In Web Scraping that means first drafts you can send, analysis you can act on, and the repetitive work handed off, so the time goes into judgement instead of typing.

Every prompt here was written for a real job and tested against the models people actually use. Nothing scraped from a thread.

How to use these prompts

Open a prompt, copy it, and replace the [bracketed] variables with your own product, audience or topic. The structure around them stays as is. That structure is the part doing the work.

Paste it into ChatGPT, Claude, Gemini, Grok or the model you already use. If the output drifts, tighten the context line instead of rewriting the whole prompt.

No account needed to copy one. No setup, no extension, nothing to install.

Which AI tool works best for Web Scraping prompts?

Text prompts here run well in ChatGPT, Claude, Gemini and Grok; image prompts target Midjourney and Nano Banana. Each card lists the models it was tested with.

Are these AI prompts free to use?

A big part of the library is free: open a prompt, copy it, use it. Premium packs and the Complete AI Bundle unlock the full collection with lifetime updates.

How do I adapt these prompts to my use case?

Start with the [variables]: niche, audience, constraints. If the result still misses, add one example of the output you want. A single good example beats three extra instructions.

For a prompt built from scratch, the Start Now card above opens the custom prompt generator.

Related resources

Get smarter on AI every week

One email a week with the best new prompts, tools, and model updates. Unsubscribe anytime.

Join 100,000+ subscribers. One email a week, real prompts, tools, and model updates. Unsubscribe anytime.