Building Resilient Scrapers With Playwright and Redis
What You’ll Need
To follow along with this tutorial, you will need the following tools and services:
- A Linux server: You can use a Hetzner VPS or DigitalOcean droplet for hosting your scraper and Redis instance.
- Alternatively, a Contabo VPS provides cheap compute if you need heavy concurrency.
- Node.js (version 18 or higher) installed on your local machine or server.
- Docker and Docker Compose installed for quick database and Redis deployment.
- n8n Cloud if you plan to orchestrate your scraper triggers using no-code workflows.
- Namecheap if you plan to register a custom domain for an admin monitoring API.
Table of Contents
- The Architecture of a Resilient Distributed Scraper
- Setting Up Infrastructure with Docker Compose
- Implementing the Redis Queue Producer
- Building the Worker with Playwright and Retries
- Handling Failures and Dead Letter Queues
- Getting Started
The Architecture of a Resilient Distributed Scraper
Web scraping at scale frequently breaks down due to dynamic JavaScript rendering, memory leaks in headless browsers, IP rate limits, and network drops. If you run a single script that loops through thousands of URLs with Playwright, a single memory leak or unhandled exception will crash your entire job, forcing you to restart from scratch.
To solve this, we decouple URL discovery from extraction using a producer-consumer architecture. Redis acts as our centralized queue and state database, while Playwright instances run across isolated stateless worker processes.
In this setup:
- The Producer discovers target URLs, checks Redis to prevent duplicate processing, and pushes jobs into a Redis List or Stream.
- The Worker pops job payloads, launches isolated Playwright browser contexts, fetches the page data, and acknowledges completion.
- Errors trigger exponential backoff retry logic. If a job fails repeatedly, it is moved to a Dead Letter Queue (DLQ) for analysis.
This pattern allows us to process thousands of pages across multiple servers without dropping jobs or double scraping. Once you extract raw content, you can push payloads down to storage layers. For instance, if you are saving high throughput data to PostgreSQL, you should review our guide on Configuring PgBouncer Connection Pooling for PostgreSQL to prevent worker connection starvation. Furthermore, if you plan to transform scraped unstructured text into vector representations, read our step-by-step walkthrough on Connecting OpenAI embeddings to vector databases.
Setting Up Infrastructure with Docker Compose
We will deploy our system using Docker Compose. We will provision a Redis container configured with AOF (Append Only File) persistence to guarantee that job queues survive container restarts.
Spin up your environment on a Hetzner VPS by creating a docker-compose.yml file:
version: '3.8'
services:
redis:
image: redis:7-alpine
container_name: scraper_redis
command: redis-server --appendonly yes --requirepass SuperSecretRedisPassword123!
ports:
- "6379:6379"
volumes:
- redis_data:/data
restart: always
volumes:
redis_data:
Start the Redis server by running the command below in your terminal:
docker compose up -d
Verify that Redis is healthy by pinging it through the Docker CLI:
docker exec -it scraper_redis redis-cli -a SuperSecretRedisPassword123! ping
If it returns PONG, your queue backbone is ready.
💡 Fast-Track Your Project: Don’t want to configure this yourself? I build custom n8n pipelines and bots. Message me with code SYS3-HUGO.
Implementing the Redis Queue Producer
The producer component adds targets to the queue. It uses a Redis Set named scraped_urls for atomic deduplication alongside a Redis List named scraping_queue for job distribution.
First, initialize a Node.js project and install the required dependencies:
npm init -y
npm install ioredis playwright
Create a file named producer.js. This script populates the Redis queue while ensuring duplicate links are ignored entirely.
const Redis = require('ioredis');
const redis = new Redis({
host: '127.0.0.1',
port: 6379,
password: 'SuperSecretRedisPassword123!'
});
const TARGET_URLS = [
'https://news.ycombinator.com/news?p=1',
'https://news.ycombinator.com/news?p=2',
'https://news.ycombinator.com/news?p=3',
'https://news.ycombinator.com/news?p=1'
];
async function enqueueJobs() {
console.log('Starting job enqueueing process...');
for (const url of TARGET_URLS) {
const isNew = await redis.sadd('scraped_urls', url);
if (isNew === 1) {
const jobPayload = {
url: url,
retries: 0,
createdAt: new Date().toISOString()
};
await redis.rpush('scraping_queue', JSON.stringify(jobPayload));
console.log(`Successfully queued: ${url}`);
} else {
console.log(`Skipped duplicate URL: ${url}`);
}
}
const queueLength = await redis.llen('scraping_queue');
console.log(`Current queue length: ${queueLength}`);
process.exit(0);
}
enqueueJobs().catch((err) => {
console.error('Producer error:', err);
process.exit(1);
});
Execute the producer script:
node producer.js
The output will confirm that duplicate entries are ignored during the enqueue phase.
Building the Worker with Playwright and Retries
Workers consume jobs from Redis using the blocking pop operation BLPOP. This allows the process to sleep until work is available, eliminating idle CPU overhead.
To prevent memory leaks common to long running headless browser instances, we reuse a single browser process while creating and destroying isolated BrowserContext objects for individual page loads.
Create a file named worker.js:
const { chromium } = require('playwright');
const Redis = require('ioredis');
const redis = new Redis({
host: '127.0.0.1',
port: 6379,
password: 'SuperSecretRedisPassword123!'
});
const MAX_RETRIES = 3;
const QUEUE_TIMEOUT_SECONDS = 5;
async function scrapePage(context, url) {
const page = await context.newPage();
try {
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30000 });
const pageTitle = await page.title();
const headings = await page.$$eval('h1, h2, td.title', elements =>
elements.map(el => el.textContent.trim()).filter(text => text.length > 0)
);
return {
title: pageTitle,
headingCount: headings.length,
sampleHeading: headings[0] || null
};
} finally {
await page.close();
}
}
async function startWorker() {
console.log('Worker started. Launching Chromium...');
const browser = await chromium.launch({ headless: true });
while (true) {
try {
const result = await redis.blpop('scraping_queue', QUEUE_TIMEOUT_SECONDS);
if (!result) {
console.log('Queue empty. Waiting for new jobs...');
continue;
}
const [queueName, rawData] = result;
const job = JSON.parse(rawData);
console.log(`Processing URL: ${job.url} (Attempt ${job.retries + 1})`);
const context = await browser.newContext({
userAgent: 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
});
try {
const scrapedData = await scrapePage(context, job.url);
console.log(`Successfully scraped ${job.url}:`, scrapedData);
await redis.hset('scraped_results', job.url, JSON.stringify(scrapedData));
} catch (scrapeError) {
console.error(`Failed to scrape ${job.url}:`, scrapeError.message);
if (job.retries < MAX_RETRIES) {
job.retries += 1;
job.lastError = scrapeError.message;
console.log(`Re-queueing ${job.url} with retry count ${job.retries}`);
await redis.rpush('scraping_queue', JSON.stringify(job));
} else {
console.error(`Max retries reached for ${job.url}. Moving to DLQ.`);
job.failedAt = new Date().toISOString();
job.finalError = scrapeError.message;
await redis.rpush('dead_letter_queue', JSON.stringify(job));
}
} finally {
await context.close();
}
} catch (workerError) {
console.error('Unexpected worker loop error:', workerError);
}
}
}
startWorker().catch(async (err) => {
console.error('Fatal worker crash:', err);
process.exit(1);
});
Run the worker process in a separate terminal tab:
node worker.js
If you plan to expose worker status metrics or push scraped payloads directly to internal microservices, ensure your HTTP endpoints are secured. You can read our detailed guide on Securing Microservice Endpoints With OAuth2 Bearer Tokens to implement proper authorization controls.
Handling Failures and Dead Letter Queues
In production systems, network instability, Cloudflare challenges, or transient DNS failures will cause some requests to fail. Instead of discarding failed URLs, our worker logic routes permanent failures into the dead_letter_queue list after three attempts.
You can create a management utility named dlq_inspector.js to inspect and replay failed jobs manually:
const Redis = require('ioredis');
const redis = new Redis({
host: '127.0.0.1',
port: 6379,
password: 'SuperSecretRedisPassword123!'
});
async function inspectAndReplayDLQ() {
const dlqLength = await redis.llen('dead_letter_queue');
console.log(`Total jobs in Dead Letter Queue: ${dlqLength}`);
if (dlqLength === 0) {
process.exit(0);
}
for (let i = 0; i < dlqLength; i++) {
const rawJob = await redis.lpop('dead_letter_queue');
if (!rawJob) break;
const job = JSON.parse(rawJob);
console.log(`Inspecting failed job: ${job.url}`);
console.log(`Failure Reason: ${job.finalError}`);
job.retries = 0;
delete job.finalError;
delete job.failedAt;
await redis.rpush('scraping_queue', JSON.stringify(job));
console.log(`Replayed ${job.url} back to main queue.`);
}
process.exit(0);
}
inspectAndReplayDLQ().catch((err) => {
console.error('DLQ Inspector error:', err);
process.exit(1);
});
Run the inspector script whenever you want to re-process failed targets after updating proxy configurations or scraper rules:
node dlq_inspector.js
By decoupling target discovery, processing queues, failure contexts, and persistence layers, you achieve an enterprise grade web scraper capable of running indefinitely without manual intervention.
Getting Started
Ready to deploy your distributed scraper? Get high performance compute infrastructure and hosting setup below:
- Spin up a cloud server on Hetzner VPS or DigitalOcean.
- Allocate high memory instances for heavy parallel scraping using Contabo VPS.
- Automate scraper schedules and webhook notifications using n8n Cloud.
- Register custom monitoring domains using Namecheap.
Outsource Your Automation
Don’t have time? I build production n8n workflows, WhatsApp bots, and fully automated YouTube Shorts pipelines. Hire me on Fiverr, mention SYS3-HUGO for priority. Or DM at chasebot.online.
Want to automate this yourself?
Start with n8n Cloud (free tier available) or self-host on a Hetzner VPS for full control.
Want this engine running on your own VPS?
This blog publishes itself — daily, unattended, on free API tiers. The full engine, Hugo theme, and setup guide are available as System 3.
Get System 3