Blog
Ideas for building on the open web.
Product updates, engineering notes, and practical guides for reliable web data.
Latest dispatches
From the blog
4 stories
- 01
Avoiding HTTP 429 Errors and Rate Limiting
A delay in each worker does not control the total traffic your scraper sends. Servers may enforce limits by IP address, API token, account, endpoint, time window, or the number of requests running at once. This guide explains how to diagnose 429 responses, interpret Retry-After, coordinate workers around a shared request budget, and retry without creating another burst. The goal is to match the limit the server actually enforces, rather than keep increasing a guessed delay.
Read story - 02
Scraping Behind Logins: Authentication Challenges
Scraping content behind a login means handling sign-in forms, keeping sessions active, and protecting the credentials that grant access. This guide explains how to recognize authentication failures, reuse sessions, and keep secrets out of code and logs. It focuses on authorized access through accounts you own or have explicit permission to use, with practical steps for keeping the scraper reliable and the account secure.
Read story - 03
Web Scraping for Market Research: Turn Website Changes Into Decisions
Web scraping for market research helps you compare prices, track supplier availability, and spot changes in products or purchasing terms. A useful dataset does more than collect values: it shows what changed and why it matters. Through illustrative examples—including a 5.6% price increase and a four-week lead-time extension—this guide explains what to scrape, how often to refresh it, and how to turn website changes into comparable records, useful trends, and actionable alerts.
Read story - 04
Bypassing Cloudflare and Anti-Bot Protections
Bypassing Cloudflare's anti-bot protections requires replicating a genuine browser fingerprint and using clean residential IPs, as the system blocks 99.9% of naive scrapers but not sophisticated adversaries who address both signals. A 2024 study found most commercial bypass services fail because they cannot maintain browser consistency, while successful evasion depends on solving fingerprint replication before IP reputation. This approach violates Cloudflare's Terms of Service and may breach the Computer Fraud and Abuse Act, making such methods legally risky.
Read story