I build web scraping infrastructure at scale.
I specialize in high-performance crawlers, data pipelines, and scraping systems that handle hundreds of millions of requests. Currently building tools for web technology detection and large-scale data extraction.
Rust β High-perf crawlers, domain scanners, async networking
Python β aiohttp, httpx, FastAPI, data processing
Data β DuckDB, Parquet, Polars, Pandas
Infra β Cloudflare Workers, Hetzner, Oracle Cloud ARM
Approach β Pure HTTP, no browsers. Speed & scale first.
- Domain Scanner β Scanning 200M+ domains for web technology detection (Rust)
- Google Maps Scraper β Extracted hundreds of millions of records at ~100 req/s, no proxies
- Tech Detection SaaS β BuiltWith/Wappalyzer alternative powered by DuckDB + Cloudflare Workers
- Lead Generation Pipeline β Crawling β email extraction β SMTP verification, near-zero cost
| Metric | Value |
|---|---|
| Domains scanned | 200M+ |
| Google Maps records | 100M+ |
| Avg. request rate | ~100 req/s per instance |
| Infrastructure cost | ~$5/month |
I write about web scraping, data engineering, and building lean infrastructure β in Spanish and English.
- π¦ X / Twitter β Scraping tips, war stories, and data drops
- π Blog β Coming soon
Browsers are a last resort. HTTP requests are king. Scale with smart architecture, not bigger servers. Near-zero infra cost is not a limitation β it's a design goal.
Open to collaborations on scraping tools, data products, and open source infra.