@harshvz/crawler is an open-source web scraping tool with over 1,200+ downloads on npm. It crawls websites using BFS (breadth-first) or DFS (depth-first) search algorithms and saves clean, structured content into Markdown, JSON, or CSV files.
Built with stealth features and tracker blocking, it helps developers quickly turn online documentation, articles, and knowledge bases into clean local datasets without downloading heavy browser engines.
Key Highlights
- 1,200+ npm Downloads: Used by developers for fast, local website extraction.
- Smart Crawling Strategies: Choose between Breadth-First Search (BFS) for wide scans or Depth-First Search (DFS) for deep link journeys.
- Multiple Output Formats: Export clean pages directly as Markdown, JSON, or CSV files.
- Domain Scoped: Automatically stays inside the original website so it never wanders off to external links.
- Anti-Detection & Privacy: Masks browser automation flags, spoofs realistic user-agents, and blocks third-party trackers for faster page loads.
- Lightweight Engine: Runs on Obscura via Chrome DevTools Protocol, making it much faster and smaller than standard Chromium setups.
- Interactive CLI & Node API: Run it directly from the terminal with easy prompts or import it as a TypeScript library in your code.
Quick Start
Install via npm
npm install -g @harshvz/crawler
Run the Interactive CLI
crawler
Simply enter the target URL and follow the guided prompts to set depth, output format, and filters.
Programmatic Usage
import Scraper from '@harshvz/crawler';
const scraper = new Scraper('https://example.com', {
depth: 2,
format: 'md',
delay: 500,
tags: 'h1, p, a, img',
});
// Crawl using Breadth-First Search
const results = await scraper.bfs('/');
await scraper.close();
Tech Stack
- Runtime: Node.js, TypeScript
- Engine: Chrome DevTools Protocol (CDP), Obscura Headless Engine
- Distribution: npm package, CLI executable