A Chrome extension that clones websites by crawling pages and downloading all assets, similar to mirror_playwright.py but running natively in Chrome for better bot detection bypass.
- Same-domain crawling: Crawls all pages on the same domain as the start URL
- Asset collection: Downloads images, CSS, JavaScript, SVGs, and other assets
- Link rewriting: Rewrites HTML links to point to local downloaded assets
- Rate limiting: Configurable delay between requests with randomization
- Depth limiting: Limits crawl depth to prevent infinite loops
- Max pages limit: Stops after reaching a specified number of pages
- Robots.txt support: Optional robots.txt checking (can be disabled)
- Progress tracking: Real-time progress display showing pages saved, assets saved, and queue length
- Open Chrome and navigate to
chrome://extensions/ - Enable "Developer mode" (toggle in the top right)
- Click "Load unpacked"
- Select the
cloner_extensionfolder - The extension should now appear in your extensions list
- Click the extension icon in the Chrome toolbar
- Enter the following settings:
- Start URL: The URL to start cloning from (e.g.,
https://example.com) - Max Pages: Maximum number of pages to download (default: 200)
- Max Depth: Maximum crawl depth (default: 4)
- Delay: Delay in seconds between requests (default: 1.0)
- Ignore robots.txt: Check this to bypass robots.txt restrictions
- Start URL: The URL to start cloning from (e.g.,
- Click "Start Cloning"
- Monitor progress in the status section
- Click "Stop" to stop the crawler at any time
- The extension opens tabs in the background to visit each page
- Content scripts extract links and assets from each page
- Assets are downloaded using the Chrome Downloads API
- HTML is rewritten to point to local asset files
- New links are added to the queue for crawling
- The process continues until max pages or depth is reached
The cloned website will be saved in your default Downloads folder with the following structure:
Downloads/
├── example.com/
│ ├── index.html
│ ├── page1.html
│ ├── page2/
│ │ └── index.html
│ └── ...
└── assets/
├── image1.png
├── stylesheet.css
├── script.js
└── ...
- Chrome's Downloads API saves files to the default Downloads folder, not a custom location
- Large websites may take a long time to clone
- Some JavaScript-heavy sites may not be fully rendered (Chrome extension limitations)
- Cloudflare and other bot protection may still block the crawler in some cases
- The extension requires access to all websites (
<all_urls>permission) to crawl them - It uses Chrome's Storage API to save your settings locally
- No data is sent to external servers
Files in this extension:
manifest.json: Extension configuration and permissionsbackground.js: Main crawler logic (service worker)content.js: Content script for extracting page datapopup.html/js/css: User interfaceutils.js: Utility functions for URL handlingrobots.js: Robots.txt parser and checker
- The extension runs as a service worker, which may be suspended by Chrome
- For very large sites, consider using the Python script (
mirror_playwright.py) instead - Make sure you have permission to crawl websites you're cloning
- Use the "Ignore robots.txt" option responsibly