Skip to content
vakandiPublic

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

Website Cloner Chrome Extension

A Chrome extension that clones websites by crawling pages and downloading all assets, similar to mirror_playwright.py but running natively in Chrome for better bot detection bypass.

Features

  • Same-domain crawling: Crawls all pages on the same domain as the start URL
  • Asset collection: Downloads images, CSS, JavaScript, SVGs, and other assets
  • Link rewriting: Rewrites HTML links to point to local downloaded assets
  • Rate limiting: Configurable delay between requests with randomization
  • Depth limiting: Limits crawl depth to prevent infinite loops
  • Max pages limit: Stops after reaching a specified number of pages
  • Robots.txt support: Optional robots.txt checking (can be disabled)
  • Progress tracking: Real-time progress display showing pages saved, assets saved, and queue length

Installation

  1. Open Chrome and navigate to chrome://extensions/
  2. Enable "Developer mode" (toggle in the top right)
  3. Click "Load unpacked"
  4. Select the cloner_extension folder
  5. The extension should now appear in your extensions list

Usage

  1. Click the extension icon in the Chrome toolbar
  2. Enter the following settings:
    • Start URL: The URL to start cloning from (e.g., https://example.com)
    • Max Pages: Maximum number of pages to download (default: 200)
    • Max Depth: Maximum crawl depth (default: 4)
    • Delay: Delay in seconds between requests (default: 1.0)
    • Ignore robots.txt: Check this to bypass robots.txt restrictions
  3. Click "Start Cloning"
  4. Monitor progress in the status section
  5. Click "Stop" to stop the crawler at any time

How It Works

  1. The extension opens tabs in the background to visit each page
  2. Content scripts extract links and assets from each page
  3. Assets are downloaded using the Chrome Downloads API
  4. HTML is rewritten to point to local asset files
  5. New links are added to the queue for crawling
  6. The process continues until max pages or depth is reached

File Structure

The cloned website will be saved in your default Downloads folder with the following structure:

Downloads/
├── example.com/
│   ├── index.html
│   ├── page1.html
│   ├── page2/
│   │   └── index.html
│   └── ...
└── assets/
    ├── image1.png
    ├── stylesheet.css
    ├── script.js
    └── ...

Limitations

  • Chrome's Downloads API saves files to the default Downloads folder, not a custom location
  • Large websites may take a long time to clone
  • Some JavaScript-heavy sites may not be fully rendered (Chrome extension limitations)
  • Cloudflare and other bot protection may still block the crawler in some cases

Privacy

  • The extension requires access to all websites (<all_urls> permission) to crawl them
  • It uses Chrome's Storage API to save your settings locally
  • No data is sent to external servers

Development

Files in this extension:

  • manifest.json: Extension configuration and permissions
  • background.js: Main crawler logic (service worker)
  • content.js: Content script for extracting page data
  • popup.html/js/css: User interface
  • utils.js: Utility functions for URL handling
  • robots.js: Robots.txt parser and checker

Notes

  • The extension runs as a service worker, which may be suspended by Chrome
  • For very large sites, consider using the Python script (mirror_playwright.py) instead
  • Make sure you have permission to crawl websites you're cloning
  • Use the "Ignore robots.txt" option responsibly

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages