|
| 1 | +# TAGLINE |
| 2 | + |
| 3 | +Find duplicate and near-duplicate code with tree-sitter |
| 4 | + |
| 5 | +# TLDR |
| 6 | + |
| 7 | +**Find duplicate blocks** in a project (default ruleset, exact match after normalization) |
| 8 | + |
| 9 | +```treepeat detect [path/to/project]``` |
| 10 | + |
| 11 | +**Near-duplicates** at an 80 percent threshold |
| 12 | + |
| 13 | +```treepeat detect --similarity [80] [path/to/project]``` |
| 14 | + |
| 15 | +**Structural clones**, with identifiers and constants anonymized |
| 16 | + |
| 17 | +```treepeat --ruleset loose detect [path/to/project]``` |
| 18 | + |
| 19 | +Compare **raw syntax trees**, without normalization |
| 20 | + |
| 21 | +```treepeat --ruleset none detect [path/to/project]``` |
| 22 | + |
| 23 | +Show a **side-by-side diff** of the first two hits in each group |
| 24 | + |
| 25 | +```treepeat detect --diff --min-lines [10] [path/to/project]``` |
| 26 | + |
| 27 | +Write **SARIF** for a CI consumer |
| 28 | + |
| 29 | +```treepeat detect --format sarif --output [results.sarif] [path/to/project]``` |
| 30 | + |
| 31 | +**Fail the run** when any similar group is found |
| 32 | + |
| 33 | +```treepeat detect --fail [path/to/project]``` |
| 34 | + |
| 35 | +**List the rules** in a ruleset, optionally for one language |
| 36 | + |
| 37 | +```treepeat list-ruleset default --language [python]``` |
| 38 | + |
| 39 | +See how one file is **normalized** |
| 40 | + |
| 41 | +```treepeat treesitter [path/to/file.py]``` |
| 42 | + |
| 43 | +# SYNOPSIS |
| 44 | + |
| 45 | +**treepeat** [**-l** _level_] [**-r** _ruleset_] _command_ |
| 46 | + |
| 47 | +**treepeat detect** [**-s** _percent_] [**--min-lines** _n_] [**-f** _format_] [**-o** _file_] [**--diff**] [**--fail**] _path_ |
| 48 | + |
| 49 | +**treepeat list-ruleset** [**-l** _language_] _ruleset_ |
| 50 | + |
| 51 | +**treepeat treesitter** [**-t**] _file_ |
| 52 | + |
| 53 | +# DESCRIPTION |
| 54 | + |
| 55 | +**treepeat** scans source code for duplicate and near-duplicate regions. It parses files with tree-sitter, extracts language-meaningful regions such as functions and classes, normalizes those trees, and groups regions that meet a similarity threshold. It is a code-clone finder, not a line-oriented diff. |
| 56 | + |
| 57 | +Three built-in rulesets control normalization. **none** compares raw syntax trees. **default** (the usual setting) drops whitespace, string contents, and some high-level nodes such as names, so near-copies still match. **loose** also anonymizes identifiers and constants, so the same structure with different names still matches. Pass **--ruleset** before the subcommand. |
| 58 | + |
| 59 | +**detect** walks a file or directory. Matches shorter than **--min-lines** (default 5) are dropped. **--similarity** is a percent from 5 to 100; the default of 100 means the normalized trees must match exactly. Console output lists groups and a summary. **--format sarif** writes a SARIF report instead, which is the form most CI security dashboards ingest. |
| 60 | + |
| 61 | +Paths are skipped when they match **--ignore** globs or a gitignore-style file whose name matches `.*ignore` (including **.gitignore**). **--add-regions** and **--exclude-regions** change which syntax-node types count as a region, per language, using the form `language:node1,node2`. |
| 62 | + |
| 63 | +**list-ruleset** prints the rules in **none**, **default**, or **loose**. **treesitter** prints one file beside the tokens treepeat will compare, or beside the transformed source with **--transformed**. |
| 64 | + |
| 65 | +Supported languages include Astro, Bash, CSS, Go, HTML, Java, JavaScript, JSX, Kotlin, Lua, Markdown, Python, Rust, SQL, TSX, TypeScript, and YAML. JSX is parsed with the JavaScript grammar. |
| 66 | + |
| 67 | +The tool requires **Python 3.11** or newer. Install it with `pip install treepeat`. |
| 68 | + |
| 69 | +# PARAMETERS |
| 70 | + |
| 71 | +**detect** _path_ |
| 72 | +> Scan a file or directory for similar regions. |
| 73 | +
|
| 74 | +**list-ruleset** _ruleset_ |
| 75 | +> Print rules for **none**, **default**, or **loose**. |
| 76 | +
|
| 77 | +**treesitter** _file_ |
| 78 | +> Show one file next to its normalized tree-sitter view. |
| 79 | +
|
| 80 | +**-r** _ruleset_, **--ruleset** _ruleset_ |
| 81 | +> Normalization profile: **none**, **default**, or **loose**. Default is **default**. Set this on the main command, before the subcommand. |
| 82 | +
|
| 83 | +**-l** _level_, **--log-level** _level_ |
| 84 | +> Log level: DEBUG, INFO, WARNING, ERROR, or CRITICAL. Default is WARNING. On **list-ruleset**, **-l** is the language filter instead. |
| 85 | +
|
| 86 | +**-s** _percent_, **--similarity** _percent_ |
| 87 | +> Minimum similarity, from 5 to 100. Default is 100. |
| 88 | +
|
| 89 | +**--min-lines** _n_ |
| 90 | +> Ignore regions shorter than _n_ lines. Default is 5. |
| 91 | +
|
| 92 | +**-f** _format_, **--format** _format_ |
| 93 | +> **console** (default) or **sarif**. |
| 94 | +
|
| 95 | +**-o** _file_, **--output** _file_ |
| 96 | +> Write the report to _file_. Default is standard output. SARIF uses this path; console output is printed directly. |
| 97 | +
|
| 98 | +**--diff** |
| 99 | +> In console output, show a side-by-side diff of the first two regions in each group. |
| 100 | +
|
| 101 | +**--fail** |
| 102 | +> Exit with status 1 when any similar group is reported. |
| 103 | +
|
| 104 | +**-i** _globs_, **--ignore** _globs_ |
| 105 | +> Comma-separated globs to skip, for example `*.test.py,**/node_modules/**`. |
| 106 | +
|
| 107 | +**--ignore-files** _globs_ |
| 108 | +> Comma-separated globs that locate ignore files. Default is `**/.*ignore`. |
| 109 | +
|
| 110 | +**--add-regions** _spec_ |
| 111 | +> Extra region node types, as `language:node1,node2`. Repeatable. |
| 112 | +
|
| 113 | +**--exclude-regions** _spec_ |
| 114 | +> Region labels to drop, as `language:label1,label2`. Repeatable. |
| 115 | +
|
| 116 | +**--ignore-node-types** _types_ |
| 117 | +> Comma-separated AST node types to skip while extracting regions. |
| 118 | +
|
| 119 | +**-v**, **--verbose** |
| 120 | +> After a console run, print timing and which node types were used. |
| 121 | +
|
| 122 | +**-p**, **--progress** |
| 123 | +> Progress bars on standard error while a long scan runs. |
| 124 | +
|
| 125 | +**-l** _language_, **--language** _language_ |
| 126 | +> On **list-ruleset** only: show rules that apply to one language. |
| 127 | +
|
| 128 | +**-t**, **--transformed** |
| 129 | +> On **treesitter**: show rewritten source on the right instead of tokens. |
| 130 | +
|
| 131 | +**--version** |
| 132 | +> Print the version and exit. |
| 133 | +
|
| 134 | +# CAVEATS |
| 135 | + |
| 136 | +The project is a proof of concept. Language coverage is limited, and two files in an unsupported language are simply not compared. |
| 137 | + |
| 138 | +**--ruleset** and **--log-level** belong on the main command (`treepeat --ruleset loose detect …`). **list-ruleset** reuses **-l** for a language name, so a log level has to be set before that subcommand. |
| 139 | + |
| 140 | +**--similarity** rejects values below 5, even though some examples elsewhere say the range starts at 1. **--diff** and **--verbose** affect console output only. **--fail** exits 1 for any reported group, which is easy to trip on a large tree at the default ruleset. If **detect** cannot parse any file, it exits 1 before printing groups. |
| 141 | + |
| 142 | +# HISTORY |
| 143 | + |
| 144 | +**treepeat** was written by **Dane Summers** and first published on PyPI in **November 2025**. It is released under the Apache-2.0 license. |
| 145 | + |
| 146 | +# SEE ALSO |
| 147 | + |
| 148 | +[ast-grep](/man/ast-grep)(1), [semgrep](/man/semgrep)(1), [diff](/man/diff)(1) |
| 149 | + |
| 150 | +# RESOURCES |
| 151 | + |
| 152 | +```[Source code](https://github.com/dsummersl/treepeat)``` |
| 153 | + |
| 154 | +```[Documentation](https://github.com/dsummersl/treepeat#usage)``` |
| 155 | + |
| 156 | +<!-- verified: 2026-09-24 --> |
0 commit comments