Skip to content

fix: perch coverage reached half the methods and over-counted kills - #345

Closed
jbrown9513 wants to merge 6 commits into
mainfrom
fix/coverage-reach
Closed

jbrown9513 wants to merge 6 commits into
mainfrom
fix/coverage-reach

Conversation

@jbrown9513

Copy link
Copy Markdown
Contributor

perch coverage on perch itself said 43% of methods had a test. The tests reach about 98% of them. Two things were wrong, and both were the call graph's, not the model's.

The walk stopped three calls from each test, a number nobody measured. On perch, where tests call main() and the work is five calls down, that hid about 130 methods; past depth 5 the count stops moving, so there is no depth to pick. The walk now follows every call until there are none.

The graph could not follow four things every JavaScript codebase does, and most others:

  • A parameter. analyzeFiles(files, { analyzer }) calls analyzer.analyzeSource(), and analyzer held nothing, so the whole parser looked untested. Every method now records its parameters, every call records what it passes (positionally, by name, and the fields of an object literal passed whole), and a method called on a parameter resolves to a method of whatever each caller passes, read at that caller's own call site. Callers found this way are callers too, so the rounds repeat until nothing new resolves.
  • An object a function returns. const files = liveCounter(io); files.update(3): the member of the object literal the factory returns, written out or through the local it returns.
  • A table. commands[name](io), and const command = commands[name]; command(io): every function in the table, since the key is the run's to choose.
  • A callback. A function written inside another can only run if its parent ran, so the parent reaches it: a visitor handed to walk, a worker's handler, a retry's body.

Also: a factory declared to return an interface (createAnalyzer(): Analyzer) is read as returning the class it builds when the interface has no members; a property read that hits a getter (node.type) is the call it is, in JavaScript, TypeScript and Python as in Scala; and references.ts had a local named using, which the TypeScript grammar reads as a keyword, so the file had not parsed in full and none of its 72 functions could be reached. On perch, reach goes from 483 of 1,114 methods to 1,105 of 1,129; what is left is Node's private methods, a ternary between two tables, and functions no test calls.

With reach honest, "Methods tested" is gone and the score is counted the way Stryker and PIT count it: every method gets its mutants, a method no test reaches has its mutants made and asked about nothing, they are No coverage and count against the score, and the covered score beside it leaves them out. --depth is gone with the walk's limit.

With every test in reach, the score went to 99%, and that was false too. A mutant was killed when the product of its tests' chances of missing it fell under a half, so eight tests each put at 0.3 made a near-certain kill out of a mutant none of them was likely to catch. On 149 mutants of perch run for real, 59 of 76 it called killed had survived. A mutant is now killed when one test is likely enough to fail against it, and that likelihood is KILLED, 0.7: against the same runs, a test the model put at 0.5 to 0.7 failed one time in five, and one at 0.7 or over two times in three. At 0.5 the rule claimed 68 kills for 38 real ones; at 0.7 it claims 41, 30 of them right, and 100 of the 108 it calls survived had survived. Redundancy and a test's kill list use the same threshold. Each mutant now records fails, the chance each test in asked fails against it, which is what the check was made from.

On the shop fixture, 27 of 27 mutants and 95 of 95 per-test predictions match the real runs. On perch: 1,129 methods, 1,105 reached, 20,374 mutants, a score of 28% with 7,406 survived and 180 with no coverage, for $0 on cached answers.

Closes #344

A mutant was killed when the product of its tests' chances of missing it fell under a half, so eight tests each put at 0.3 made a near-certain kill. On 149 mutants of perch run for real, 59 of 76 it called killed had survived. A mutant is now killed when one test is at KILLED, 0.7, or over: a test the model put at 0.5 to 0.7 failed one time in five, one at 0.7 or over two times in three, and the rule claims 41 kills for 38 real ones against 68 at 0.5. Each mutant records fails, the chance each asked test fails against it.
… and sets no probability

The redundancy rule selected the file's top level through the 'redundant' column name in DIFF_TOTALS and read it broken at 66% on main; it now selects code that uses redundantWith. The probabilities rule names KILLED beside min as a floor on an answer. The docs page's longest sentences are split, and the module header states the redundancy rule as the code applies it.
@perchcode-bot

perchcode-bot Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Perch Scan: no issues in 486 methods at 12fd8b2.

See the scan

@jbrown9513

Copy link
Copy Markdown
Contributor Author

Folded into #355.

@jbrown9513 jbrown9513 closed this Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

perch coverage reported 43% of perch's methods tested when the tests reach 98%

1 participant