Repository navigation
feat: perch coverage runs your tests and asks the model only about survivors - #355
jbrown9513 wants to merge 34 commits into
Conversation
…a table or a callback
A mutant was killed when the product of its tests' chances of missing it fell under a half, so eight tests each put at 0.3 made a near-certain kill. On 149 mutants of perch run for real, 59 of 76 it called killed had survived. A mutant is now killed when one test is at KILLED, 0.7, or over: a test the model put at 0.5 to 0.7 failed one time in five, one at 0.7 or over two times in three, and the rule claims 41 kills for 38 real ones against 68 at 0.5. Each mutant records fails, the chance each asked test fails against it.
… and sets no probability The redundancy rule selected the file's top level through the 'redundant' column name in DIFF_TOTALS and read it broken at 66% on main; it now selects code that uses redundantWith. The probabilities rule names KILLED beside min as a floor on an answer. The docs page's longest sentences are split, and the module header states the redundancy rule as the code applies it.
… what to add The summary, each file's page and the terminal listed one row per survived mutant, with the same eight test names under every one, and the code came after the rows. They now rank methods with survived mutants, most first, each with a sentence from its surest mutant saying what a new test has to tell apart, and the mutants fold under the row. A file's page shows its code first with a count in the gutter where mutants survived, one note open at a time, and next and previous survived. Source files are a tree worst first with scores coloured against 60 and 80 as Stryker colours them; test files are listed only when something needs fixing.
…tterns as Stryker does The mutant table stopped at operators, conditions, literals, statement calls and bodies; a method's own calls were never touched, so a test that never checks a filter was never found. mutants.js now swaps a method for its opposite or drops it from a chain in every language's own names, empties list and object literals and fills empty ones, makes an optional chain unconditional, blanks an arrow function, edits regular expressions, forces ternary and loop conditions, turns every assignment operator, and makes ?? an &&.
The method tables were plain objects, so a call to a name Object has a property for, toString, valueOf, constructor, hasOwnProperty, found that property and handed a function to Buffer.from. The tables have no prototype now.
…rloads A call on an interface, abstract class or trait resolved to a declaration with no body, so gson's adapters and serde_json's serializer counted as untested. A Rust impl was named by its whole type, a decorated Python definition had no class, a value handed to a call on an untyped parameter reached nothing, only one overload of a name was reached, and a Python method's receiver counted as its first argument. Each is followed now, in every language perch reads.
…ro reached nothing A value of a type parameter is a value of the trait that bounds it, a struct's field is bound as this.field with the same substitution, and a trait's method with no body runs its implementations. A macro the repository defines is a call by its name, inside another macro's arguments too. A value of a repository class handed to any call the graph cannot see, from code that is not a test, reaches its methods.
…ves to A trait's fn with no body was no method, so a call on a bounded type parameter had nothing to resolve to and ran none of the implementations. It is a declaration now, as a Java abstract method is, and the class hierarchy runs every implementation beside it. The resolution rule says so.
A call read from a macro's tokens, tri!(T::deserialize(&mut de)), had no arguments, so what it handed over reached nothing. Each top-level argument that is a name is recorded as the local it is.
There was a problem hiding this comment.
Perch found 2 issues in code this pull request changed. See the scan
| return !target && /^[A-Z]/.test(name) && !name.includes('.') && classMethods(file, name).length ? { name, path: file.path, from } : null; | ||
| }; | ||
| /** What a method makes: the class of what it returns, through a function it returns when it returns one. */ | ||
| const madeBy = (id, depth) => { |
There was a problem hiding this comment.
Lint: test-calls-resolve-by-language (54%)
Perch believes buildGraph.madeBy has this problem. It is a probability, not a located defect: the cause may sit a few lines away.
Check a fix with perch check src/graph.js::buildGraph.madeBy. See the scan
There was a problem hiding this comment.
No change: madeBy resolves through resolve and what a method returns, which the rule allows. At 54% this reads as noise.
| return owner.length > 1 && /^[A-Z]/.test(owner.at(-2)) ? { name: owner.at(-2), full: owner.slice(0, -1).join('.'), path: node.path, from } : null; | ||
| }; | ||
| /** A local that names a class rather than holding a value: `cls = Command`, an instance of Command when it is called. */ | ||
| const classNamed = (file, name, from, depth) => { |
There was a problem hiding this comment.
Lint: test-calls-resolve-by-language (52%)
Perch believes buildGraph.classNamed has this problem. It is a probability, not a located defect: the cause may sit a few lines away.
Check a fix with perch check src/graph.js::buildGraph.classNamed. See the scan
There was a problem hiding this comment.
No change: classNamed falls back only to a class declared in the same file, which the rule allows. At 52% this reads as noise.
|
Perch Scan: 24 issues in 884 methods at |
… call graph A mutant was asked of the eight tests nearest the method by call depth and called survived when those missed it. On click and gson a thousand tests reach a core method, and the nearest eight were generic ones that never check it. Every test that may run a method is now asked, most likely first: those calling it through written calls, and those reaching it through an interface or a handed value whose code names its method, class or file. They go a batch to a request until one kills the mutant, and a survivor has been asked of all of them. Checks-nothing and duplicates are judged on first-batch answers, and duplicates only within one test file over three mutants or more.
… already settled Once a test was likely enough to fail that a mutant could no longer be listed as survived, or the edit changed nothing a caller could observe, every remaining test was still asked: on gson that was more than half of 37,000 requests and changed no listed finding. Asking stops there. An equivalent mutant is left out of the score, as mutation testing defines it, and one some test is likely but not sure to kill is counted as undecided. Tests are matched to a method by their file name and code, not their directory, and a word most tests name no longer counts.
…only of the tests that ran it The JUnit, LCOV, Cobertura, JaCoCo and coverage.py readers and the test-name matching for Python, JavaScript, the JVM, C and C++ and Rust are back, with --junit, --lcov, --cobertura, --jacoco, --contexts and coverage_reports in perch.yaml. What a run measured replaces the call graph's guess: a mutant whose statements no test ran has no coverage and is not asked, and with per-test coverage a mutant is asked only of the tests that ran its statements. Each mutant carries the first lines of the statements it edits, which is where coverage tools record them. --contexts also reads coverage.py's .coverage data file, which says test by test what its JSON says in gigabytes.
There was a problem hiding this comment.
Perch found 7 issues in code this pull request changed. See the scan
…out survivors perch runs pytest itself, in copies of the repository at the commit: once over the suite with per-test coverage, to know which tests run each line, then once per mutant over exactly those tests. A mutant a test fails on is caught, one that runs past three times the tests' own time is caught by its timeout, one the tests cannot be collected against is invalid and left out of the score, and one no test ran has no coverage. The model is asked nothing about any of those; it is asked only whether a survivor is a change a caller could observe. Checks-nothing and duplicate tests come from the real results. A framework perch cannot run is still estimated from the call graph, and says so. The report-file flags and coverage_reports are gone: perch does not read coverage reports lying around.
…ver run coverage.py records a multi-line condition by the lines it executes, not the if's first line, and a default value runs at import, under no test. A mutant now counts as run when any of its own lines or its statement's first line ran, and one in a function's signature is run by every test that runs the function. The docs say perch runs the tests.
A run perch starts is now held to its copy, its scratch directory and the temporary directory: sandbox-exec on macOS, bubblewrap on Linux when installed, and a logged warning without either. Whatever a run started is killed when it ends, not only when it times out. During a click run, files of the original checkout were deleted after perch had exited, by a process a mutated test left behind.
There was a problem hiding this comment.
Perch found 4 issues in code this pull request changed. See the scan
perch coverage no longer falls back to predicting kills from the call graph. A repository whose tests it cannot run is an error naming what is missing. The kills question is gone; the model is asked only whether a survivor matters. A survivor the model could not be asked about is still listed, with no probability. --verbose names the pytest config it read. CI installs pytest and pytest-cov so the tests that run pytest for real run there.
There was a problem hiding this comment.
Perch found 3 issues in code this pull request changed. See the scan
…CI might leave perch runs the tests itself, so the readers for LCOV, Cobertura, JaCoCo and coverage.py's JSON, the merging of shards and report kinds, and the matching of report test names across five language families are gone. A runner now returns the lines each test ran by perch's test id; pytest's maps its node ids itself, from coverage.py's data file. A test the run says ran code in scope is kept whatever the graph found, and a test the runner ran that perch has no test for is logged and counted. Test judging is exact: no batch rounds, no probabilities from guesses. Invalid mutants leave each method's score as they leave the totals. The report page no longer says predicted.
There was a problem hiding this comment.
Perch found 7 issues in code this pull request changed. See the scan
… and runs mutants without restarting pytest collects the suite once and runs each mutant in a forked child with the mutated function swapped in. JavaScript and TypeScript get mutant schemata: every mutant written into the code once behind a switch, and per-test coverage recorded by the switches, exact to the mutant. Each run stops at its first failing test, as Stryker's does; tests that killed nothing and tests that run the same mutants get only the extra runs it takes to settle them. Timeouts follow Stryker's 1.5x plus five seconds. Each session is sandboxed once, not each mutant: a long-lived server or command driver runs everything under it. Outcomes are saved and reused while a mutant's code, its tests, what they run and the manifests are unchanged; --since runs only the changed code's mutants and those its changed tests reach. Mutants inside types are no longer made, nor an emptied expression body.
| export function exec(command, args, { cwd, env = {}, timeout = 0, writable = null } = {}) { | ||
| const run = writable ? sandboxed(command, args, writable) : { command, args }; | ||
| return new Promise(resolve => { | ||
| const child = spawn(run.command, run.args, { cwd, env: { ...process.env, PWD: cwd, ...env }, stdio: ['ignore', 'pipe', 'pipe'], detached: true }); |
There was a problem hiding this comment.
spawn gets an argument array, not a shell string, and the sandbox profile takes its paths as separate -D parameters, so nothing here goes through a shell.
There was a problem hiding this comment.
Perch found 25 issues in code this pull request changed. See the scan
| * Commands come in on file descriptor 3 and results go out on 4, a JSON object a line. A child writes a line per test as it | ||
| * finishes, so a child that dies mid-run names the test it died in. | ||
| */ | ||
| export const SERVER = String.raw` |
There was a problem hiding this comment.
P2 Security: code injection (75%)
Perch believes <top-level> has this problem. It is a probability, not a located defect: the cause may sit a few lines away.
Check a fix with perch check src/runners/pytest-server.js::<top-level>. See the scan
There was a problem hiding this comment.
By design: the fork server compiles the mutated function perch sent it, inside the sandbox. Nothing outside perch supplies that source.
Rust and C# get mutant schemata, built once: a mutant the compiler rejects is traced by its error's line and column to the edit there, taken out, and the code built again, and is invalid. libtest's thread names and a stack walk to the test attribute in .NET, through an async test's state machine, say which test reached each switch. Go has no expression that chooses between two values of any type, so each Go mutant is written into a copy of the module per worker and rebuilt from Go's build cache, with per-test coverage from each test and subtest run alone. .NET scope follows a test project's ProjectReferences and reads .slnx.
There was a problem hiding this comment.
Perch found 11 issues in code this pull request changed. See the scan
| * The mutants an error at `line` and `column` of an instrumented file falls in: the edit around that place, innermost first, else | ||
| * the choice around it. Without a column, anything on the line. | ||
| */ | ||
| function locate(text, line, column) { |
There was a problem hiding this comment.
P1 Defect: off by one (67%)
Perch believes locate has this problem. It is a probability, not a located defect: the cause may sit a few lines away.
Check a fix with perch check src/runners/schemata.js::locate. See the scan
There was a problem hiding this comment.
No change: every compiler perch reads gives 1-based columns, and the overlap test is right.
Each worker loads the framework once, sandboxed once, and runs every mutant in-process; one that hangs is killed and replaced. Vitest's caches go to scratch, and a framework error stops the run rather than reading as a result. Mocha keeps the flags the test script gives it. A run a mutant crashes before its tests report is a kill. Computed and .each test titles, and tests a helper or a conditional suite declares, are matched to perch's tests by template and by their name's ending. Test groups are parted on mutants one of them kills before every outcome is filled in. The command driver answers on its own descriptor and escapes Unicode line separators.
There was a problem hiding this comment.
Perch found 6 issues in code this pull request changed. See the scan
perch runs the tests itself, so the reports/ directories, the scripts that regenerated them, and the jest-junit and mocha-junit-reporter settings in the fixtures had nothing left to read them.
… sbt Mutant schemata for Java, Kotlin and Scala, built once by the project's own build; a mutant the compiler rejects is traced by the error's line and column and taken out. A JVM per core runs each mutant through JUnit 5, JUnit 4, Parameterized and MUnit included, TestNG, or ScalaTest, with the project's classes loaded afresh for each so no state carries between mutants. A for loop's counter is switched as a compound assignment, and a + that joins strings is no longer turned into a -. A mutant the compiler rejects is invalid whether or not a test reaches it. JavaScript: what a suite's hooks and a test file's loading reach is put down to that file's tests, and Mocha's worker loads the repository's modules afresh for each mutant.
There was a problem hiding this comment.
Perch found 3 issues in code this pull request changed. See the scan
Karma runs in headless Chrome from a config perch wraps around the project's: its hooks load first, and what they record comes back as marked console lines, the mutant and the name filter going in as client arguments. The sources are the files the Karma config loads. Cucumber's scenarios are read from the feature files and are the tests, each named in Cucumber's own Before hook. A run that reported its tests skipped is no crash; only one that reported nothing is.
There was a problem hiding this comment.
Perch found 5 issues in code this pull request changed. See the scan
| }, | ||
|
|
||
| /** The suite once, every switch recording which test reached it: what each test ran, by mutant and by line, and each result. */ | ||
| async coverageRun({ copy, scratch, tool }) { |
There was a problem hiding this comment.
P1 Defect: unhandled null (62%)
Perch believes javascriptRunner.runner.coverageRun has this problem. It is a probability, not a located defect: the cause may sit a few lines away.
Check a fix with perch check src/runners/javascript.js::javascriptRunner.runner.coverageRun. See the scan
There was a problem hiding this comment.
No change: every id a switch reports was assigned in this copy's instrumentation, so keys.get cannot miss.
With a worker per core, tests that open real connections can run the machine out of local ports, and a test fails for the machine rather than the mutant. A failing test now runs again by itself against the same mutant before it counts. A warm worker whose run leaves a server, socket or timer open is replaced, so the next mutant does not inherit it. A nested function's mutants were also made again by the function around it, so each was counted and run twice; each now belongs to the innermost method. A command that left a child holding its output open was waited on until the child ended; its process group now goes when it exits.
There was a problem hiding this comment.
Perch found an issue in code this pull request changed. See the scan
… wrong tests A repository with Cucumber and Mocha got two JavaScript runners on one copy: the second wrote its switches over the first's, its hooks file replaced the first's, every JavaScript mutant went to the first runner, and the scope left the feature files out. The copy is now instrumented once and shared, each runner's files are named for its framework, each mutant's tests go to the runner that ran them, and Cucumber's feature files and step definitions are in the scope. The cargo, go, dotnet and JVM runners kept a run's state in the module, so two runs in one process overwrote each other; each is now made per run. Karma's file patterns are read from basePath, and its framework count no longer includes source files. Catches that swallowed every error now ignore only the one they expect, the worker pool closes busy workers too, the driver stops a run that prints past 256 MB, and the outcome cache is written through the store. The repository's coverage rules said perch never runs tests; they now say kills come from runs, and a run-measured finding may carry 1.
The repository allows only GitHub's own actions, so the workflow with sbt/setup-sbt failed before any job ran. sbt now installs from its own apt repository.
There was a problem hiding this comment.
Perch found 11 issues in code this pull request changed. See the scan
| }; | ||
|
|
||
| /** A cargo runner: what preparing one copy found is its own, so two runs in one process keep theirs apart. */ | ||
| export function cargoRunner() { |
There was a problem hiding this comment.
P1 Defect: wrong return value (89%)
Perch believes cargoRunner has this problem. It is a probability, not a located defect: the cause may sit a few lines away.
Check a fix with perch check src/runners/cargo.js::cargoRunner. See the scan
There was a problem hiding this comment.
No change: the factory returns every property perch reads from a runner (name, languages, copiesFor, available, prepare, coverageRun, session), each bound to this run's state.
| const copiesFor = () => 1; | ||
|
|
||
| /** A JVM runner: what preparing one copy found is its own, so two runs in one process keep theirs apart. */ | ||
| export function jvmRunner() { |
There was a problem hiding this comment.
P1 Defect: wrong return value (88%)
Perch believes jvmRunner has this problem. It is a probability, not a located defect: the cause may sit a few lines away.
Check a fix with perch check src/runners/jvm.js::jvmRunner. See the scan
There was a problem hiding this comment.
No change: the factory returns every property perch reads from a runner (name, languages, copiesFor, available, prepare, coverageRun, session), each bound to this run's state.
| `; | ||
|
|
||
| /** A dotnet runner: what preparing one copy found is its own, so two runs in one process keep theirs apart. */ | ||
| export function dotnetRunner() { |
There was a problem hiding this comment.
P1 Defect: wrong return value (87%)
Perch believes dotnetRunner has this problem. It is a probability, not a located defect: the cause may sit a few lines away.
Check a fix with perch check src/runners/dotnet.js::dotnetRunner. See the scan
There was a problem hiding this comment.
No change: the factory returns every property perch reads from a runner (name, languages, copiesFor, available, prepare, coverageRun, session), each bound to this run's state.
| const plain = test => test.replace(/#\d+(?=\/|$)/g, ''); | ||
|
|
||
| /** A go runner: what one run found is its own, so two runs in one process keep theirs apart. */ | ||
| export function goRunner() { |
There was a problem hiding this comment.
P1 Defect: wrong return value (86%)
Perch believes goRunner has this problem. It is a probability, not a located defect: the cause may sit a few lines away.
Check a fix with perch check src/runners/go.js::goRunner. See the scan
There was a problem hiding this comment.
No change: the factory returns every property perch reads from a runner (name, languages, copiesFor, available, prepare, coverageRun, session), each bound to this run's state.
| * The copy made ready: each crate root given the switch, every mutant written into its file, and the tests built, again without | ||
| * whatever the compiler rejects, until they build. | ||
| */ | ||
| async function prepare(state, { copies: [copy], generated, graph }) { |
There was a problem hiding this comment.
P1 Defect: wrong lookup (68%)
Perch believes prepare has this problem. It is a probability, not a located defect: the cause may sit a few lines away.
Check a fix with perch check src/runners/cargo.js::prepare. See the scan
There was a problem hiding this comment.
No change: an integration test outside every crate root's directory is its own crate, so an empty module path is the name libtest gives it.
| } | ||
|
|
||
| /** Each built test binary once, every switch recording which test's thread reached it. */ | ||
| async function coverageRun({ names, keys, unplaced, binaries, testOfNode }, { copy, scratch }) { |
There was a problem hiding this comment.
P1 Defect: unhandled null (66%)
Perch believes coverageRun has this problem. It is a probability, not a located defect: the cause may sit a few lines away.
Check a fix with perch check src/runners/cargo.js::coverageRun. See the scan
There was a problem hiding this comment.
No change: every id in the hits file was written by this copy's switches, so keys.get cannot miss.
| * A Karma config, loaded with a stand-in for Karma's own config object so its `files` are what it set: the patterns of every | ||
| * file the browser loads, the code under test and the tests alike. Printed on the last line, after a marker. | ||
| */ | ||
| async function loadKarma({ root, config, node, paths }) { |
There was a problem hiding this comment.
P1 Defect: wrong lookup (66%)
Perch believes loadKarma has this problem. It is a probability, not a located defect: the cause may sit a few lines away.
Check a fix with perch check src/test-scope.js::loadKarma. See the scan
There was a problem hiding this comment.
No change: the stand-in config object answers an unknown key, LOG_INFO and the like, with its own name, which is all a config reads from it.
| } | ||
|
|
||
| /** The project made ready: the switch in each module with mutants, every mutant written in, and the tests built until they build. */ | ||
| async function prepare(state, { copies: [copy], generated, graph, tool }) { |
There was a problem hiding this comment.
P1 Defect: unhandled null (61%)
Perch believes prepare has this problem. It is a probability, not a located defect: the cause may sit a few lines away.
Check a fix with perch check src/runners/jvm.js::prepare. See the scan
There was a problem hiding this comment.
No change: tool here is what this runner's available returned, which always carries build.
| } | ||
| throw new IncompleteCheckError(`method ${node.qualified_name} does not fit the ${budget}-token budget`); | ||
| }; | ||
| return (await answer({ subject: 'method', node, build, typed: () => compile(mattersAsked), questions: mattersAsked })).answers.matters; |
There was a problem hiding this comment.
Lint: coverage-kills-come-from-runs (56%)
Perch believes triage.<anonymous>.typed has this problem. It is a probability, not a located defect: the cause may sit a few lines away.
Check a fix with perch check src/coverage.js::triage.<anonymous>.typed. See the scan
There was a problem hiding this comment.
No change: askCoverage asks whether a survivor matters; it never decides whether a mutant was killed.
go list -json prints its packages one after another, and they were split with a pattern on closing and opening braces, which also matched inside a string such as a package's doc line. They are now split where the braces close outside a string.
…o Go on a fresh machine The copy links in the repository's node_modules, which were found with find -type d and so were missed when node_modules is itself a link, as a monorepo tool or a shared install leaves it: the copy had none, and the framework could not load. A sandboxed run also resolved every directory it may write to, and Go's module cache does not exist until Go first writes it; such a directory is now made first.
There was a problem hiding this comment.
Perch found an issue in code this pull request changed. See the scan
…ailed to start A fork server that exited while it collected the suite rejected the session without killing the process group pytest had started. The group is now killed before the error goes back.
perch coverageasked a model to predict which tests would fail against each mutant, from a call graph's guess at which tests run it. Both halves were guesses. On click, gson and serde_json the graph missed or invented reach in every language, and the predictions cost up to $54 a run and were wrong four times in ten about what survived.perch now runs the tests itself, as Stryker and PIT do, and the model judges only what the run leaves.
--sinceruns only the mutants of the code a branch changed.Runners: pytest; Vitest, Jest, Mocha, Jasmine, Karma, Cucumber and node:test; cargo; go test; dotnet test with xUnit, NUnit or MSTest; JUnit 5, JUnit 4, TestNG, ScalaTest and MUnit on Gradle, Maven or sbt. A repository whose tests perch cannot run is an error that says what is missing. Nothing is estimated.
On click, the run takes 124 seconds on 24 workers, against 12 minutes before, for a 70.6% mutation score, and the model's part costs about a cent.
Closes #344, closes #346, closes #349, closes #351, closes #352, closes #353, closes #354