ci: scheduled audits use the Docker backend — native under-kills subprocess-tested code - #69
Merged
Merged
Conversation
…rocess-tested code The first scheduled audit scored 77.5% vs the Docker-verified 98.7% baseline; forensics show 72 of its 88 survivors are Docker-verified kills, clustered in the crates whose tests spawn the built lash binary. Reproduced locally in flawd's native mode and filed upstream (tasks.native-subprocess-false-survivors); that audit's score is invalidated for threshold-setting. Score history must be trustworthy, so the weekly audit switches to the Docker config; PR runs stay native for speed while the gate is report-only, with the caveat documented.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The first scheduled audit (33077879750) scored 77.5% vs the Docker-verified local baseline of 98.7%. Forensics: 72 of its 88 survivors are Docker-verified kills, 66 of them in lash-cli/lash-tui — the crates whose tests spawn the built
lashbinary. Reproduced in flawd's native mode locally (no runner, no cache): a flawd defect, filed upstream with acceptance criteria.Until that fix lands: the weekly audit (which feeds the future gate threshold) runs on the Docker backend for trustworthy numbers; PR runs stay native for fast feedback while the gate is report-only, with the known caveat that survivors in subprocess-tested code may be false. That first audit's score is invalidated for threshold purposes.