Unrot the flawd mutation-testing config - #35
Merged
Merged
Conversation
Two problems, both of which stopped `flawd run` dead on a fresh clone. The committed `flawd.toml` still carried an `[llm]` table. Flawd removed LLM semantic operators in v0.7.0 and now treats a leftover `[llm]` as a hard config error, so the command aborted before doing any work. Nothing had exercised the config since March, so this rotted unnoticed. The coverage command was also unrunnable. It set LLVM_COV and LLVM_PROFDATA to `/opt/homebrew/...`, which are macOS paths that cannot resolve inside a Linux container, and the Dockerfile never installed `cargo-llvm-cov` in the first place. Since flawd defaults to docker isolation, the documented invocation built the image, failed coverage with cargo exit 101, and wrote nothing. The Dockerfile now installs `cargo-llvm-cov` from its prebuilt release (selected by `uname -m`, so this works on both arm64 and x86_64 hosts) plus the `llvm-tools-preview` component, which supplies llvm-cov and llvm-profdata. rust-toolchain.toml is copied before the component is added so it lands on the toolchain the build actually uses rather than the image's default. With the tools present the command needs no path overrides at all, matching what CI already does. The command also creates `coverage/` first. The directory is gitignored, so it exists on a developer's machine but never in a fresh container, and llvm-cov will not create the output directory itself. Verified with `flawd run` under default (docker) isolation: the image builds, coverage is collected in-container, and per-test targeting reports 6/6 mutants targeted with no full-suite fallback.
This was referenced Aug 9, 2026
fohara
added a commit
that referenced
this pull request
Aug 9, 2026
Two bugs, both of which made `lash format` unsafe to run on a file it did not fully model. `format_file` rebuilt the file from the parsed model alone, and the model holds the header, the Description section and the task tree and nothing else. Every other section was silently deleted: a file with `## Notes` and `## References` came back with only the header and the tasks, exit code 0, no warning. This turned out to be broader than filed. A section *above* `## Tasks` went too, because the parser folds it into "overview" text that `TaskFile` never stored and the formatter never emitted. `format_file` now takes the source alongside the parsed file. It regenerates only the spans it owns — the H1 plus annotation block, the Description section, the Tasks section — and copies every other line through unchanged. Section boundaries come from `parser::header::section_span`, which runs through pulldown-cmark, so a `##` inside a fenced code block neither opens nor closes a section. Passing an empty source keeps the old model-only behaviour, which is what the existing unit tests want. The second bug: the parser records inline labels in metadata without removing them from the title, and the formatter wrote both. Every run appended another copy of every label, so `- [ ] one #docs` became `#docs #docs` and then `#docs #docs #docs`, and `format --check` reported the file as needing formatting forever. The formatter strips inline labels from the title and writes the sorted metadata list as the only place they appear. The parsed title keeps them, so `lash list` and search behave as they do today. Each generated block ends with its own blank separator and the source's is skipped rather than copied. Emitting both would add a blank line per run, which is the same non-idempotence in a different coat. Tests: sections above and below `## Tasks` survive and keep their order, a fenced heading does not split a section, labels are emitted once, an already-canonical file comes back byte-identical, and an idempotence property test over six shapes asserting `format(format(x)) == format(x)`. All six fail against the old behaviour. Also ticks off three items in lash.index.md that were fixed in #34 and #35 but never marked done.
fohara
added a commit
that referenced
this pull request
Aug 9, 2026
Two bugs, both of which made `lash format` unsafe to run on a file it did not fully model. `format_file` rebuilt the file from the parsed model alone, and the model holds the header, the Description section and the task tree and nothing else. Every other section was silently deleted: a file with `## Notes` and `## References` came back with only the header and the tasks, exit code 0, no warning. This turned out to be broader than filed. A section *above* `## Tasks` went too, because the parser folds it into "overview" text that `TaskFile` never stored and the formatter never emitted. `format_file` now takes the source alongside the parsed file. It regenerates only the spans it owns — the H1 plus annotation block, the Description section, the Tasks section — and copies every other line through unchanged. Section boundaries come from `parser::header::section_span`, which runs through pulldown-cmark, so a `##` inside a fenced code block neither opens nor closes a section. Passing an empty source keeps the old model-only behaviour, which is what the existing unit tests want. The second bug: the parser records inline labels in metadata without removing them from the title, and the formatter wrote both. Every run appended another copy of every label, so `- [ ] one #docs` became `#docs #docs` and then `#docs #docs #docs`, and `format --check` reported the file as needing formatting forever. The formatter strips inline labels from the title and writes the sorted metadata list as the only place they appear. The parsed title keeps them, so `lash list` and search behave as they do today. Each generated block ends with its own blank separator and the source's is skipped rather than copied. Emitting both would add a blank line per run, which is the same non-idempotence in a different coat. Tests: sections above and below `## Tasks` survive and keep their order, a fenced heading does not split a section, labels are emitted once, an already-canonical file comes back byte-identical, and an idempotence property test over six shapes asserting `format(format(x)) == format(x)`. All six fail against the old behaviour. Also ticks off three items in lash.index.md that were fixed in #34 and #35 but never marked done.
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two problems, both of which stopped
flawd rundead on a fresh clone. Foundwhile running Flawd v0.7.0 against this repo on 2026-08-08.
The
[llm]tableOur committed
flawd.tomlstill carried an[llm]table (mode = "off",max_edit_chars = 120). Flawd removed LLM semantic operators in v0.7.0 and nowtreats a leftover
[llm]as a hard config error, soflawd runaborted beforedoing any work. The config had not been exercised since the last run in March,
so this rotted unnoticed.
Deleting those three lines is the whole fix. With
[llm]gone the restvalidates clean against v0.7.0 (
flawd config show, no other stale keys).The coverage command
[coverage] commandsetLLVM_COVandLLVM_PROFDATAto/opt/homebrew/...,macOS Homebrew paths that cannot resolve inside a Linux container. The
Dockerfile also never installed
cargo-llvm-cov. Flawd defaults to dockerisolation, so the documented invocation built the image, failed coverage with
cargo exit 101, and wrote nothing.
The Dockerfile now installs
cargo-llvm-covfrom its prebuilt release,selected by
uname -mso it works on both arm64 and x86_64 hosts, plus thellvm-tools-previewcomponent that supplies llvm-cov and llvm-profdata.rust-toolchain.tomlis copied before the component is added, so it lands onthe toolchain the build actually uses rather than the image's default. With the
tools present the command needs no path overrides, matching what the CI
coverage job already does.
The command also creates
coverage/first. That directory is gitignored, so itexists on a developer's machine but never in a fresh container, and llvm-cov
does not create its own output directory.
Verification
flawd rununder default (docker) isolation, with a small budget:Per-test collection succeeds rather than falling back, which is the behaviour
the config was supposed to enable.
Not done here
The ticket also floated a CI job running
flawd config showso version drift inour own tooling gets caught by the build rather than by a human months later.
That needs flawd installed on the runner, which is a bigger change than this
fix; leaving it for a follow-up.