Skip to content

Unrot the flawd mutation-testing config - #35

Merged
fohara merged 1 commit into
mainfrom
fix/flawd-config-rot
Aug 9, 2026
Merged

fohara merged 1 commit into
mainfrom
fix/flawd-config-rot

Conversation

@fohara

@fohara fohara commented Aug 9, 2026

Copy link
Copy Markdown
Contributor

Two problems, both of which stopped flawd run dead on a fresh clone. Found
while running Flawd v0.7.0 against this repo on 2026-08-08.

The [llm] table

Our committed flawd.toml still carried an [llm] table (mode = "off",
max_edit_chars = 120). Flawd removed LLM semantic operators in v0.7.0 and now
treats a leftover [llm] as a hard config error, so flawd run aborted before
doing any work. The config had not been exercised since the last run in March,
so this rotted unnoticed.

Deleting those three lines is the whole fix. With [llm] gone the rest
validates clean against v0.7.0 (flawd config show, no other stale keys).

The coverage command

[coverage] command set LLVM_COV and LLVM_PROFDATA to /opt/homebrew/...,
macOS Homebrew paths that cannot resolve inside a Linux container. The
Dockerfile also never installed cargo-llvm-cov. Flawd defaults to docker
isolation, so the documented invocation built the image, failed coverage with
cargo exit 101, and wrote nothing.

The Dockerfile now installs cargo-llvm-cov from its prebuilt release,
selected by uname -m so it works on both arm64 and x86_64 hosts, plus the
llvm-tools-preview component that supplies llvm-cov and llvm-profdata.
rust-toolchain.toml is copied before the component is added, so it lands on
the toolchain the build actually uses rather than the image's default. With the
tools present the command needs no path overrides, matching what the CI
coverage job already does.

The command also creates coverage/ first. That directory is gitignored, so it
exists on a developer's machine but never in a fresh container, and llvm-cov
does not create its own output directory.

Verification

flawd run under default (docker) isolation, with a small budget:

Baseline: passed
Per-test targeted: 6/6 (100.0%)
Full-suite fallback: 0/6 (0.0%)
Per-test coverage: 3312 tests loaded
Killed: 6 (100.00%)

Per-test collection succeeds rather than falling back, which is the behaviour
the config was supposed to enable.

Not done here

The ticket also floated a CI job running flawd config show so version drift in
our own tooling gets caught by the build rather than by a human months later.
That needs flawd installed on the runner, which is a bigger change than this
fix; leaving it for a follow-up.

Two problems, both of which stopped `flawd run` dead on a fresh clone.

The committed `flawd.toml` still carried an `[llm]` table. Flawd removed
LLM semantic operators in v0.7.0 and now treats a leftover `[llm]` as a
hard config error, so the command aborted before doing any work. Nothing
had exercised the config since March, so this rotted unnoticed.

The coverage command was also unrunnable. It set LLVM_COV and
LLVM_PROFDATA to `/opt/homebrew/...`, which are macOS paths that cannot
resolve inside a Linux container, and the Dockerfile never installed
`cargo-llvm-cov` in the first place. Since flawd defaults to docker
isolation, the documented invocation built the image, failed coverage
with cargo exit 101, and wrote nothing.

The Dockerfile now installs `cargo-llvm-cov` from its prebuilt release
(selected by `uname -m`, so this works on both arm64 and x86_64 hosts)
plus the `llvm-tools-preview` component, which supplies llvm-cov and
llvm-profdata. rust-toolchain.toml is copied before the component is
added so it lands on the toolchain the build actually uses rather than
the image's default. With the tools present the command needs no path
overrides at all, matching what CI already does.

The command also creates `coverage/` first. The directory is gitignored,
so it exists on a developer's machine but never in a fresh container, and
llvm-cov will not create the output directory itself.

Verified with `flawd run` under default (docker) isolation: the image
builds, coverage is collected in-container, and per-test targeting
reports 6/6 mutants targeted with no full-suite fallback.
@fohara
fohara merged commit 3d9d62c into main Aug 9, 2026
21 checks passed
@fohara
fohara deleted the fix/flawd-config-rot branch August 9, 2026 14:40
fohara added a commit that referenced this pull request Aug 9, 2026
Two bugs, both of which made `lash format` unsafe to run on a file it did
not fully model.

`format_file` rebuilt the file from the parsed model alone, and the model
holds the header, the Description section and the task tree and nothing
else. Every other section was silently deleted: a file with `## Notes`
and `## References` came back with only the header and the tasks, exit
code 0, no warning. This turned out to be broader than filed. A section
*above* `## Tasks` went too, because the parser folds it into "overview"
text that `TaskFile` never stored and the formatter never emitted.

`format_file` now takes the source alongside the parsed file. It
regenerates only the spans it owns — the H1 plus annotation block, the
Description section, the Tasks section — and copies every other line
through unchanged. Section boundaries come from
`parser::header::section_span`, which runs through pulldown-cmark, so a
`##` inside a fenced code block neither opens nor closes a section.
Passing an empty source keeps the old model-only behaviour, which is what
the existing unit tests want.

The second bug: the parser records inline labels in metadata without
removing them from the title, and the formatter wrote both. Every run
appended another copy of every label, so `- [ ] one #docs` became
`#docs #docs` and then `#docs #docs #docs`, and `format --check` reported
the file as needing formatting forever. The formatter strips inline
labels from the title and writes the sorted metadata list as the only
place they appear. The parsed title keeps them, so `lash list` and search
behave as they do today.

Each generated block ends with its own blank separator and the source's
is skipped rather than copied. Emitting both would add a blank line per
run, which is the same non-idempotence in a different coat.

Tests: sections above and below `## Tasks` survive and keep their order,
a fenced heading does not split a section, labels are emitted once, an
already-canonical file comes back byte-identical, and an idempotence
property test over six shapes asserting `format(format(x)) ==
format(x)`. All six fail against the old behaviour.

Also ticks off three items in lash.index.md that were fixed in #34 and
#35 but never marked done.
fohara added a commit that referenced this pull request Aug 9, 2026
Two bugs, both of which made `lash format` unsafe to run on a file it did
not fully model.

`format_file` rebuilt the file from the parsed model alone, and the model
holds the header, the Description section and the task tree and nothing
else. Every other section was silently deleted: a file with `## Notes`
and `## References` came back with only the header and the tasks, exit
code 0, no warning. This turned out to be broader than filed. A section
*above* `## Tasks` went too, because the parser folds it into "overview"
text that `TaskFile` never stored and the formatter never emitted.

`format_file` now takes the source alongside the parsed file. It
regenerates only the spans it owns — the H1 plus annotation block, the
Description section, the Tasks section — and copies every other line
through unchanged. Section boundaries come from
`parser::header::section_span`, which runs through pulldown-cmark, so a
`##` inside a fenced code block neither opens nor closes a section.
Passing an empty source keeps the old model-only behaviour, which is what
the existing unit tests want.

The second bug: the parser records inline labels in metadata without
removing them from the title, and the formatter wrote both. Every run
appended another copy of every label, so `- [ ] one #docs` became
`#docs #docs` and then `#docs #docs #docs`, and `format --check` reported
the file as needing formatting forever. The formatter strips inline
labels from the title and writes the sorted metadata list as the only
place they appear. The parsed title keeps them, so `lash list` and search
behave as they do today.

Each generated block ends with its own blank separator and the source's
is skipped rather than copied. Emitting both would add a blank line per
run, which is the same non-idempotence in a different coat.

Tests: sections above and below `## Tasks` survive and keep their order,
a fenced heading does not split a section, labels are emitted once, an
already-canonical file comes back byte-identical, and an idempotence
property test over six shapes asserting `format(format(x)) ==
format(x)`. All six fail against the old behaviour.

Also ticks off three items in lash.index.md that were fixed in #34 and
#35 but never marked done.
@fohara fohara mentioned this pull request Aug 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant