Pre-flight checks
What happened?
When every doc, paper and image in the corpus hits the semantic cache, the generated skill tells the agent "If all files are cached, skip to Part C directly". Part C reads graphify-out/.graphify_semantic.json unconditionally, but the only Part B step that writes that file is the Step B3 merge. The second /graphify run on an unchanged mixed corpus (code plus docs) therefore fails in Part C with FileNotFoundError.
The obvious workaround is to write an empty .graphify_semantic.json, as the code-only fast path does. That gets past Part C but silently drops every cached semantic node, because B3 is also the step that merges .graphify_cached.json in.
Expected: Part C runs, and the extraction contains the cached nodes.
Where it lives:
tools/skillgen/fragments/core/core.md, the sentence right after the Step B0 block.
- Rendered into every split host (
graphify/skill.md, skill-codex.md, skill-windows.md and the rest).
- The monolith hosts
skill-aider.md and skill-devin.md carry the same sentence, but --monolith-roundtrip pins their bodies to the v8 baseline.
Related gap: stale chunk files. B3 merges every graphify-out/.graphify_chunk_*.json it finds, and chunks are only removed by the final cleanup step. A run that is interrupted after dispatch leaves its chunks behind, and the next run merges them as if they were fresh. That includes an all-cached run, once it goes through B3.
Proposed fix (PR to follow): the change touches only the source fragment, with artifacts regenerated by python -m tools.skillgen --bless.
-
Replace "skip to Part C directly" with: skip Steps B1 and B2 but still run Step B3's commands, and do not write an empty .graphify_semantic.json.
-
In the Step B0 block, delete any .graphify_chunk_*.json before dispatch. Nothing has been dispatched at that point, so any chunk on disk is a leftover.
-
Add regression tests:
- a text check that every split host routes the all-cached case through B3;
- a behavioural test that runs the rendered B0, B3 and Part C Python blocks against an all-cached corpus with a stale chunk on disk.
Both tests fail on current v8 and pass with the fix.
Out of scope: the aider and devin monoliths. Changing them needs a sanctioned round-trip change, which can be a follow-up.
Steps to reproduce
# Corpus with at least one document, e.g. notes.md
/graphify ./corpus # first run: notes.md is extracted and cached
/graphify ./corpus # second run, nothing changed
# Step B0 prints: Cache: 1 files hit, 0 files need extraction
# Runbook says: "If all files are cached, skip to Part C directly"
# Part C reads graphify-out/.graphify_semantic.json -> FileNotFoundError
Reproduced without an agent by running the Python bodies of the rendered graphify/skill.md (v8 at 35adf43) in order: Step B0, then Part C.
Error output or graph output
Cache: 1 files hit, 0 files need extraction
FileNotFoundError: [Errno 2] No such file or directory: 'graphify-out/.graphify_semantic.json'
Graphify version
0.9.76 (v8 at 35adf43)
Operating System
macOS
Python Version
3.12
Pre-flight checks
What happened?
When every doc, paper and image in the corpus hits the semantic cache, the generated skill tells the agent "If all files are cached, skip to Part C directly". Part C reads
graphify-out/.graphify_semantic.jsonunconditionally, but the only Part B step that writes that file is the Step B3 merge. The second/graphifyrun on an unchanged mixed corpus (code plus docs) therefore fails in Part C withFileNotFoundError.The obvious workaround is to write an empty
.graphify_semantic.json, as the code-only fast path does. That gets past Part C but silently drops every cached semantic node, because B3 is also the step that merges.graphify_cached.jsonin.Expected: Part C runs, and the extraction contains the cached nodes.
Where it lives:
tools/skillgen/fragments/core/core.md, the sentence right after the Step B0 block.graphify/skill.md,skill-codex.md,skill-windows.mdand the rest).skill-aider.mdandskill-devin.mdcarry the same sentence, but--monolith-roundtrippins their bodies to the v8 baseline.Related gap: stale chunk files. B3 merges every
graphify-out/.graphify_chunk_*.jsonit finds, and chunks are only removed by the final cleanup step. A run that is interrupted after dispatch leaves its chunks behind, and the next run merges them as if they were fresh. That includes an all-cached run, once it goes through B3.Proposed fix (PR to follow): the change touches only the source fragment, with artifacts regenerated by
python -m tools.skillgen --bless.Replace "skip to Part C directly" with: skip Steps B1 and B2 but still run Step B3's commands, and do not write an empty
.graphify_semantic.json.In the Step B0 block, delete any
.graphify_chunk_*.jsonbefore dispatch. Nothing has been dispatched at that point, so any chunk on disk is a leftover.Add regression tests:
Both tests fail on current
v8and pass with the fix.Out of scope: the aider and devin monoliths. Changing them needs a sanctioned round-trip change, which can be a follow-up.
Steps to reproduce
Reproduced without an agent by running the Python bodies of the rendered
graphify/skill.md(v8at35adf43) in order: Step B0, then Part C.Error output or graph output
Graphify version
0.9.76 (
v8at35adf43)Operating System
macOS
Python Version
3.12