fix(skill): run Step B3 when every semantic file is cached - #4117
brunovima83 wants to merge 1 commit into
Conversation
The split-host runbook told the agent to skip to Part C when every doc, paper and image hit the semantic cache. Part C reads .graphify_semantic.json unconditionally and only the Step B3 merge writes it, so a second run on an unchanged mixed corpus failed with FileNotFoundError; writing an empty file instead dropped every cached node. The all-cached case now skips B1/B2 but still runs B3. Step B0 also clears .graphify_chunk_*.json before dispatch: B3 merges every chunk on disk, and a leftover from an interrupted run would otherwise be merged as fresh. The aider and devin monoliths keep the old sentence; they are pinned to the v8 baseline by --monolith-roundtrip. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Thanks for the pull request, @brunovima83. A maintainer will review it soon. Want to talk it through while it is in review? Come join us on our Discord server. For longer-form discussion there is also GitHub Discussions. A couple of things that speed up review: make sure the test suite passes on Python 3.10 and 3.13, and that the change keeps extraction deterministic. |
There was a problem hiding this comment.
Graphify reviewed this change.
Looks safe to merge — no coupling regressions and no blocking issues, checked against the code graph (not a self-assessment).
Graphify review — findings
Fixes the all-cached path in the generated skill docs: Part B's cache check now deletes leftover .graphify_chunk_*.json files from interrupted runs before anything is dispatched. Previously Step B3 would merge those stale chunks. When every file is cached, the instructions now skip only Steps B1 and B2 and still run Step B3's merge, because that merge is the only step that writes .graphify_semantic.json, which Part C reads unconditionally. They also warn against writing an empty file instead, since that would drop every cached node.
No blocking issues surfaced. 15 lower-confidence candidates did not survive cross-model review.
Analysis details — impact, health, verification
Impact & health
Graphify review
Impact — 736 functions depend on the 736 functions this change touches.
Health — grade A; no new coupling hotspots.
Verification — 736 functions in the blast radius were not formally verified this run (proofs are advisory here).
Gate & verification
graphify gate
PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.
Advisory (not blocking):
- verification_scope: 736 function(s) in the blast radius were not formally verified this run
Test selection
Test selection
330 of 330 test file(s) selected (100%) via static blast radius.
Escalated to a full run for safety — the selection is not trustworthy on its own (see below). CI should run the whole suite.
tests/test_affected_cli.py— full-run-safetytests/test_affected_member_seed.py— full-run-safetytests/test_agents_platform.py— full-run-safetytests/test_analyze.py— full-run-safetytests/test_anthropic_custom_endpoint.py— full-run-safetytests/test_antigravity_install.py— full-run-safetytests/test_apm_fallback_version.py— full-run-safetytests/test_architecture_doc.py— full-run-safetytests/test_astro_extraction.py— full-run-safetytests/test_astro_import_ids.py— full-run-safetytests/test_atomic_canvas_export.py— full-run-safetytests/test_atomic_version_stamp.py— full-run-safetytests/test_atomic_writes.py— full-run-safetytests/test_backend_env_isolation.py— full-run-safetytests/test_backend_extras.py— full-run-safetytests/test_benchmark.py— full-run-safetytests/test_benchmark_raw_graph.py— full-run-safetytests/test_blade_extractor.py— full-run-safetytests/test_build.py— full-run-safetytests/test_build_merge_dedup_scope.py— full-run-safetytests/test_build_merge_hyperedges_and_prune.py— full-run-safetytests/test_build_merge_shrink_guard.py— full-run-safetytests/test_builtin_global_type_refs.py— full-run-safetytests/test_cache.py— full-run-safetytests/test_callflow_html.py— full-run-safetytests/test_cargo_introspect.py— full-run-safetytests/test_cargo_missing_manifest.py— full-run-safetytests/test_carried_hyperedge_remap.py— full-run-safetytests/test_case_sensitive_resolution.py— full-run-safetytests/test_charmap_encoding.py— full-run-safetytests/test_chunking.py— full-run-safetytests/test_cjs_module_extension.py— full-run-safetytests/test_claude_cli_backend.py— full-run-safetytests/test_claude_md.py— full-run-safetytests/test_cli_broken_pipe.py— full-run-safetytests/test_cli_export.py— full-run-safetytests/test_cli_help.py— full-run-safetytests/test_cluster.py— full-run-safetytests/test_cluster_exclude_hubs.py— full-run-safetytests/test_cobol_extractor.py— full-run-safetytests/test_codebuddy.py— full-run-safetytests/test_community_hub_labels.py— full-run-safetytests/test_community_labels_skill.py— full-run-safetytests/test_confidence.py— full-run-safetytests/test_corrupt_graph_json.py— full-run-safetytests/test_cpp_method_declarations.py— full-run-safetytests/test_cpp_nested_and_cli.py— full-run-safetytests/test_cpp_objc_cross_file_calls.py— full-run-safetytests/test_cpp_preprocess.py— full-run-safetytests/test_cross_extension_reexport_self_cycle.py— full-run-safety- … and 280 more
non-code file(s) changed (
graphify/skill-agents.md,graphify/skill-amp.md,graphify/skill-claw.md,graphify/skill-codex.md,graphify/skill-copilot.md…) → running the full suite for safety (a code graph can't see config/fixture/data deps)
changed code file(s) with no mapped test (
graphify/skill-agents.md,graphify/skill-amp.md,graphify/skill-claw.md,graphify/skill-codex.md,graphify/skill-copilot.md…) — a coverage gap or a missing link — running the full suite rather than only the selected tests
Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.
|
Landed in v0.9.77 via an authorship-preserving cherry-pick, so your commit is on |
What does this PR do?
Fixes #4116.
When every doc, paper and image hits the semantic cache, the split-host runbook says "skip to Part C directly". Part C reads
graphify-out/.graphify_semantic.jsonunconditionally, and only the Step B3 merge writes it. As a result, a second/graphifyon an unchanged mixed corpus fails withFileNotFoundError. Writing an empty file instead hides the error but drops every cached node.Changes, all in
tools/skillgen/fragments/core/core.md, with artifacts regenerated by skillgen:.graphify_semantic.json.graphify-out/.graphify_chunk_*.jsonbefore dispatch. B3 merges every chunk file it finds, and nothing has been dispatched yet at B0. So a chunk left by an interrupted run would otherwise be merged as fresh, and with change 1 that now includes all-cached runs.New test file:
tests/test_skill_semantic_all_cached.py.Type of change
Verification & Invariants
Invariant: every route into Part C leaves a
.graphify_semantic.jsonthat contains the cached semantic nodes plus this run's chunks, and nothing else.How this is proven:
test_all_cached_run_is_routed_through_the_b3_mergechecks that every rendered split host routes the all-cached case through B3.test_all_cached_run_reaches_part_c_with_cached_nodes_and_no_stale_chunksrenders the claude body and executes its real Step B0, three Step B3 and Part C Python blocks in a temp dir. The corpus is all-cached (seeded throughsave_semantic_cachewith the sameprompt_file), and a stale chunk sits on disk. The test asserts that Part C succeeds, the cached node is present and the stale node is absent.v8(35adf43) and pass with the fix.Persisted state: the semantic cache is untouched. B3 already limits cache writes to
.graphify_uncached.txt(allowed_source_files), which is empty in the all-cached case, so nothing is re-saved.Limitations:
--monolith-roundtrip, so they need a sanctioned round-trip change and are left for a follow-up.sys.executable -c, not through bash. It is cross-platform, but it does not exercise the$(cat graphify-out/.graphify_python)wrapper.How was this tested?
Graphify-specific checklist
uv run python -m tools.skillgen --bless) when changing their source fragments.🤖 Generated with Claude Code