Version: graphifyy 0.9.73 (also 0.9.6)
Summary
Hyperedge ids are chosen per extraction, so two different documents can emit hyperedges with the same id. A full build keeps both. An incremental build_merge that re-extracts only one of the two documents silently drops the other document's hyperedge, even though that document was not re-extracted.
The cause is that build_merge carries forward the existing hyperedges of untouched files through export.attach_hyperedges. That function de-duplicates by id alone (seen_ids), so a carried hyperedge loses to any new-chunk hyperedge with the same id from a different source_file.
Seen in the wild on a ~14k-node graph: two docs (an ADR and a plan describing the same mechanism) both produced public_pool_invariant_enforcement.
Reproduction (self-contained)
import json, os, tempfile
from pathlib import Path
from graphify.build import build_merge
from graphify.export import to_json
from graphify.cluster import cluster
d = Path(tempfile.mkdtemp()); os.chdir(d)
(d / "a.md").write_text("# A\n\nAlpha.\n"); (d / "b.md").write_text("# B\n\nBeta.\n")
Path("graphify-out").mkdir()
def node(i, sf): return {"id": i, "label": i, "file_type": "concept", "source_file": sf}
def he(sf, members): return {"id": "shared_flow", "label": f"Flow in {sf}", "relation": "participate_in",
"nodes": members, "source_file": sf, "confidence": "INFERRED", "confidence_score": 0.8}
a, b = ["a_one", "a_two", "a_three"], ["b_one", "b_two", "b_three"]
# 1. full build: both hyperedges kept
G = build_merge([{"nodes": [node(i, "a.md") for i in a] + [node(i, "b.md") for i in b], "edges": [],
"hyperedges": [he("a.md", a), he("b.md", b)]}],
graph_path="graphify-out/graph.json", root=".", dedup=False)
to_json(G, cluster(G), "graphify-out/graph.json", force=True)
print([(h["id"], h["source_file"]) for h in json.load(open("graphify-out/graph.json"))["hyperedges"]])
# [('shared_flow', 'a.md'), ('shared_flow', 'b.md')]
# 2. re-extract ONLY b.md
G = build_merge([{"nodes": [node(i, "b.md") for i in b], "edges": [], "hyperedges": [he("b.md", b)]}],
graph_path="graphify-out/graph.json", root=".", dedup=False)
print([(h["id"], h["source_file"]) for h in G.graph["hyperedges"]])
# [('shared_flow', 'b.md')] <- a.md's hyperedge is gone, though a.md was not touched
Expected: after step 2, both hyperedges are present: a.md's carried unchanged and b.md's replaced.
Actual: a.md's hyperedge is silently dropped.
Suggested fix
Treat a hyperedge's identity as (id, source_file) rather than id when de-duplicating carried hyperedges in attach_hyperedges / build_merge, which is consistent with the replace-per-source rule build_merge already applies. Alternatively, namespace hyperedge ids by source file at extraction time.
Version: graphifyy 0.9.73 (also 0.9.6)
Summary
Hyperedge ids are chosen per extraction, so two different documents can emit hyperedges with the same
id. A full build keeps both. An incrementalbuild_mergethat re-extracts only one of the two documents silently drops the other document's hyperedge, even though that document was not re-extracted.The cause is that
build_mergecarries forward the existing hyperedges of untouched files throughexport.attach_hyperedges. That function de-duplicates byidalone (seen_ids), so a carried hyperedge loses to any new-chunk hyperedge with the same id from a differentsource_file.Seen in the wild on a ~14k-node graph: two docs (an ADR and a plan describing the same mechanism) both produced
public_pool_invariant_enforcement.Reproduction (self-contained)
Expected: after step 2, both hyperedges are present:
a.md's carried unchanged andb.md's replaced.Actual:
a.md's hyperedge is silently dropped.Suggested fix
Treat a hyperedge's identity as
(id, source_file)rather thanidwhen de-duplicating carried hyperedges inattach_hyperedges/build_merge, which is consistent with the replace-per-source rulebuild_mergealready applies. Alternatively, namespace hyperedge ids by source file at extraction time.