Skip to content

build_merge: re-extracting one doc silently drops another doc's hyperedge when both share an id (attach_hyperedges de-dups by id alone) #3981

Description

@stephenc-git

Version: graphifyy 0.9.73 (also 0.9.6)

Summary

Hyperedge ids are chosen per extraction, so two different documents can emit hyperedges with the same id. A full build keeps both. An incremental build_merge that re-extracts only one of the two documents silently drops the other document's hyperedge, even though that document was not re-extracted.

The cause is that build_merge carries forward the existing hyperedges of untouched files through export.attach_hyperedges. That function de-duplicates by id alone (seen_ids), so a carried hyperedge loses to any new-chunk hyperedge with the same id from a different source_file.

Seen in the wild on a ~14k-node graph: two docs (an ADR and a plan describing the same mechanism) both produced public_pool_invariant_enforcement.

Reproduction (self-contained)

import json, os, tempfile
from pathlib import Path
from graphify.build import build_merge
from graphify.export import to_json
from graphify.cluster import cluster

d = Path(tempfile.mkdtemp()); os.chdir(d)
(d / "a.md").write_text("# A\n\nAlpha.\n"); (d / "b.md").write_text("# B\n\nBeta.\n")
Path("graphify-out").mkdir()

def node(i, sf): return {"id": i, "label": i, "file_type": "concept", "source_file": sf}
def he(sf, members): return {"id": "shared_flow", "label": f"Flow in {sf}", "relation": "participate_in",
                             "nodes": members, "source_file": sf, "confidence": "INFERRED", "confidence_score": 0.8}
a, b = ["a_one", "a_two", "a_three"], ["b_one", "b_two", "b_three"]

# 1. full build: both hyperedges kept
G = build_merge([{"nodes": [node(i, "a.md") for i in a] + [node(i, "b.md") for i in b], "edges": [],
                  "hyperedges": [he("a.md", a), he("b.md", b)]}],
                graph_path="graphify-out/graph.json", root=".", dedup=False)
to_json(G, cluster(G), "graphify-out/graph.json", force=True)
print([(h["id"], h["source_file"]) for h in json.load(open("graphify-out/graph.json"))["hyperedges"]])
# [('shared_flow', 'a.md'), ('shared_flow', 'b.md')]

# 2. re-extract ONLY b.md
G = build_merge([{"nodes": [node(i, "b.md") for i in b], "edges": [], "hyperedges": [he("b.md", b)]}],
                graph_path="graphify-out/graph.json", root=".", dedup=False)
print([(h["id"], h["source_file"]) for h in G.graph["hyperedges"]])
# [('shared_flow', 'b.md')]   <- a.md's hyperedge is gone, though a.md was not touched

Expected: after step 2, both hyperedges are present: a.md's carried unchanged and b.md's replaced.
Actual: a.md's hyperedge is silently dropped.

Suggested fix

Treat a hyperedge's identity as (id, source_file) rather than id when de-duplicating carried hyperedges in attach_hyperedges / build_merge, which is consistent with the replace-per-source rule build_merge already applies. Alternatively, namespace hyperedge ids by source file at extraction time.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions