Skip to content

feat: add CodeTurtle review validation and publishing - #23

Merged
venkatpachala merged 1 commit into
mainfrom
codeturtle-publishing
Sep 27, 2026
Merged

venkatpachala merged 1 commit into
mainfrom
codeturtle-publishing

Conversation

@venkatpachala

Copy link
Copy Markdown
Owner

No description provided.

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Your trial has ended. Reactivate Greptile to resume code reviews.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 3e9c9295b3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread benchmark/score.py
Comment on lines +55 to +58
predictions = [json.loads(path.read_text(encoding="utf-8")) for path in sorted((run_dir / "predictions").glob("pr_*.json"))]
original_count = len(predictions)
if original_count != int(manifest.get("prs_completed") or 0):
raise ValueError("Prediction count differs from sealed manifest")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Verify prediction hashes before scoring sealed runs

The scorer treats a matching file count as proof that predictions are sealed, but it loads their current contents without comparing them to hashes recorded when prediction finished. After gold labels are exposed, any pr_*.json can therefore be edited while preserving the count, and score_run will accept and report the altered results. Record prediction digests in the manifest at sealing time and verify every digest before scoring.

Useful? React with 👍 / 👎.

Comment thread benchmark/run.py
Comment on lines +139 to +141
output = pred_dir / f"pr_{index + cfg.start:03d}.json"; env = os.environ.copy(); env["OLLAMA_MODEL"] = cfg.model
env["LLM_BACKEND"] = cfg.provider
env["OLLAMA_BASE_URL"] = cfg.base_url

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Pass the configured model to the selected backend

When a benchmark config selects provider: openai, this setup still places cfg.model only in OLLAMA_MODEL. AIGateway selects OpenAI's model exclusively from OPENAI_MODEL, so the run silently uses the ambient value or gpt-4o-mini while its manifest records cfg.model; comparisons and reported provenance can therefore describe a different model from the one actually evaluated. Set the backend-specific model variable from cfg.model.

Useful? React with 👍 / 👎.

Comment on lines +62 to +63
identity = hashlib.sha256(json.dumps([target.repo, target.number, target.head_sha, target.base_sha,
event, sorted(f.fingerprint for f in result.product_findings)], sort_keys=True).encode()).hexdigest()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Include the verdict in the publication identity

For the same HEAD, a partial no-finding COMMENT run and a later complete no-finding MERGE run produce the same marker because both map to the COMMENT GitHub event and have the same empty fingerprint set. _existing consequently reports the second run as already published, leaving the earlier inconclusive recommendation visible instead of publishing the completed result. Include the decision and relevant analysis state in the identity.

Useful? React with 👍 / 👎.

@venkatpachala venkatpachala left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

CodeTurtle review

Recommendation: COMMENT
Analysis: partial
Inspection: 2/105 (2%)
Reasons: partial_analysis, insufficient_inspection

This is a review recommendation. No merge operation was performed.
Tests are supporting evidence; this run does not establish BASE/HEAD regression attribution.

  • bundle:B-001: failed (INVALID_JSON)
  • bundle:B-003: failed (INVALID_JSON)
  • bundle:B-004: failed (INVALID_JSON)
  • graph: degraded (graph_unavailable)

No substantiated defects found within the reported analysis scope.

@venkatpachala venkatpachala left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

CodeTurtle integration test: grouped summary and inline delivery. No defect is asserted.

# Self-review workflow for the CodeTurtle repo.
# For other repos, copy examples/github-action.yml instead.
# See docs/github-action.md
# Uses base-repository tooling, never installs the analyzed PR checkout.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

CodeTurtle delivery validation: this is a test annotation, not a defect claim. The summary and this inline were submitted together in one review.

@venkatpachala
venkatpachala merged commit 63d4df7 into main Sep 27, 2026
1 of 2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant