Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
49 changes: 38 additions & 11 deletions .github/workflows/pipeline-daily.yml
Original file line number Diff line number Diff line change
@@ -1,20 +1,28 @@
name: pipeline-daily

# Runs the onchain analytics pipeline daily at 01:00 UTC (1h after midnight
# to ensure yesterday's data is fully finalized on all chains), then refreshes
# the dbt models so dashboards show the latest data.
# Manual only. The daily cron trigger was removed deliberately.
#
# Required secrets:
# ENVIO_API_TOKEN -- HyperSync API token
# GCP_SA_KEY -- GCP service account JSON key (BigQuery access)
# SLACK_WEBHOOK_URL -- Slack incoming webhook for alerts (optional)
# Scheduled ingestion stays off until the readiness remediation is complete. Until then this
# workflow runs only on an explicit manual dispatch, and only against a caller supplied dataset
# that is not the production dataset. The guard below enforces that at the workflow boundary, so
# a mistyped or forgotten input stops the run instead of writing to production.
#
# Manual trigger: use workflow_dispatch to run on demand.
# Required secrets:
# ENVIO_API_TOKEN HyperSync API token
# GCP_SA_KEY GCP service account JSON key (BigQuery access)
# SLACK_WEBHOOK_URL Slack incoming webhook for alerts (optional)

on:
schedule:
- cron: '0 1 * * *' # 01:00 UTC daily
workflow_dispatch: # Manual trigger from GitHub UI
workflow_dispatch:
inputs:
dataset_id:
description: 'Target BigQuery dataset. Must not be the production dataset.'
required: true
type: string
confirm:
description: 'Type RUN-AGAINST-SANDBOX to confirm the target is not production.'
required: true
type: string

concurrency:
group: pipeline-daily
Expand All @@ -30,6 +38,24 @@ jobs:
timeout-minutes: 30

steps:
- name: Refuse a production or unconfirmed target
working-directory: .
env:
DATASET_ID: ${{ inputs.dataset_id }}
CONFIRM: ${{ inputs.confirm }}
run: |
if [ "$CONFIRM" != "RUN-AGAINST-SANDBOX" ]; then
echo "Refused: confirmation phrase missing or wrong."
exit 1
fi
case "$DATASET_ID" in
""|BlockchainEvents|Staging|Semantic|Marts)
echo "Refused: '$DATASET_ID' is a production dataset or is empty."
exit 1
;;
esac
echo "Target dataset accepted: $DATASET_ID"

- uses: actions/checkout@v4

- uses: actions/setup-node@v4
Expand All @@ -48,6 +74,7 @@ jobs:

- name: Run pipeline (daily)
env:
DATASET_ID: ${{ inputs.dataset_id }}
ENVIO_API_TOKEN: ${{ secrets.ENVIO_API_TOKEN }}
SLACK_WEBHOOK_URL: ${{ secrets.SLACK_WEBHOOK_URL }}
run: npx tsx src/index.ts daily
Expand Down
115 changes: 115 additions & 0 deletions .github/workflows/pipeline-pr.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,115 @@
name: pipeline-pr

# The pull-request gate for the ingestion pipeline. NO CREDENTIALS, BY DESIGN.
#
# Nothing in this workflow authenticates to anything. It uses no `secrets.*` value, invokes no
# cloud authentication action, and reaches no chain, so a fork pull request runs it in full and a
# compromised dependency has nothing to steal. That is a deliberate property and the reason the
# test harness is built the way it is: every edge the pipeline has is replaceable in process, so
# proving the code works needs no BigQuery and no RPC endpoint.
#
# WHAT IT DOES NOT DO, AND WHEN THAT CHANGES. It does not run the known-failures suite. That suite
# reproduces defects other phases own, so gating pull requests on it would make every pull request
# red for reasons its author cannot act on. `npm run verify:local` runs both and is the command
# that decides whether the candidate is releasable; enforcement of the full suite in CI is added
# by the phase that makes the last known failure green.

on:
pull_request:
paths:
- 'projects/onchain-analytics/pipeline-v5/**'
- 'projects/onchain-analytics/gd_dbt/**'
- '.github/workflows/pipeline-pr.yml'
workflow_dispatch:

permissions:
contents: read

concurrency:
group: pipeline-pr-${{ github.ref }}
cancel-in-progress: true

jobs:
pipeline:
name: Pipeline lint, typecheck and tests
runs-on: ubuntu-latest
timeout-minutes: 15
defaults:
run:
working-directory: projects/onchain-analytics/pipeline-v5

steps:
- uses: actions/checkout@v4

- uses: actions/setup-node@v4
with:
node-version: '20'
cache: 'npm'
cache-dependency-path: projects/onchain-analytics/pipeline-v5/package-lock.json

# `ci`, not `install`. It fails when package.json and the lockfile disagree, which is the
# only way a dependency change can be noticed in review rather than in production.
- name: Install dependencies from the lockfile
run: npm ci

- name: Lint
run: npm run lint

- name: Strict typecheck
run: npm run typecheck

- name: Unit and integration tests with coverage
run: npm run test:coverage

- name: Upload coverage
if: always()
uses: actions/upload-artifact@v4
with:
name: pipeline-coverage
path: projects/onchain-analytics/pipeline-v5/coverage
retention-days: 7

dbt:
name: dbt dependency resolution and parse
runs-on: ubuntu-latest
timeout-minutes: 15
defaults:
run:
working-directory: projects/onchain-analytics/gd_dbt

steps:
- uses: actions/checkout@v4

- uses: actions/setup-python@v5
with:
python-version: '3.11'

- name: Install dbt
run: pip install "dbt-bigquery==1.8.*"

- name: Resolve packages
run: dbt deps

# `parse` builds the manifest from the project alone. It compiles the graph, resolves every
# ref and source, and validates the YAML, all without a warehouse connection, which is what
# lets this job run with no credentials. A broken ref or a malformed schema file fails here.
- name: Parse the project
env:
# A profile must exist for the parse to resolve, and these values are never connected
# to. There is no key, so no authentication can occur even if something tried.
DBT_PROFILES_DIR: ${{ github.workspace }}/projects/onchain-analytics/gd_dbt/.ci
run: |
mkdir -p "$DBT_PROFILES_DIR"
cat > "$DBT_PROFILES_DIR/profiles.yml" <<'YAML'
gd_dbt:
target: parse_only
outputs:
parse_only:
type: bigquery
method: oauth
project: parse-only-no-such-project
dataset: parse_only
threads: 1
location: US
YAML
dbt parse --no-partial-parse
142 changes: 142 additions & 0 deletions .github/workflows/pipeline-sandbox-integration.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,142 @@
name: pipeline-sandbox-integration

# DEFINED AND DISABLED. Do not enable this workflow yet.
#
# This is the credentialled half of the test strategy: the integration checks that genuinely need
# a BigQuery sandbox, which the in-process simulator can only approximate. The definition lands
# before the identity does, so the shape is reviewable before anything can run with it.
#
# WHY IT CANNOT RUN YET, precisely. It needs a workload-identity provider and a service account
# that can write to a sandbox dataset and cannot touch `BlockchainEvents`. Neither exists:
# `iam.serviceAccounts.create` is denied to the current identity at project scope, and no
# previously disabled billable service may be activated to work around that. Creating the
# identity is an administrator action on the Google Cloud project.
#
# THREE THINGS KEEP IT OFF UNTIL THEN, so forgetting one is not enough to start it:
# 1. There is no `schedule` and no `pull_request` trigger. Manual dispatch only.
# 2. The first step refuses unless `confirm` is typed exactly, and refuses every production
# dataset name outright.
# 3. Every job is gated on `vars.SANDBOX_INTEGRATION_ENABLED == 'true'`, a repository variable
# that does not exist. Until someone creates it, dispatching this workflow does nothing.
#
# KEYLESS BY CONSTRUCTION. It authenticates through Workload Identity Federation, never a
# downloaded key. `GCP_SA_KEY` is deliberately absent from this file: a long-lived JSON key in a
# repository secret is the same class of standing credential this project is removing elsewhere.

on:
workflow_dispatch:
inputs:
dataset_id:
description: 'Sandbox BigQuery dataset to create and drop. Must not be a production dataset.'
required: true
type: string
confirm:
description: 'Type RUN-AGAINST-SANDBOX to confirm the target is not production.'
required: true
type: string

permissions:
contents: read
id-token: write

concurrency:
group: pipeline-sandbox-integration
cancel-in-progress: false

jobs:
guard:
name: Refuse unless explicitly enabled and confirmed
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- name: This workflow is disabled
if: vars.SANDBOX_INTEGRATION_ENABLED != 'true'
run: |
echo "Refused: SANDBOX_INTEGRATION_ENABLED is not set to 'true'."
echo "This workflow is defined but disabled until its service account exists."
exit 1

- name: Refuse a production or unconfirmed target
env:
DATASET_ID: ${{ inputs.dataset_id }}
CONFIRM: ${{ inputs.confirm }}
run: |
if [ "$CONFIRM" != "RUN-AGAINST-SANDBOX" ]; then
echo "Refused: confirmation phrase missing or wrong."
exit 1
fi
case "$DATASET_ID" in
""|BlockchainEvents|Staging|Semantic|Marts|dev_sandbox)
echo "Refused: '$DATASET_ID' is a production or in-use dataset, or is empty."
exit 1
;;
esac
echo "Target dataset accepted: $DATASET_ID"

integration:
name: Sandbox integration tests
needs: guard
if: vars.SANDBOX_INTEGRATION_ENABLED == 'true'
runs-on: ubuntu-latest
timeout-minutes: 30
environment: sandbox
defaults:
run:
working-directory: projects/onchain-analytics/pipeline-v5

steps:
- uses: actions/checkout@v4

- uses: actions/setup-node@v4
with:
node-version: '20'
cache: 'npm'
cache-dependency-path: projects/onchain-analytics/pipeline-v5/package-lock.json

- name: Install dependencies from the lockfile
run: npm ci

# Keyless. The provider and service account are supplied as repository variables, not
# secrets, because neither is a credential.
- name: Authenticate to GCP through Workload Identity Federation
uses: google-github-actions/auth@v2
with:
workload_identity_provider: ${{ vars.GCP_WORKLOAD_IDENTITY_PROVIDER }}
service_account: ${{ vars.GCP_SANDBOX_SERVICE_ACCOUNT }}

- name: Sandbox integration tests
env:
DATASET_ID: ${{ inputs.dataset_id }}
STAGING_DATASET_ID: ${{ inputs.dataset_id }}_staging
GCP_PROJECT_ID: ${{ vars.GCP_PROJECT_ID }}
ENVIO_API_TOKEN: ${{ secrets.ENVIO_API_TOKEN }}
# The name the pipeline actually reads. This previously set BQ_MAX_BYTES_BILLED, which
# nothing in the codebase looks at, so the cap it claimed to apply was never applied.
# 10 GiB per job.
MAX_BYTES_BILLED_PER_JOB: '10737418240'
run: npm run test:integration

# The two-process write-safety check. It cannot run in the credential-free suite: it needs
# a real table for two separate processes to contend over, which is the only way to
# demonstrate that one of them is excluded. Everything else about the write path is
# covered in-process, and in-process cannot answer this question.
- name: Write-safety gate, two concurrent processes against a real sandbox
env:
GCP_PROJECT_ID: ${{ vars.GCP_PROJECT_ID }}
ENVIO_API_TOKEN: ${{ secrets.ENVIO_API_TOKEN }}
MAX_BYTES_BILLED_PER_JOB: '10737418240'
run: npm run gate:write-safety

# Runs even when the tests fail, because a sandbox left behind is a defect this project has
# already produced twice: two undocumented datasets exist in the project today, one of them
# recorded as deleted in a maintenance note and still present.
- name: Drop the sandbox dataset
if: always()
env:
DATASET_ID: ${{ inputs.dataset_id }}
GCP_PROJECT_ID: ${{ vars.GCP_PROJECT_ID }}
run: |
echo "Cleanup for $GCP_PROJECT_ID:$DATASET_ID is supplied by the sandbox manager in"
echo "pipeline-v5/src/sandbox.ts, which checks the purpose label before deleting"
echo "anything and proves absence with a listing afterwards."
exit 1
5 changes: 5 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -153,3 +153,8 @@ queries/dune/sybil-monitor/
# one checkout. Giving the mess one address is what actually contains it.
_scratch/
**/_scratch/

# Test coverage output (added 2026-09-28 with the pipeline test harness).
# Generated by `npm run test:coverage`; the PR workflow uploads it as an artifact instead.
coverage/
**/coverage/
22 changes: 22 additions & 0 deletions .shipgate-allow
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,28 @@ queries/dune/reserve-analysis/lp-v3-positions.sql :: content:address-list :: Uni
# --- dbt seeds: address + number pairs, but the number is not a balance -------
projects/onchain-analytics/gd_dbt/seeds/contract_deployments.csv :: content:holder-table :: contract address paired with its deployment BLOCK, not a holder balance
projects/onchain-analytics/gd_dbt/seeds/tokens.csv :: content:holder-table :: token contract address paired with its DECIMALS, not a holder balance
# event_surface.csv is the event ABI catalogue: one row per (contract, era, event).
# Checked 2026-09-28 before waiving, not assumed. All 347 distinct addresses in the
# file are contract addresses already present in contract_deployments.csv above,
# including the 11 that appear outside the two address columns (all in abi_source,
# all of the form bytecode_equality_with_celo_0x..., naming a registry contract).
# Its only numeric columns are chain_id, era_index, indexed_positions and the two
# era block bounds. No balance column and no wallet address exists in it.
# Receipt: specs/_scratch/coord-unit-01/out/01-event-surface-address-provenance.json
projects/onchain-analytics/gd_dbt/seeds/event_surface.csv :: content:holder-table :: contract address paired with its era BLOCK BOUNDS, not a holder balance
# era_intervals.csv and era_boundary_evidence.csv are the era validity-interval table
# and the evidence behind it: which implementation was live between which blocks, and
# how that was established. Checked 2026-09-28 before waiving, not assumed.
# era_intervals.csv 262 rows, 278 distinct addresses, ALL already present in
# contract_deployments.csv above. Numeric columns are chain_id,
# valid_from_block, valid_to_block, era_index.
# era_boundary_evidence.csv 41 rows, 44 distinct addresses, ALL already present above.
# Numeric columns are block numbers plus COUNTS of verification
# checkpoints, events and chunks.
# Neither file has an amount column, a balance, or a wallet address.
# Receipt: specs/_scratch/coord-unit-02-gate/verify-new-seeds.mjs
projects/onchain-analytics/gd_dbt/seeds/era_intervals.csv :: content:holder-table :: contract address paired with its era BLOCK BOUNDS, not a holder balance
projects/onchain-analytics/gd_dbt/seeds/era_boundary_evidence.csv :: content:holder-table :: contract address paired with block numbers and verification COUNTS, not a holder balance

# --- reader-facing docs that trip a filename or attribution rule -------------
projects/goodwidget-components/spec.md :: name:spec.md :: real component specification, reader-facing product doc, not a pipeline artifact
Expand Down
Loading
Loading