Skip to content

About

dlt destination for Altertable

Resources

Stars

0 stars

Watchers

1 watching

Forks

Repository files navigation

dlt-altertable

CI PyPI Python 3.12+ License: Apache 2.0

Load data into Altertable with dlt.

Install and configure

Requires Python 3.12+ and dlt >= 1.30, < 2.

uv add dlt-altertable

Set .dlt/secrets.toml:

[destination.altertable]
host = "api.altertable.ai"
catalog = "lakehouse"
dataset_name = "crm"
username = "..."
password = "..."

Use the HTTP API host. Defaults: HTTPS, port 443, compute size XS. Set dataset_name on the destination. dlt's pipeline setting is ignored. The catalog must exist. Schemas are created automatically.

Settings also accept dlt environment variables (DESTINATION__ALTERTABLE__HOST, etc.), existing ALTERTABLE_* credentials, or arguments to altertable(...).

Use

import dlt
from dlt_altertable import altertable


@dlt.resource(
    name="contacts",
    write_disposition="merge",
    primary_key="id",
    columns={"lastmodifieddate": {"dedup_sort": "desc"}},
)
def contacts():
    yield [
        {"id": 1, "email": "ada@example.com", "lastmodifieddate": 100},
        {"id": 2, "email": "grace@example.com", "lastmodifieddate": 100},
    ]


pipeline = dlt.pipeline(pipeline_name="crm", destination=altertable)
pipeline.run(contacts())

Loading

Disposition Behavior
append Adds rows.
replace Recreates the table, then appends subsequent files.
merge Upserts by primary_key; supports child tables and hard_delete.

Merge requires a non-nullable primary key. In the example, dedup_sort keeps the highest lastmodifieddate per id, comparing incoming and stored rows. delete-insert, insert-only, scd2, merge_key, and ascending dedup_sort are unsupported.

New columns are added automatically. Nested data defaults to JSON strings (max_table_nesting=0). Set altertable(naming_convention="snake_case", max_table_nesting=None) to flatten objects into columns and lists into child tables.

Nested merges replace children and require one root per primary key per load. With child tables or hard_delete, send complete records when inserting or updating: omitted fields become null and omitted lists are cleared. Flat merges without hard_delete preserve columns absent from incoming files.

Use columns={"deleted": {"hard_delete": True}} to delete records and their children (dlt deletion rules).

Append and replace commit one file at a time. Replacement recreates the table on its first file, then appends the rest. Retried appends can duplicate rows. Do not set LOAD__PARALLELISM_STRATEGY=parallel: replacement requires sequential files per table.

Partitioning

Partition hints require append or merge:

from dlt_altertable import altertable_adapter, altertable_partition

resource = altertable_adapter(contacts(), partition=altertable_partition.bucket(16, "id"))
pipeline.run(resource)

Pass partition=[] to reset. Sort hints are not supported yet.

Performance

Override defaults with dlt configuration:

export DATA_WRITER__FILE_MAX_BYTES=268435456  # 256 MiB (0 disables byte-based rotation)
export DATA_WRITER__COMPRESSION=zstd

altertable(max_parallel_load_jobs=2) allows parallel uploads, subject to LOAD__WORKERS.

Read and verify

from dlt_altertable import verify_catalog, verify_load

rows = pipeline.dataset().contacts.limit(10).fetchall()
print(verify_catalog())
print(verify_load(pipeline))

Verification helpers return a list of problems. [] means none found. For Arrow inputs with append or merge, set NORMALIZE__PARQUET_NORMALIZER__ADD_DLT_LOAD_ID=true before loading to use verify_load.

Reads buffer results in memory. Use .arrow() for Arrow or .df() for pandas.

Operational notes

Destructive refresh requires altertable(allow_destructive_refresh=True). Use refresh="drop_resources" to drop and recreate selected resource tables and reset their state. Use refresh="drop_data" to clear their rows and reset their state while preserving the schema.

  • Fresh runners restore incremental state with the same pipeline name, catalog, and schema. Persist pipelines_dir to resume unfinished loads.
  • Keep the one-hour upload timeout: a timed-out request can still complete on the server.

Develop

uv run --locked pytest
uv run --locked ruff check .
uv run --locked ruff format --check .
uvx ty check src

Install the Git hooks with uvx pre-commit install.

Run HTTP integration tests against altertable-mock:

docker run --rm -d --name dlt-altertable-mock -p 127.0.0.1:15100:15000 \
  -e ALTERTABLE_MOCK_USERS=integration:integration \
  ghcr.io/altertable-ai/altertable-mock:latest
DLT_ALTERTABLE_INTEGRATION=1 uv run --locked pytest tests/integration
docker stop dlt-altertable-mock

Apache 2.0.

About

dlt destination for Altertable

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages