This version is in beta. Some features may change before release.

Transferring data (transferdata)

Stream-copy every row from one umbral database to another, preserving primary and foreign keys. Resumable across interruptions.

transferdata copies rows from one umbral database to another — env1 → env2, a VPS-to-VPS move, or the data half of an inspectdb port once the target schema is stood up. Rows land with their source primary keys and foreign keys intact, so the object graph is identical on the far side.

Info

Both databases must share your app's schema. transferdata copies data, not DDL — run migrate on the target first so its tables exist. For a Django → umbral port, inspectdb generates the schema; this moves the rows afterward.

Run it

Code
bash
# SQLite file to SQLite file:
cargo run -- transferdata --from ./old.sqlite3 --to ./new.sqlite3
 
# Any URL works on either side:
cargo run -- transferdata \
--from postgres://user:pass@vps1/app \
--to postgres://user:pass@vps2/app \
--batch 5000

The source is opened read-only; the target read-write. The two ends need not share a backend — a SQLite → Postgres move works (values coerce to the target's native types: SQLite 0/1 → a Postgres boolean, a UUID id → the native uuid), which is the typical "upgrade my dev SQLite to production Postgres" path.

What it guarantees

PKs preserved

Each row keeps its source id — no autoincrement re-numbering — so foreign keys still point where they did.

Parents first

Tables copy in FK-topological order, so a child never lands before its parent (the target's FK constraints would reject it).

Resumable

Interrupted mid-30GB? Rerun the same command — it resumes at the last committed page, no duplicates.

Streamed

Keyset-paginated batches; only one batch is ever in memory, so table size doesn't matter.

M2M links

Many-to-many junction rows copy after both endpoints, so relations like book.tags survive the move intact.

Resume is exact, not best-effort

Each batch's row inserts and its progress checkpoint commit in one target transaction (tracked in a umbral_transfer_state table the tool owns). A crash rolls back both, so a rerun re-reads that page cleanly — never a half-written batch, never a primary-key collision on restart. Kill a transfer over a flaky link, run the identical command again, and it continues where it stopped.

Options

FlagDefaultMeaning
--from <url>—Source database (URL or SQLite file path).
--to <url>—Target database (URL or SQLite file path).
--batch <n>1000Rows per page / per target transaction.
--only <a,b>allLimit to these tables (FK order still respected).
--map <framework\|file.json>noneTranslate a foreign-shaped source's column names — a framework preset or a custom JSON map file (see below).
--workers <n>1Copy this many independent tables concurrently (per FK level).
--dry-runoffPrint the copy order + source row counts; write nothing.
Code
bash
# See what would move, and in what order, without touching the target:
cargo run -- transferdata --from ./old.sqlite3 --to ./new.sqlite3 --dry-run

Porting from another ORM: --map

transferdata normally assumes both ends share the umbral schema. --map <framework> lets the source be a database shaped by another ORM — closing the loop after inspectdb generated your models. It undoes that ORM's column conventions in reverse: a source FK column maps to the clean umbral field, and an M2M through table's FK columns map to umbral's parent_id / child_id. Table names are unchanged (inspectdb keeps them), so only the FK and junction columns translate.

--mapFK columnJunction columns
djangoauthor_id → authorcommunity_id / software_id
rails (activerecord)author_id → author<model>_id
laravel (eloquent)author_id → author<model>_id
prisma (typeorm)authorId → author<model>Id

Django, Rails and Laravel share the snake-case <field>_id convention; Prisma (and camelCase JS ORMs) use <field>Id.

Code
bash
# Schema first (inspectdb -> migrate), then the data straight from Django:
cargo run -- transferdata --from postgres://user:pass@host/django_db \
--to ./umbral.sqlite3 --map django

A custom map: --map <file.json>

When the source doesn't fit a single preset — a Prisma database that @map'd most columns to snake_case but left a handful camelCase, say — pass --map a path to a JSON file instead of a framework name. Each key is an umbral field name and each value is the source column it reads from. A per-table entry under tables wins over the same key in the global columns map, so a mixed-case source is expressed exactly:

Code
json
{
"columns": {
"some_field": "someSourceColumn"
},
"tables": {
"verification_attempt": { "created_at": "createdAt" },
"witness": { "witness_merkle_root": "Witness_merkle_root" }
}
}
Code
bash
cargo run -- transferdata --from postgres://... --to ./umbral.sqlite3 --map ./mapping.json

Only the columns you list are renamed; every other column reads through unchanged. --map is a framework name or a file path — if the value isn't one of the presets, it's treated as a JSON file and a clear error is printed if it can't be read or parsed.

Warning

Passwords need one extra step. Django's password hashes (pbkdf2_sha256$…, argon2$…) are in a format umbral-auth can't verify — left as-is, those users can't log in and can't get a clean rejection. After the transfer, run resetforeignpasswords (from umbral-auth) to replace every unverifiable hash with a valid argon2 hash of an unknown password. Those accounts then recover through the normal password-reset (password-forgot) flow.

Code
bash
cargo run -- resetforeignpasswords
# -> Neutralized 1,204 of 1,210 user password hash(es) umbral-auth could not verify.
# Those users must set a new password via the password-forgot / reset flow …

It's idempotent (umbral-native hashes are left untouched), so it's safe to re-run. The command targets whichever user model your AuthPlugin is configured with — the built-in AuthUser, or a custom AuthPlugin. See reset_unverifiable_passwords for the generic API.

Going faster: --workers

Tables that don't depend on each other copy concurrently. transferdata groups tables into FK levels (a level's tables reference only earlier levels), and --workers <n> runs up to n tables from the same level at once — parents still finish before children. This mostly helps a Postgres target (SQLite serializes writers); junctions run concurrently after all models.

Code
bash
cargo run -- transferdata --from ./old.sqlite3 --to postgres://…/app --workers 8

Not yet

Registered models (single-column PK), their M2M junction rows, cross-backend moves (SQLite ↔ Postgres), Django-source mapping (--map django), parallel workers, and circular foreign keys are covered. A self-referential table (a parent_id tree) or a mutually-referential pair copies under a single FK-deferred transaction so the cycle resolves at commit. On Postgres this uses SET CONSTRAINTS ALL DEFERRED, which works because umbral emits FK constraints DEFERRABLE INITIALLY IMMEDIATE — so a target created by migrate supports it out of the box (a pre-existing database whose FKs aren't deferrable would need those constraints re-created). Source mapping covers Django, Rails, Laravel and Prisma (see the --map table above). The design lives in docs/decisions/2026-08-16-data-transfer-engine.md.

clidatamigration