Transferring data (transferdata)
Stream-copy every row from one umbral database to another, preserving primary and foreign keys. Resumable across interruptions.
transferdata copies rows from one umbral database to another — env1 → env2, a
VPS-to-VPS move, or the data half of an inspectdb port once the target schema
is stood up. Rows land with their source primary keys and foreign keys
intact, so the object graph is identical on the far side.
Both databases must share your app's schema. transferdata copies data, not
DDL — run migrate on the target first so its tables exist. For a Django →
umbral port, inspectdb generates the
schema; this moves the rows afterward.
Run it
# SQLite file to SQLite file:cargo run -- transferdata --from ./old.sqlite3 --to ./new.sqlite3 # Any URL works on either side:cargo run -- transferdata \ --from postgres://user:pass@vps1/app \ --to postgres://user:pass@vps2/app \ --batch 5000The source is opened read-only; the target read-write. The two ends need not
share a backend — a SQLite → Postgres move works (values coerce to the
target's native types: SQLite 0/1 → a Postgres boolean, a UUID id → the
native uuid), which is the typical "upgrade my dev SQLite to production
Postgres" path.
What it guarantees
PKs preserved
Each row keeps its source id — no autoincrement re-numbering — so foreign keys still point where they did.
Parents first
Tables copy in FK-topological order, so a child never lands before its parent (the target's FK constraints would reject it).
Resumable
Interrupted mid-30GB? Rerun the same command — it resumes at the last committed page, no duplicates.
Streamed
Keyset-paginated batches; only one batch is ever in memory, so table size doesn't matter.
M2M links
Many-to-many junction rows copy after both endpoints, so relations like book.tags survive the move intact.
Resume is exact, not best-effort
Each batch's row inserts and its progress checkpoint commit in one target
transaction (tracked in a umbral_transfer_state table the tool owns). A crash
rolls back both, so a rerun re-reads that page cleanly — never a half-written
batch, never a primary-key collision on restart. Kill a transfer over a flaky
link, run the identical command again, and it continues where it stopped.
Options
| Flag | Default | Meaning |
|---|---|---|
--from <url> | — | Source database (URL or SQLite file path). |
--to <url> | — | Target database (URL or SQLite file path). |
--batch <n> | 1000 | Rows per page / per target transaction. |
--only <a,b> | all | Limit to these tables (FK order still respected). |
--map <framework\|file.json> | none | Translate a foreign-shaped source's column names — a framework preset or a custom JSON map file (see below). |
--workers <n> | 1 | Copy this many independent tables concurrently (per FK level). |
--dry-run | off | Print the copy order + source row counts; write nothing. |
# See what would move, and in what order, without touching the target:cargo run -- transferdata --from ./old.sqlite3 --to ./new.sqlite3 --dry-runPorting from another ORM: --map
transferdata normally assumes both ends share the umbral schema. --map <framework>
lets the source be a database shaped by another ORM — closing the loop after
inspectdb generated your models. It undoes
that ORM's column conventions in reverse: a source FK column maps to the clean
umbral field, and an M2M through table's FK columns map to umbral's parent_id /
child_id. Table names are unchanged (inspectdb keeps them), so only the FK and
junction columns translate.
--map | FK column | Junction columns |
|---|---|---|
django | author_id → author | community_id / software_id |
rails (activerecord) | author_id → author | <model>_id |
laravel (eloquent) | author_id → author | <model>_id |
prisma (typeorm) | authorId → author | <model>Id |
Django, Rails and Laravel share the snake-case <field>_id convention; Prisma
(and camelCase JS ORMs) use <field>Id.
# Schema first (inspectdb -> migrate), then the data straight from Django:cargo run -- transferdata --from postgres://user:pass@host/django_db \ --to ./umbral.sqlite3 --map djangoA custom map: --map <file.json>
When the source doesn't fit a single preset — a Prisma database that @map'd
most columns to snake_case but left a handful camelCase, say — pass --map a
path to a JSON file instead of a framework name. Each key is an umbral
field name and each value is the source column it reads from. A per-table
entry under tables wins over the same key in the global columns map, so a
mixed-case source is expressed exactly:
{ "columns": { "some_field": "someSourceColumn" }, "tables": { "verification_attempt": { "created_at": "createdAt" }, "witness": { "witness_merkle_root": "Witness_merkle_root" } }}cargo run -- transferdata --from postgres://... --to ./umbral.sqlite3 --map ./mapping.jsonOnly the columns you list are renamed; every other column reads through unchanged.
--map is a framework name or a file path — if the value isn't one of the
presets, it's treated as a JSON file and a clear error is printed if it can't be
read or parsed.
Passwords need one extra step. Django's password hashes
(pbkdf2_sha256$…, argon2$…) are in a format umbral-auth can't verify — left
as-is, those users can't log in and can't get a clean rejection. After the
transfer, run resetforeignpasswords (from umbral-auth) to replace every
unverifiable hash with a valid argon2 hash of an unknown password. Those
accounts then recover through the normal password-reset (password-forgot)
flow.
cargo run -- resetforeignpasswords# -> Neutralized 1,204 of 1,210 user password hash(es) umbral-auth could not verify.# Those users must set a new password via the password-forgot / reset flow …It's idempotent (umbral-native hashes are left untouched), so it's safe to
re-run. The command targets whichever user model your AuthPlugin is configured
with — the built-in AuthUser, or a custom AuthPlugin. See
reset_unverifiable_passwords for the generic API.
Going faster: --workers
Tables that don't depend on each other copy concurrently. transferdata groups
tables into FK levels (a level's tables reference only earlier levels), and
--workers <n> runs up to n tables from the same level at once — parents still
finish before children. This mostly helps a Postgres target (SQLite serializes
writers); junctions run concurrently after all models.
cargo run -- transferdata --from ./old.sqlite3 --to postgres://…/app --workers 8Not yet
Registered models (single-column PK), their M2M junction rows, cross-backend
moves (SQLite ↔ Postgres), Django-source mapping (--map django), parallel
workers, and circular foreign keys are covered. A self-referential table (a
parent_id tree) or a mutually-referential pair copies under a single
FK-deferred transaction so the cycle resolves at commit. On Postgres this uses
SET CONSTRAINTS ALL DEFERRED, which works because umbral emits FK constraints
DEFERRABLE INITIALLY IMMEDIATE — so a target created by migrate supports it
out of the box (a pre-existing database whose FKs aren't deferrable would need
those constraints re-created). Source mapping covers Django, Rails, Laravel and
Prisma (see the --map table above). The design lives in
docs/decisions/2026-08-16-data-transfer-engine.md.