mirror of
https://github.com/kevinveenbirkenbach/docker-volume-backup.git
synced 2026-08-24 23:04:34 +00:00
The --empty pre-clean is a catalog-wide sweep: it drops every non-template database and every non-pg_ role of the instance. On a dedicated instance that is exactly right, because the dump recreates all of it. On a shared one it destroys databases the dump does not carry, with nothing to restore them from - and no test ever executed that sweep, because the e2e dropped the cluster by hand first and left the pre-clean with zero rows to generate. Scoping the sweep to the dump's own inventory looks like the fix and is worse. A surviving database that owns or merely grants to one of the dump's roles pins that role in pg_shdepend; DROP OWNED BY only reaches the control database the pre-clean is connected to, so DROP ROLE fails - after phase 1 has already dropped the dump's databases. ON_ERROR_STOP aborts, the replay never starts, and the instance is left half emptied. So the instance is checked instead. --empty now refuses when the instance holds a database the dump does not carry, names it, and touches nothing. The sweep stays as it was, safe behind that refusal. Reading the dump's inventory needs a real identifier parser: a quoted name may hold spaces, and psql options precede the target of a \\connect line. The e2e no longer drops the cluster itself, so --empty has to do it and the replay has to put it back; a second pass then adds a foreign database and requires the refusal to leave both it and the restored data alone. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
34 lines
1.5 KiB
SQL
34 lines
1.5 KiB
SQL
-- Pre-clean for `restore cluster --empty`. A pg_dumpall stream recreates roles
|
|
-- and databases, so replaying it into a populated cluster dies on the first
|
|
-- CREATE ROLE. Emitted as one DROP per row and run via \gexec so each executes
|
|
-- as its own top-level statement: DROP DATABASE cannot run inside a
|
|
-- transaction block, which rules out a single DO block.
|
|
-- The phase column pins the order: databases must be gone before their owners
|
|
-- can be dropped, and DROP OWNED BY releases what a role still holds in the
|
|
-- control database. Template databases, the control database itself, the pg_*
|
|
-- system roles and the connecting role are kept - the dump does not recreate
|
|
-- them and dropping them would end the session.
|
|
-- The sweep stays catalog-wide on purpose: a scoped one leaves databases that
|
|
-- pin a dumped role in pg_shdepend, and phase 3 then fails after phase 1 has
|
|
-- already dropped. assert_instance_matches_dump refuses before this runs.
|
|
SELECT statement
|
|
FROM (
|
|
SELECT 1 AS phase,
|
|
format('DROP DATABASE IF EXISTS %I', datname) AS statement
|
|
FROM pg_database
|
|
WHERE NOT datistemplate
|
|
AND datname <> current_database()
|
|
UNION ALL
|
|
SELECT 2, format('DROP OWNED BY %I', rolname)
|
|
FROM pg_roles
|
|
WHERE NOT starts_with(rolname, 'pg_')
|
|
AND rolname <> current_user
|
|
UNION ALL
|
|
SELECT 3, format('DROP ROLE IF EXISTS %I', rolname)
|
|
FROM pg_roles
|
|
WHERE NOT starts_with(rolname, 'pg_')
|
|
AND rolname <> current_user
|
|
) drops
|
|
ORDER BY phase
|
|
\gexec
|