mirror of
https://github.com/kevinveenbirkenbach/docker-volume-backup.git
synced 2026-08-24 23:04:34 +00:00
feat(restore): replay pg_dumpall cluster dumps
A databases.csv row asking for every database of an instance (database = '*') makes the backup side write <instance>.cluster.backup.sql via pg_dumpall, and nothing could read it back: the restore CLI knew files, postgres and mariadb. That dump was stored and unrestorable - a format whose producer had no consumer. Adds `baudolo-restore cluster`. Three properties of a cluster stream shape it, and each one bit during development: - It recreates databases, and CREATE DATABASE cannot run inside a transaction block. So unlike the single-database replay this one must NOT be wrapped in --single-transaction. The unit tests now pin both contracts against each other. - It recreates every role including the one the replay connects as, and the pre-clean cannot drop the role holding its own session. That single CREATE ROLE is filtered out of the stream while its ALTER ROLE is kept, because that is what carries the attributes and the password. Found by running it: the first replay died on `role "postgres" already exists`. - --empty means more than for one database: the cluster's databases go first, then DROP OWNED BY releases what a role still holds in the control database, then the roles themselves. The order is pinned by a phase column because \gexec would otherwise emit them interleaved, and a role cannot be dropped while it still owns a database. Without --empty the replay stops at the first object that already exists. Recreating a cluster over a populated one is a decision, not a default. The e2e test drills the real thing: two databases and their owning role are dropped outright and have to come back with their payload and their ownership intact. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
30
src/baudolo/restore/db/cluster_preclean.sql
Normal file
30
src/baudolo/restore/db/cluster_preclean.sql
Normal file
@@ -0,0 +1,30 @@
|
||||
-- Pre-clean for `restore cluster --empty`. A pg_dumpall stream recreates roles
|
||||
-- and databases, so replaying it into a populated cluster dies on the first
|
||||
-- CREATE ROLE. Emitted as one DROP per row and run via \gexec so each executes
|
||||
-- as its own top-level statement: DROP DATABASE cannot run inside a
|
||||
-- transaction block, which rules out a single DO block.
|
||||
-- The phase column pins the order: databases must be gone before their owners
|
||||
-- can be dropped, and DROP OWNED BY releases what a role still holds in the
|
||||
-- control database. Template databases, the control database itself, the pg_*
|
||||
-- system roles and the connecting role are kept - the dump does not recreate
|
||||
-- them and dropping them would end the session.
|
||||
SELECT statement
|
||||
FROM (
|
||||
SELECT 1 AS phase,
|
||||
format('DROP DATABASE IF EXISTS %I', datname) AS statement
|
||||
FROM pg_database
|
||||
WHERE NOT datistemplate
|
||||
AND datname <> current_database()
|
||||
UNION ALL
|
||||
SELECT 2, format('DROP OWNED BY %I', rolname)
|
||||
FROM pg_roles
|
||||
WHERE NOT starts_with(rolname, 'pg_')
|
||||
AND rolname <> current_user
|
||||
UNION ALL
|
||||
SELECT 3, format('DROP ROLE IF EXISTS %I', rolname)
|
||||
FROM pg_roles
|
||||
WHERE NOT starts_with(rolname, 'pg_')
|
||||
AND rolname <> current_user
|
||||
) drops
|
||||
ORDER BY phase
|
||||
\gexec
|
||||
Reference in New Issue
Block a user