10 Commits

Author SHA1 Message Date
ec3d1a5046 Release version 3.1.4 2026-07-31 11:40:03 +02:00
37b735cf7b fix(backup): verify the post-stop volume pass by content
Each volume is copied twice into the same destination: once hot with the
container running, once cold after it is stopped. The cold pass is meant to
correct the hot one, but it uses rsync's quick check, which compares size and
whole-second mtime.

A pre-allocated 16 MiB WAL segment never changes size. When its last pre-stop
write and postgres' shutdown checkpoint fall in the same whole second, the cold
pass skips it while still replacing global/pg_control, whose previous write was
seconds earlier. The generation then holds a post-shutdown pg_control pointing at
a checkpoint LSN whose WAL bytes are still zero. Restoring it gives

    invalid record length at 0/3CB7A20: expected at least 24, got 0
    invalid checkpoint record
    PANIC: could not locate a valid checkpoint record at 0/3CB7A20

and postgres crash-loops until the restore test gives up after 1200s. Observed in
infinito-nexus-core CI run 30586213932, debian leg; the centos leg of the same
run restored cleanly, and the two generations differ by exactly one 16 MiB WAL
segment.

authoritative is a required keyword rather than an optional flag, so each of the
four call sites states which pass it is. Set, it adds --checksum, so the cold
pass compares content. The hot passes keep the quick check.

Reproduced with real rsync in both directions: without --checksum the WAL stays
zero while pg_control is post-shutdown; with it the pair is consistent. Hardlink
dedup against --link-dest is unaffected, measured at 20/20 with generation size
unchanged.

Rejected alternatives, both measured rather than argued: --ignore-times destroys
dedup (20/20 to 0/20 hardlinks, generation 36 to 72 MiB), and --modify-window=-1
makes correctness depend on the destination filesystem preserving sub-second
mtimes, which silently degrades on NFS or ext3.

COST. The quick check scales with file count; --checksum scales with bytes, read
on both sides, and it runs while the container is stopped. Measured warm-cache on
215 MiB: 0.10s to 0.39s, a factor of 3.9. Derived for disk-bound runs, the extra
stop time stays under a minute up to roughly 4 GB on spinning disk, 15 GB on a
SATA SSD, 60 GB on NVMe. At 1 TB it is 17 minutes on NVMe and over three hours on
spinning disk; at 10 TB it is hours to days. No size threshold is built in - a
guessed one would drop the guarantee exactly where an unrestorable backup costs
most. Volumes at that scale need filesystem snapshots or pg_basebackup rather
than two rsync passes over a live tree, which is a limit of this approach and not
of this change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 11:32:37 +02:00
9dc57c3235 Release version 3.1.3 2026-07-20 20:01:26 +02:00
8409843ff9 fix(restore): wrap the postgres dump replay in a single transaction
The dump replay ran statement-by-statement with autocommit. When restore --empty runs against a LIVE database, the pre-clean drops a table, the replay recreates it and autocommits, and a concurrent writer (discourse's mini_scheduler upserting scheduler_stats(id=1)) inserts the same primary key into the empty table before the dump's COPY loads it. The COPY then aborts with a duplicate-key violation under ON_ERROR_STOP and the whole restore fails. Running the replay with --single-transaction keeps the recreated table invisible to other sessions until commit, so the writer can never insert the racing row.

The --empty pre-clean stays multi-statement (\gexec, one DROP per statement): running every DROP in one transaction exhausts max_locks_per_transaction on large schemas (e.g. gitlab).

Extract the pre-clean SQL from the inline string into restore/db/empty_preclean.sql (loaded via dirname(__file__)) and declare it as package-data so it ships in the wheel. Add a unit test guarding the single-transaction/multi-statement split and an e2e that reproduces the live-writer race.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 20:00:31 +02:00
d2ba2eb5ae Release version 3.1.2 2026-07-18 00:39:10 +02:00
6a016d7a58 chore(claude): ignore local runtime state under .claude, keep settings.json
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-18 00:38:38 +02:00
d1d5445b1d chore(claude): pause for confirmation before CHANGELOG or pyproject edits
Release metadata (version bump and changelog entry) stays a manual step; agent edits to CHANGELOG.md and pyproject.toml now require explicit operator approval. The local .mcp.json is ignored.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-18 00:37:28 +02:00
bd267cc280 fix(restore): drop text search objects in the postgres --empty pre-clean
A schema shipping a custom text search dictionary (taiga's english_stem_nostop) survived the pre-clean and aborted the dump replay with a duplicate pg_ts_dict_dictname_index violation under ON_ERROR_STOP. The discovery SELECT now also enumerates user-owned pg_ts_config and pg_ts_dict entries. The string-assertion unit test is replaced by real scenario data in the e2e: the seeded schema now contains an overloaded f()/f(int) pair and the nostop dictionary plus configuration, and the restored database is queried to prove each survives the backup, pre-clean and replay cycle exactly once.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-18 00:34:53 +02:00
53460242d8 Release version 3.1.1 2026-07-17 17:28:54 +02:00
01a00dd791 fix(restore): drop overloaded functions by identity signature in the --empty pre-clean
A signature-less DROP FUNCTION aborts with 'function name is not unique' as soon as a schema overloads a name, killing the whole --empty replay under ON_ERROR_STOP. The function branch now emits name(identity args) via pg_get_function_identity_arguments; because that compound must not be identifier-quoted as a whole, quoting moves from the outer format into each branch, guarded by a unit test that pins the per-branch %I quoting.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 17:27:55 +02:00
13 changed files with 444 additions and 106 deletions

3
.claude/.gitignore vendored Normal file
View File

@@ -0,0 +1,3 @@
*
!.gitignore
!settings.json

10
.claude/settings.json Normal file
View File

@@ -0,0 +1,10 @@
{
"permissions": {
"ask": [
"Edit(CHANGELOG.md)",
"Write(CHANGELOG.md)",
"Edit(pyproject.toml)",
"Write(pyproject.toml)"
]
}
}

1
.gitignore vendored
View File

@@ -3,3 +3,4 @@ artifacts/
*.egg-info
dist/
build/
.mcp.json

View File

@@ -1,5 +1,78 @@
# Changelog
## [3.1.4] - 2026-07-31
- Backup: the post-stop volume pass now compares by content (*--checksum*), so
it can no longer skip a file the hot pass had copied from a live source. Each
volume is rsynced twice into the same destination — once with the container
running, once after it is stopped — and the second pass used rsync's quick
check (size plus whole-second mtime). A pre-allocated 16 MiB WAL segment never
changes size, so when its last pre-stop write and postgres' shutdown
checkpoint fell in the same whole second, the cold pass skipped it while still
replacing *global/pg_control*, whose previous write was seconds earlier. The
generation then paired a post-shutdown *pg_control* with a WAL segment still
zeroed at the recorded checkpoint LSN, and restoring it crash-looped postgres
with "invalid record length … expected at least 24, got 0" followed by "PANIC:
could not locate a valid checkpoint record". *backup_volume* takes
*authoritative* as a required keyword rather than an optional flag, so each of
the four call sites states which pass it is; the hot passes keep the quick
check. *--ignore-times* was rejected because it destroys the *--link-dest*
hardlink dedup (measured 20/20 to 0/20, generation size doubled), and
*--modify-window=-1* because it makes correctness depend on the destination
filesystem preserving sub-second mtimes, which degrades silently on NFS or
ext3.
- Cost: the quick check scales with file count, *--checksum* with bytes read on
both sides, and it runs while the container is stopped. The extra stop time
stays under a minute up to roughly 4 GB on spinning disk, 15 GB on a SATA SSD
and 60 GB on NVMe; at 1 TB it is 17 minutes on NVMe and over three hours on
spinning disk. No size threshold is built in, since a guessed one would drop
the guarantee exactly where an unrestorable backup costs most — volumes at
that scale want filesystem snapshots or *pg_basebackup* rather than two rsync
passes over a live tree.
## [3.1.3] - 2026-07-20
- Restore: the postgres dump replay now runs under *--single-transaction*,
so a concurrent writer on a live database can no longer interleave a row
between the replay's table re-create and its *COPY* and trip a "duplicate
key value violates unique constraint" abort under ON_ERROR_STOP. This is
the discourse restore-drill race (*mini_scheduler* upserting
*scheduler_stats(id=1)* mid-restore) that failed the whole restore. The
*--empty* pre-clean stays multi-statement (*\gexec*, one DROP per
statement) because a single DROP transaction exhausts
*max_locks_per_transaction* on large schemas (e.g. gitlab).
- Refactor: the *--empty* pre-clean SQL moves out of the inline Python
string into *src/baudolo/restore/db/empty_preclean.sql* (loaded via
*dirname(__file__)*, declared as package-data so it ships in the wheel).
- Tests: a unit test guards the single-transaction / multi-statement split
(replay carries *--single-transaction*, pre-clean does not); a new e2e
reproduces the live-writer race and asserts the restore survives it.
## [3.1.2] - 2026-07-18
- Restore: the postgres *--empty* pre-clean also drops user-owned text
search configurations and dictionaries (*pg_ts_config*, *pg_ts_dict*),
so a schema shipping a custom dictionary (e.g. taiga's
*english_stem_nostop*) no longer aborts the replay with "duplicate key
value violates unique constraint pg_ts_dict_dictname_index" under
ON_ERROR_STOP.
- Tests: the string-assertion unit test for the pre-clean SQL is replaced
by real scenario data in the e2e: the seeded schema contains an
overloaded *f()/f(int)* pair and the nostop dictionary plus
configuration, and the restored database is queried to prove each
survives the backup, pre-clean and replay cycle exactly once.
## [3.1.1] - 2026-07-17
- Restore: the postgres *--empty* pre-clean drops functions and procedures
by their identity signature (*pg_get_function_identity_arguments*), so a
schema that overloads a function name (e.g. discourse) no longer aborts
the replay with "function name is not unique" under ON_ERROR_STOP.
Identifier quoting moves from the outer DROP format into each object
branch, since the *name(args)* compound must not be quoted as a whole; a
unit test pins the per-branch *%I* quoting so future branches cannot
regress unquoted.
## [3.1.0] - 2026-07-15
- Restore: the postgres *--empty* pre-clean emits one DROP per object and

View File

@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
[project]
name = "backup-docker-to-local"
version = "3.1.0"
version = "3.1.4"
description = "Backup Docker volumes to local with rsync and optional DB dumps."
readme = "README.md"
requires-python = ">=3.9"
@@ -27,3 +27,6 @@ package-dir = { "" = "src" }
[tool.setuptools.packages.find]
where = ["src"]
exclude = ["tests*"]
[tool.setuptools.package-data]
"baudolo.restore.db" = ["*.sql"]

View File

@@ -226,19 +226,19 @@ def main() -> int:
if args.everything:
# "everything": always do pre-rsync, then stop + rsync again
stoppable = filter_stoppable(containers)
backup_volume(versions_dir, volume_name, vol_dir)
backup_volume(versions_dir, volume_name, vol_dir, authoritative=False)
change_containers_status(stoppable, "stop")
backup_volume(versions_dir, volume_name, vol_dir)
backup_volume(versions_dir, volume_name, vol_dir, authoritative=True)
if not args.shutdown:
change_containers_status(stoppable, "start")
continue
# default: rsync, and if needed stop + rsync
backup_volume(versions_dir, volume_name, vol_dir)
backup_volume(versions_dir, volume_name, vol_dir, authoritative=False)
if requires_stop(containers, args.images_no_stop_required):
stoppable = filter_stoppable(containers)
change_containers_status(stoppable, "stop")
backup_volume(versions_dir, volume_name, vol_dir)
backup_volume(versions_dir, volume_name, vol_dir, authoritative=True)
if not args.shutdown:
change_containers_status(stoppable, "start")

View File

@@ -24,16 +24,26 @@ def get_last_backup_dir(
return None
def backup_volume(versions_dir: str, volume_name: str, volume_dir: str) -> None:
"""Perform incremental file backup of a Docker volume."""
def backup_volume(
versions_dir: str, volume_name: str, volume_dir: str, *, authoritative: bool
) -> None:
"""Perform incremental file backup of a Docker volume.
Args:
authoritative: compare source and destination by content instead of by
size and whole-second mtime. Required on a pass whose destination was
already written from a live source, where a file can differ while
both attributes still agree.
"""
dest = os.path.join(volume_dir, "files") + "/"
pathlib.Path(dest).mkdir(parents=True, exist_ok=True)
last = get_last_backup_dir(versions_dir, volume_name, dest)
link_dest = f"--link-dest='{last}'" if last else ""
source = get_storage_path(volume_name)
verify = "--checksum " if authoritative else ""
cmd = f"rsync -abP --delete --delete-excluded {link_dest} {source} {dest}"
cmd = f"rsync -abP --delete --delete-excluded {verify}{link_dest} {source} {dest}"
try:
execute_shell_command(cmd)

View File

@@ -0,0 +1,64 @@
-- Owner-filtered pre-clean for `restore --empty`. Emitted as one DROP per row and
-- run via \gexec so each executes as its own top-level statement: a single DO-block
-- would run every DROP in one transaction and exhaust max_locks_per_transaction on
-- large schemas (e.g. gitlab). Also drops user-owned non-public schemas so a dump
-- that CREATE SCHEMAs (e.g. discourse's discourse_functions) does not fail on an
-- already-existing schema. Extension members (pg_trgm's set_limit) are
-- superuser-owned; IF EXISTS absorbs the CASCADE fallout.
SELECT format('DROP %s IF EXISTS public.%s CASCADE', obj.type, obj.name)
FROM (
SELECT format('%I', c.relname) AS name,
CASE c.relkind
WHEN 'v' THEN 'VIEW'
WHEN 'm' THEN 'MATERIALIZED VIEW'
WHEN 'f' THEN 'FOREIGN TABLE'
ELSE 'TABLE'
END AS type
FROM pg_class c JOIN pg_namespace n ON n.oid = c.relnamespace
WHERE n.nspname = 'public' AND c.relkind IN ('r', 'p', 'v', 'm', 'f')
AND pg_get_userbyid(c.relowner) = current_user
UNION ALL
-- Overloaded functions share a proname; DROP needs the identity
-- signature or psql aborts with "function name is not unique".
SELECT format('%I(%s)', p.proname, pg_get_function_identity_arguments(p.oid)) AS name,
CASE p.prokind WHEN 'p' THEN 'PROCEDURE' ELSE 'FUNCTION' END AS type
FROM pg_proc p JOIN pg_namespace n ON n.oid = p.pronamespace
WHERE n.nspname = 'public' AND p.prokind IN ('f', 'p', 'w')
AND pg_get_userbyid(p.proowner) = current_user
UNION ALL
SELECT format('%I', c.relname) AS name, 'SEQUENCE' AS type
FROM pg_class c JOIN pg_namespace n ON n.oid = c.relnamespace
WHERE n.nspname = 'public' AND c.relkind = 'S'
AND pg_get_userbyid(c.relowner) = current_user
UNION ALL
SELECT format('%I', t.typname) AS name, 'TYPE' AS type
FROM pg_type t JOIN pg_namespace n ON n.oid = t.typnamespace
WHERE n.nspname = 'public'
AND pg_get_userbyid(t.typowner) = current_user
AND (t.typtype IN ('e', 'd')
OR (t.typtype = 'c' AND EXISTS (
SELECT 1 FROM pg_class c2
WHERE c2.oid = t.typrelid AND c2.relkind = 'c')))
UNION ALL
SELECT format('%I', col.collname) AS name, 'COLLATION' AS type
FROM pg_collation col JOIN pg_namespace n ON n.oid = col.collnamespace
WHERE n.nspname = 'public'
AND pg_get_userbyid(col.collowner) = current_user
UNION ALL
SELECT format('%I', ts.cfgname) AS name, 'TEXT SEARCH CONFIGURATION' AS type
FROM pg_ts_config ts JOIN pg_namespace n ON n.oid = ts.cfgnamespace
WHERE n.nspname = 'public'
AND pg_get_userbyid(ts.cfgowner) = current_user
UNION ALL
SELECT format('%I', d.dictname) AS name, 'TEXT SEARCH DICTIONARY' AS type
FROM pg_ts_dict d JOIN pg_namespace n ON n.oid = d.dictnamespace
WHERE n.nspname = 'public'
AND pg_get_userbyid(d.dictowner) = current_user
) obj
UNION ALL
SELECT format('DROP SCHEMA IF EXISTS %I CASCADE', n.nspname)
FROM pg_namespace n
WHERE NOT starts_with(n.nspname, 'pg_')
AND n.nspname NOT IN ('public', 'information_schema')
AND pg_get_userbyid(n.nspowner) = current_user
\gexec

View File

@@ -7,6 +7,7 @@ from collections.abc import Iterable, Iterator
from ..run import docker_exec
_SUPERUSER_ONLY_PREFIXES = (b"COMMENT ON EXTENSION", b"ALTER DEFAULT PRIVILEGES")
_EMPTY_PRECLEAN_SQL = os.path.join(os.path.dirname(__file__), "empty_preclean.sql")
def filter_superuser_only_lines(lines: Iterable[bytes]) -> Iterator[bytes]:
@@ -53,59 +54,8 @@ def restore_postgres_sql(
docker_env = {"PGPASSWORD": password}
if empty:
# Owner-filtered pre-clean emitted as one DROP per row and run via \gexec so each
# executes as its own top-level statement: a single DO-block runs every DROP in one
# transaction and exhausts max_locks_per_transaction on large schemas (e.g. gitlab).
# Also drop user-owned non-public schemas so a dump that CREATE SCHEMAs (e.g.
# discourse's discourse_functions) does not fail on an already-existing schema.
# Extension members (pg_trgm's set_limit) are superuser-owned; IF EXISTS absorbs CASCADE fallout.
drop_sql = r"""
SELECT format('DROP %s IF EXISTS public.%I CASCADE', obj.type, obj.name)
FROM (
SELECT c.relname AS name,
CASE c.relkind
WHEN 'v' THEN 'VIEW'
WHEN 'm' THEN 'MATERIALIZED VIEW'
WHEN 'f' THEN 'FOREIGN TABLE'
ELSE 'TABLE'
END AS type
FROM pg_class c JOIN pg_namespace n ON n.oid = c.relnamespace
WHERE n.nspname = 'public' AND c.relkind IN ('r', 'p', 'v', 'm', 'f')
AND pg_get_userbyid(c.relowner) = current_user
UNION ALL
SELECT p.proname AS name,
CASE p.prokind WHEN 'p' THEN 'PROCEDURE' ELSE 'FUNCTION' END AS type
FROM pg_proc p JOIN pg_namespace n ON n.oid = p.pronamespace
WHERE n.nspname = 'public' AND p.prokind IN ('f', 'p', 'w')
AND pg_get_userbyid(p.proowner) = current_user
UNION ALL
SELECT c.relname AS name, 'SEQUENCE' AS type
FROM pg_class c JOIN pg_namespace n ON n.oid = c.relnamespace
WHERE n.nspname = 'public' AND c.relkind = 'S'
AND pg_get_userbyid(c.relowner) = current_user
UNION ALL
SELECT t.typname AS name, 'TYPE' AS type
FROM pg_type t JOIN pg_namespace n ON n.oid = t.typnamespace
WHERE n.nspname = 'public'
AND pg_get_userbyid(t.typowner) = current_user
AND (t.typtype IN ('e', 'd')
OR (t.typtype = 'c' AND EXISTS (
SELECT 1 FROM pg_class c2
WHERE c2.oid = t.typrelid AND c2.relkind = 'c')))
UNION ALL
SELECT col.collname AS name, 'COLLATION' AS type
FROM pg_collation col JOIN pg_namespace n ON n.oid = col.collnamespace
WHERE n.nspname = 'public'
AND pg_get_userbyid(col.collowner) = current_user
) obj
UNION ALL
SELECT format('DROP SCHEMA IF EXISTS %I CASCADE', n.nspname)
FROM pg_namespace n
WHERE NOT starts_with(n.nspname, 'pg_')
AND n.nspname NOT IN ('public', 'information_schema')
AND pg_get_userbyid(n.nspowner) = current_user
\gexec
"""
with open(_EMPTY_PRECLEAN_SQL, encoding="utf-8") as preclean:
drop_sql = preclean.read()
docker_exec(
container,
["psql", "-v", "ON_ERROR_STOP=1", "-U", user, "-d", db_name],
@@ -122,7 +72,16 @@ SELECT format('DROP SCHEMA IF EXISTS %I CASCADE', n.nspname)
filtered.seek(0)
docker_exec(
container,
["psql", "-v", "ON_ERROR_STOP=1", "-U", user, "-d", db_name],
[
"psql",
"--single-transaction",
"-v",
"ON_ERROR_STOP=1",
"-U",
user,
"-d",
db_name,
],
stdin=filtered,
docker_env=docker_env,
)

View File

@@ -21,6 +21,9 @@ from .helpers import (
# --empty runs against the still-populated DB (no wipe), so the pre-clean must
# drop discourse_functions too or the dump's CREATE SCHEMA aborts the replay
# under ON_ERROR_STOP; the old public-only DROP left it and broke discourse.
# The f()/f(int) pair reproduces discourse's overload abort ("function name
# is not unique") and english_stem_nostop reproduces taiga's text search
# dictionary abort (duplicate pg_ts_dict_dictname_index).
SCENARIO_SQL = (
"CREATE SCHEMA discourse_functions;"
"CREATE TABLE discourse_functions.helper (id int);"
@@ -31,7 +34,13 @@ SCENARIO_SQL = (
"CREATE SEQUENCE public.s;"
"CREATE TYPE public.mood AS ENUM ('ok', 'bad');"
"CREATE FUNCTION public.f() RETURNS int LANGUAGE sql AS 'SELECT 1';"
"CREATE FUNCTION public.f(i int) RETURNS int LANGUAGE sql AS 'SELECT i';"
"CREATE COLLATION public.c (locale = 'C');"
"CREATE TEXT SEARCH DICTIONARY public.english_stem_nostop"
" (Template = snowball, Language = english);"
"CREATE TEXT SEARCH CONFIGURATION public.english_nostop (COPY = english);"
"ALTER TEXT SEARCH CONFIGURATION public.english_nostop"
" ALTER MAPPING FOR asciiword WITH public.english_stem_nostop;"
)
@@ -152,6 +161,26 @@ class TestE2EPostgresEmptyDropHard(unittest.TestCase):
"1",
)
def test_overloaded_functions_restored_once_each(self) -> None:
self.assertEqual(
self._scalar("SELECT count(*) FROM pg_proc WHERE proname='f';"), "2"
)
self.assertEqual(self._scalar("SELECT public.f(41) + public.f();"), "42")
def test_text_search_dictionary_restored_once(self) -> None:
self.assertEqual(
self._scalar(
"SELECT count(*) FROM pg_ts_dict WHERE dictname='english_stem_nostop';"
),
"1",
)
self.assertEqual(
self._scalar(
"SELECT count(*) FROM pg_ts_config WHERE cfgname='english_nostop';"
),
"1",
)
if __name__ == "__main__":
unittest.main()

View File

@@ -0,0 +1,185 @@
# tests/e2e/test_e2e_postgres_single_transaction_live_writer.py
import unittest
from .helpers import (
POSTGRES_IMAGE,
POSTGRES_DATA_DIR,
backup_run,
cleanup_docker,
create_minimal_compose_dir,
ensure_empty_dir,
latest_version_dir,
require_docker,
run,
unique,
wait_for_postgres,
write_databases_csv,
)
# The discourse restore-drill race: `restore --empty` replays the dump into a
# LIVE database while a background writer keeps touching a primary-key row
# (discourse's mini_scheduler upserts scheduler_stats(id=1)). Without a
# single-transaction replay, the pre-clean drops the table, the replay recreates
# it and auto-commits, the writer wins the gap and inserts id=1, and the dump's
# COPY of the same id then aborts with a duplicate-key violation under
# ON_ERROR_STOP -> the whole restore fails. The --single-transaction replay keeps
# the recreated table invisible until commit, so the writer can never insert the
# racing row and the restore completes. A wide filler table makes the COPY slow
# enough that the non-transactional variant loses the race deterministically.
SEED_SQL = (
"CREATE TABLE public.scheduler_stats (id int primary key, v text);"
"INSERT INTO public.scheduler_stats VALUES (1, 'from-dump');"
"CREATE TABLE public.filler (id serial primary key, blob text);"
"INSERT INTO public.filler (blob)"
" SELECT repeat('x', 512) FROM generate_series(1, 100000);"
)
WRITER_LOOP = (
"while true; do "
"psql -h 127.0.0.1 -U postgres -d appdb "
"-c \"INSERT INTO public.scheduler_stats(id, v) VALUES (1, 'live') "
'ON CONFLICT (id) DO NOTHING;" >/dev/null 2>&1; '
"done"
)
class TestE2EPostgresSingleTransactionLiveWriter(unittest.TestCase):
@classmethod
def setUpClass(cls) -> None:
require_docker()
cls.prefix = unique("baudolo-e2e-pg-single-txn")
cls.backups_dir = f"/tmp/{cls.prefix}/Backups"
ensure_empty_dir(cls.backups_dir)
cls.compose_dir = create_minimal_compose_dir(f"/tmp/{cls.prefix}")
cls.repo_name = cls.prefix
cls.pg_container = f"{cls.prefix}-pg"
cls.pg_volume = f"{cls.prefix}-pg-vol"
cls.writer = f"{cls.prefix}-writer"
cls.containers = [cls.pg_container, cls.writer]
cls.volumes = [cls.pg_volume]
run(["docker", "volume", "create", cls.pg_volume])
run(
[
"docker",
"run",
"-d",
"--name",
cls.pg_container,
"-e",
"POSTGRES_PASSWORD=pgpw",
"-e",
"POSTGRES_DB=appdb",
"-e",
"POSTGRES_USER=postgres",
"-v",
f"{cls.pg_volume}:{POSTGRES_DATA_DIR}",
POSTGRES_IMAGE,
]
)
wait_for_postgres(cls.pg_container, user="postgres", timeout_s=90)
run(
[
"docker",
"exec",
cls.pg_container,
"sh",
"-lc",
f'psql -U postgres -d appdb -v ON_ERROR_STOP=1 -c "{SEED_SQL}"',
]
)
cls.databases_csv = f"/tmp/{cls.prefix}/databases.csv"
write_databases_csv(
cls.databases_csv, [(cls.pg_container, "appdb", "postgres", "pgpw")]
)
backup_run(
backups_dir=cls.backups_dir,
repo_name=cls.repo_name,
compose_dir=cls.compose_dir,
databases_csv=cls.databases_csv,
database_containers=[cls.pg_container],
images_no_stop_required=[POSTGRES_IMAGE],
)
cls.hash, cls.version = latest_version_dir(cls.backups_dir, cls.repo_name)
run(
[
"docker",
"run",
"-d",
"--name",
cls.writer,
"--network",
f"container:{cls.pg_container}",
"-e",
"PGPASSWORD=pgpw",
POSTGRES_IMAGE,
"sh",
"-lc",
WRITER_LOOP,
]
)
cls.restore = run(
[
"baudolo-restore",
"postgres",
cls.pg_volume,
cls.hash,
cls.version,
"--backups-dir",
cls.backups_dir,
"--repo-name",
cls.repo_name,
"--container",
cls.pg_container,
"--db-name",
"appdb",
"--db-user",
"postgres",
"--db-password",
"pgpw",
"--empty",
],
capture=True,
check=False,
)
run(["docker", "rm", "-f", cls.writer], capture=True, check=False)
@classmethod
def tearDownClass(cls) -> None:
cleanup_docker(containers=cls.containers, volumes=cls.volumes)
def _scalar(self, sql: str) -> str:
p = run(
[
"docker",
"exec",
self.pg_container,
"sh",
"-lc",
f'psql -U postgres -d appdb -t -A -c "{sql}"',
]
)
return (p.stdout or "").strip()
def test_restore_survived_the_live_writer(self) -> None:
self.assertEqual(
self.restore.returncode,
0,
f"restore aborted (duplicate-key race not contained):\n{self.restore.stderr}",
)
def test_primary_key_row_restored(self) -> None:
self.assertEqual(
self._scalar("SELECT count(*) FROM public.scheduler_stats WHERE id=1;"),
"1",
)
if __name__ == "__main__":
unittest.main()

View File

@@ -1,43 +0,0 @@
"""The --empty pre-clean drop must cover every object class a dump can
re-CREATE ahead of tables; a missed class aborts the replay under
ON_ERROR_STOP (OpenProject's ICU collation public.versions_name did)."""
import tempfile
import unittest
from pathlib import Path
from unittest import mock
from baudolo.restore.db import postgres
class TestPostgresEmptyDrop(unittest.TestCase):
def _drop_sql(self) -> str:
with tempfile.NamedTemporaryFile(suffix=".sql") as fh:
Path(fh.name).write_bytes(b"SELECT 1;\n")
with mock.patch.object(postgres, "docker_exec") as run:
postgres.restore_postgres_sql(
container="c",
db_name="db",
user="u",
password="p",
sql_path=fh.name,
empty=True,
)
first_call = run.call_args_list[0]
return first_call.kwargs["stdin"].decode()
def test_drop_covers_collations(self) -> None:
sql = self._drop_sql()
self.assertIn("pg_collation", sql)
self.assertIn("'COLLATION' AS type", sql)
self.assertIn("pg_get_userbyid(col.collowner) = current_user", sql)
def test_drop_still_covers_the_other_classes(self) -> None:
sql = self._drop_sql()
for marker in ("pg_class", "pg_proc", "'SEQUENCE' AS type", "'TYPE' AS type"):
self.assertIn(marker, sql)
self.assertIn("DROP %s IF EXISTS public.%I CASCADE", sql)
if __name__ == "__main__":
unittest.main()

View File

@@ -0,0 +1,44 @@
import tempfile
import unittest
from unittest.mock import MagicMock, patch
from baudolo.restore.db import postgres as pg_mod
class TestPostgresSingleTransaction(unittest.TestCase):
def test_replay_is_single_transaction_but_preclean_is_not(self) -> None:
calls = []
def _capture(container, argv, **kwargs):
calls.append(argv)
return MagicMock()
with tempfile.NamedTemporaryFile(suffix=".sql") as sql:
sql.write(b"CREATE TABLE t (id int);\nINSERT INTO t VALUES (1);\n")
sql.flush()
with patch.object(pg_mod, "docker_exec", side_effect=_capture):
pg_mod.restore_postgres_sql(
container="db",
db_name="discourse",
user="discourse",
password="pw",
sql_path=sql.name,
empty=True,
)
self.assertEqual(len(calls), 2, f"expected pre-clean + replay: {calls}")
preclean, replay = calls[0], calls[1]
self.assertNotIn(
"--single-transaction",
preclean,
"pre-clean must stay multi-statement or it exhausts max_locks on large schemas",
)
self.assertIn(
"--single-transaction",
replay,
"dump replay must be atomic so a live concurrent writer cannot trip a duplicate-key abort",
)
if __name__ == "__main__":
unittest.main()