Compare commits

...

8 Commits

Author SHA1 Message Date
e0f89c86ec Release version 7.0.0 2026-08-18 05:13:36 +02:00
f97efb10c4 fix(backup)!: read the instance from one engine-name set, and trust it
Two shapes fell through the inline regex, which knew `database`, `db` and
`postgres` only. A container named exactly after its engine - what a compose
file writes as `container_name: postgres` - carries no separator before the
token, so it resolved to nothing and 6.0.0 stopped dumping it without saying
so. And a swarm task of a central MariaDB reads `mariadb_mariadb.1.<id>`,
where `_mariadb` was no token at all, so that database has never been dumped
under swarm at all.

ENGINE_NAMES states the set once and serves both readings: carried as a
suffix it makes the rest the instance, being one outright makes the container
its own instance.

backup_mariadb_or_postgres stops calling an application container a database.
container_engine recognises an engine by its client tools, which an
application image often ships, so refusing the dump alone would have recorded
the volume as `database: true, dumped: false` - the exact shape a restore
drill reads as a database that was missed. Without an instance there is no
database to record.

BREAKING CHANGE: `mariadb` and `mysql` join the suffix tokens, so a container
named `<app>-mariadb` resolves to the instance `<app>` rather than to its own
name. A databases.csv keyed on the full container name has to move to the
application name, or name the container in --database-containers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 05:11:46 +02:00
704481a505 Release version 6.0.0 2026-08-18 04:32:09 +02:00
b1ee8f5fac fix(backup)!: dump the container that holds the database, with its password
Two defects kept dedicated Postgres databases out of the backup.

docker exec never forwarded PGPASSWORD. execute_to_file set it on baudolo's
own process, but nothing carried it across the container boundary, so an
engine whose pg_hba demands a password on TCP loopback refused the dump.
forward_env passes a bare `-e NAME`, letting docker copy the value out of
this process's environment instead of spelling it into argv, where the
host's process list would publish it.

get_instance returned the container name unchanged when that name carried no
database token, claiming an instance it had never derived. An application
container therefore answered the same databases.csv row as its own dedicated
engine, and application images often ship the engine's client tools, so the
dump command started and wrote a file that looked like a backup and held
none of the data. Discourse is the live case: its launcher names the
container `discourse`, and the image ships pg_dumpall.

The regex stays a normaliser - `<app>-database` from compose and
`<app>_database.1.<task>` from swarm still resolve to the same instance.
Only the fallthrough changes.

BREAKING CHANGE: a database container whose name carries no `database`, `db`
or `postgres` token must now be named in --database-containers. Without that
declaration its rows no longer match and no dump is written.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 04:26:24 +02:00
8dbd5e89ea Release version 5.0.0 2026-08-18 01:55:45 +02:00
efcfe88e7f style!: adopt the core lint bar and migrate to it
The package carried no ruff configuration at all, so it ran on the defaults (E4+E7+E9+F) while infinito-nexus-core, its only consumer, holds itself to a far wider selection. Measured against that selection this tree had 224 findings. It now has none.

The selector list is core's verbatim so both repositories answer to one bar. target-version stays py39 rather than core's py311, because requires-python still declares >=3.9 and pyupgrade would otherwise propose syntax the declared minimum cannot run. Every ignore carries its reason: S603/S607 in particular, since running docker and dump binaries from PATH in list form is this tool's whole job and is already injection-safe.

Two conversions are judgement rather than mechanics. os.path.join(dir, '') was the rsync idiom for a trailing separator, which Path drops, so it becomes an explicit os.sep. os.path.abspath stays where Path.resolve() would follow symlinks and let a symlinked volume test as inside the snapshot subject.

BREAKING CHANGE: VersionMismatch is renamed VersionMismatchError.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 01:51:18 +02:00
94637c32aa feat(manifest)!: record per volume what the run established
A finished generation cannot show whether a volume held a database, nor whether a dump was produced for it: under --only-sql a failed dump falls back to a file copy, and the resulting files/ tree looks like any other copy. The run knows both and threw the knowledge away as a printed warning, leaving every reader to guess from file names.

Each generation now carries a manifest.json stating its layout and, per volume, database / dumped / engine. baudolo.generation is the single place those names are spelled; restore/paths.py, backup/db.py and backup/volume.py stop repeating them. It is deliberately import-free so a consumer can read the manifest with nothing but json, on hosts where this package is not installed.

BREAKING CHANGE: BackupException is renamed BackupError. The rename is atomic across the ten modules that define or import it, three of which also carry the manifest change, so it lands in this commit rather than a separate one that could not import.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 01:48:09 +02:00
03da186a06 build: keep the image context clean instead of wiping the worktree
Every test target required 'clean', which is 'git clean -fdX .' - so running the unit tests deleted every git-ignored file the operator had, venv and caches included. It was compensating for a missing .dockerignore: the Dockerfile's COPY . . otherwise drags __pycache__, egg-info and build output into the image context.

The ignore file fixes that where it belongs, so the prerequisite can go. 'clean' remains available as its own target.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 01:44:12 +02:00
47 changed files with 1265 additions and 213 deletions

13
.dockerignore Normal file
View File

@@ -0,0 +1,13 @@
.git
.github
__pycache__
**/__pycache__
*.egg-info
**/*.egg-info
artifacts/
dist/
build/
.venv
.ruff_cache
.pytest_cache
.mcp.json

View File

@@ -1,5 +1,125 @@
# Changelog # Changelog
## [7.0.0] - 2026-08-18
Breaking:
- Backup: *mariadb* and *mysql* join the suffix tokens, so a container named
*<app>-mariadb* or *<app>-mysql* now resolves to the instance *<app>* instead
of to its own name. A *databases.csv* keyed on the full container name has to
move to the application name, or name the container in
*--database-containers*. This narrows what 6.0.0 broke rather than widening
it: a container named exactly *postgres*, *mariadb*, *mysql*, *db* or
*database* resolves again without any declaration, which is the shape a
compose file writes as *container_name: postgres* and the most common
configuration there is.
Fixed:
- Backup: a container named exactly after its engine is dumped again. The
suffix match needs a hyphen or underscore in front of the token, which a bare
name does not carry, so 6.0.0 resolved *container_name: postgres* to nothing
and stopped dumping it without saying so. *ENGINE_NAMES* now states the set
once and serves both readings — carried as a suffix it makes the rest the
instance, being one outright makes the container its own instance.
- Backup: a central MariaDB under swarm is dumped for the first time. Swarm
names its task *mariadb_mariadb.1.<id>*, which matches neither the static
*mariadb* passed through *--database-containers* nor any token the suffix
match knew, since *_mariadb* is not *_db*. The database was silently absent
from every swarm backup this tool has ever written, before 6.0.0 as well.
- Backup: an application container is no longer recorded as a database.
*container_engine* recognises an engine by its client tools, which an
application image frequently ships, so refusing its dump alone would have
written the volume to the manifest as *database: true, dumped: false* — the
exact shape a restore drill reads as a database that was missed. Without an
instance there is no database to record, and the volume is a file backup like
any other.
## [6.0.0] - 2026-08-18
Breaking:
- Backup: a database container whose name carries no *database*, *db* or
*postgres* token — preceded by a hyphen or underscore — must now be named in
*--database-containers*. Without that declaration its *databases.csv* rows no
longer match and no dump is written, silently, because nothing fails. The
shape this hits hardest is a container named exactly *postgres* or *mariadb*:
the token needs a separator in front of it, which a bare name does not have.
*app-database*, *app_database.1.<task>* from swarm and *app-postgres-1* are
unaffected, as is any container already declared.
Fixed:
- Backup: *docker exec* now forwards *PGPASSWORD* into the container.
*execute_to_file* set the variable on baudolo's own process, but nothing
carried it across the container boundary, so an engine whose *pg_hba* demands
a password on TCP loopback refused every dump — which is every dedicated
Postgres instance on a real host. The name travels as a bare *-e NAME* so
docker copies the value out of this process's environment; spelling
*-e NAME=value* instead would publish the secret in the host's process list.
- Backup: *get_instance* no longer claims an instance it never derived. It
returned the container name unchanged when that name carried no database
token, so an application container answered the same *databases.csv* row as
its own dedicated engine. Application images frequently ship the engine's
client tools, so the dump command started and wrote a file that looked like a
backup and held none of the data: measured against Discourse, 1,680 bytes
from the application where the engine produced 10,469,439. The regex stays a
normaliser — *<app>-database* from compose and *<app>_database.1.<task>* from
swarm still resolve to one instance. Only the fallthrough changed.
New:
- Tests: *get_instance* has unit coverage for the first time. Eleven cases pin
the container names that compose, swarm and explicitly-named engines produce,
so a future change to the regex has to state which shape it gives up.
- Tests: two e2e modules cover shapes the suite structurally could not see.
Every fixture passed its container in *--database-containers*, which left the
regex branch — the only one a dedicated database ever takes — dead code under
test, and no scenario made a password mandatory, because stock
*postgres:alpine* grants trust on loopback.
*test_e2e_postgres_password_required* starts an engine with
*--auth-host=scram-sha-256* and carries a negative control asserting the
server refuses an unauthenticated dump; without it the module would pass
whether or not the password is forwarded at all.
*test_e2e_app_container_ships_client_tools* places an application container
beside its engine with neither declared, and requires the engine dumped, the
application volume copied as files, and no dump written from the application.
## [5.0.0] - 2026-08-18
**[5.0.0] - 2026-08-18**
Breaking:
- Library: *BackupException* is now *BackupError* and *VersionMismatch* is now
*VersionMismatchError*. Both names violated the convention that an exception
class ends in *Error*; the first is imported by six modules, so the rename is
atomic across the package.
- Library: *backup_dumps_for_volume* and *backup_mariadb_or_postgres* return a
*VolumeOutcome* instead of a *(bool, bool)* tuple. The pair could not carry
the detected engine, which the caller needs for the manifest.
New:
- Backup: every generation carries a *manifest.json* stating its layout and,
per volume, *database* (it held one), *dumped* (a dump was produced) and
*engine* (which one was detected). A volume with *database* and no *dumped*
was copied as raw engine files — under *--only-sql* that fallback is the
documented behaviour, and until now nothing in the finished tree said it had
happened. Restoring such a volume replays engine files instead of a dump.
- Library: *baudolo.generation* states the generation layout once — *files*,
*sql*, the dump suffixes, the manifest name. *BackupPaths*, the dump writer
and the volume copier stop spelling them out separately. The module is
import-free on purpose, so a consumer can read a manifest with nothing but
*json* on a host where this package is not installed.
Changed:
- Build: the test targets no longer depend on *clean*. *clean* is
*git clean -fdX .*, so running the unit tests deleted every git-ignored file
in the working tree. It was compensating for a missing *.dockerignore*, which
now keeps *__pycache__*, egg-info and build output out of the image context
where that belongs. *clean* remains available as its own target.
- Lint: a ruff configuration is declared. The package ran on ruff's defaults
while its consumer held itself to a far wider selection; measured against
that selection the tree had 224 findings and now has none. Includes a full
*os.path* to *pathlib* migration, with two deliberate exceptions: *abspath*
stays where *Path.resolve()* would follow symlinks and let a symlinked volume
test as inside the snapshot subject, and the rsync trailing separator is kept
explicit where *Path* would drop it.
## [4.0.0] - 2026-08-17 ## [4.0.0] - 2026-08-17
Breaking: Breaking:

View File

@@ -51,19 +51,19 @@ ruff-fix: install-lint
lint: ruff lint: ruff
# clean + build run once and in order, then lint and the three suites run # build runs once, then lint and the three suites run concurrently via -j4; the
# concurrently via -j4; the *-run targets carry no clean/build prereq so the # *-run targets carry no build prereq so the sub-make cannot race a second build.
# sub-make cannot race a second clean against build. # `clean` is deliberately not a prerequisite; .dockerignore keeps the image
# context clean instead.
test: test:
@$(MAKE) clean
@$(MAKE) build @$(MAKE) build
@$(MAKE) -j4 lint test-unit-run test-integration-run test-e2e-run @$(MAKE) -j4 lint test-unit-run test-integration-run test-e2e-run
test-unit: clean build test-unit-run test-unit: build test-unit-run
test-integration: clean build test-integration-run test-integration: build test-integration-run
test-e2e: clean build test-e2e-run test-e2e: build test-e2e-run
test-unit-run: test-unit-run:
@echo ">> Running unit tests" @echo ">> Running unit tests"

View File

@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
[project] [project]
name = "backup-docker-to-local" name = "backup-docker-to-local"
version = "4.0.0" version = "7.0.0"
description = "Backup Docker volumes to local with rsync and optional DB dumps." description = "Backup Docker volumes to local with rsync and optional DB dumps."
readme = "README.md" readme = "README.md"
requires-python = ">=3.9" requires-python = ">=3.9"
@@ -35,3 +35,52 @@ exclude = ["tests*"]
[tool.setuptools.package-data] [tool.setuptools.package-data]
"baudolo.restore.db" = ["*.sql"] "baudolo.restore.db" = ["*.sql"]
[tool.ruff]
respect-gitignore = true
# The package still declares >=3.9, so pyupgrade must not propose 3.10+ syntax.
target-version = "py39"
exclude = ["build", "dist", "*.egg-info", ".venv", "venv"]
[tool.ruff.lint]
# Adopted from infinito-nexus-core so both repositories are held to one bar;
# see that project's pyproject.toml for what each selector buys.
select = [
"E", "F", "I", "B", "UP", "RUF", "SIM", "C4", "PERF", "RET", "PIE",
"T10", "PGH", "EXE", "RSE", "ICN", "DTZ",
"TID", "LOG", "G",
"S",
"PTH",
"FURB", "W", "FA", "YTT", "A", "ISC", "SLOT", "FLY",
"PYI",
"TC", "N",
"PLE0605",
"PLW1510", "PLW2901", "PLW0108", "PLW0603",
"PLR5501", "PLC0207", "PLR1722", "PLR1714",
"TRY002", "TRY004", "TRY300", "TRY301",
"BLE001",
]
# E501: `ruff format` reflows what it can; the rest is unsplittable literals.
# RUF001/002/003: the prose uses em-dashes deliberately, not homoglyphs.
# S603/S607: this tool's whole job is running `docker` / dump binaries from
# PATH in list form, which is already injection-safe.
# PTH207/PTH208: changing `glob.glob`/`os.listdir` return shapes needs a
# per-call-site review, not a blanket rewrite.
ignore = [
"E501",
"RUF001", "RUF002", "RUF003",
"S603", "S607",
"PTH207", "PTH208",
]
[tool.ruff.lint.per-file-ignores]
# Test code legitimately uses what flake8-bandit flags in production code:
# asserts, dummy credentials, /tmp fixtures, broad excepts in teardown, and
# SQL built from fixture names (S608) to set the databases under test up.
"tests/**" = [
"S101", "S102", "S105", "S106", "S108", "S110", "S112", "S608", "BLE001",
]
# The e2e helpers package is a deliberate re-export aggregator.
"tests/e2e/helpers/__init__.py" = ["F403"]

View File

@@ -2,9 +2,9 @@
from __future__ import annotations from __future__ import annotations
import os
from contextlib import ExitStack from contextlib import ExitStack
from datetime import datetime from datetime import datetime
from pathlib import Path
from .cli import parse_args from .cli import parse_args
from .compose import handle_docker_compose_services from .compose import handle_docker_compose_services
@@ -14,12 +14,13 @@ from .docker import (
docker_volume_names, docker_volume_names,
filter_stoppable, filter_stoppable,
) )
from .dumps import backup_dumps_for_volume, load_databases_df from .dumps import VolumeOutcome, backup_dumps_for_volume, load_databases_df
from .layout import ( from .layout import (
create_version_directory, create_version_directory,
create_volume_directory, create_volume_directory,
get_machine_id, get_machine_id,
stamp_directory, stamp_directory,
write_manifest,
) )
from .policy import requires_stop, volume_is_fully_ignored from .policy import requires_stop, volume_is_fully_ignored
from .snapshot import snapshot_source, volume_snapshot from .snapshot import snapshot_source, volume_snapshot
@@ -34,13 +35,15 @@ def main() -> int:
# order new ones before the existing ones wherever the offset is positive. # order new ones before the existing ones wherever the offset is positive.
backup_time = datetime.now().strftime("%Y%m%d%H%M%S") # noqa: DTZ005 backup_time = datetime.now().strftime("%Y%m%d%H%M%S") # noqa: DTZ005
versions_dir = os.path.join(args.backups_dir, machine_id, args.repo_name) versions_dir = str(Path(args.backups_dir) / machine_id / args.repo_name)
version_dir = create_version_directory(versions_dir, backup_time) version_dir = create_version_directory(versions_dir, backup_time)
databases_df = None if args.only_files else load_databases_df(args.databases_csv) databases_df = None if args.only_files else load_databases_df(args.databases_csv)
print("💾 Start volume backups...", flush=True) print("💾 Start volume backups...", flush=True)
outcomes: dict[str, VolumeOutcome] = {}
with ExitStack() as stack: with ExitStack() as stack:
resolve_source = None resolve_source = None
if args.snapshot: if args.snapshot:
@@ -69,17 +72,18 @@ def main() -> int:
vol_dir = create_volume_directory(version_dir, volume_name) vol_dir = create_volume_directory(version_dir, volume_name)
found_db = dumped_any = False outcome = VolumeOutcome(database=False, dumped=False)
if not args.only_files: if not args.only_files:
found_db, dumped_any = backup_dumps_for_volume( outcome = backup_dumps_for_volume(
containers=containers, containers=containers,
vol_dir=vol_dir, vol_dir=vol_dir,
databases_df=databases_df, databases_df=databases_df,
database_containers=args.database_containers, database_containers=args.database_containers,
) )
outcomes[volume_name] = outcome
if args.only_sql and found_db: if args.only_sql and outcome.database:
if not dumped_any: if not outcome.dumped:
print( print(
f"WARNING: only-sql requested but no DB dump was produced for DB volume '{volume_name}'. " f"WARNING: only-sql requested but no DB dump was produced for DB volume '{volume_name}'. "
"Falling back to file backup.", "Falling back to file backup.",
@@ -129,6 +133,7 @@ def main() -> int:
if not args.shutdown: if not args.shutdown:
change_containers_status(stoppable, "start") change_containers_status(stoppable, "start")
write_manifest(version_dir, outcomes)
stamp_directory(version_dir) stamp_directory(version_dir)
print("Finished volume backups.", flush=True) print("Finished volume backups.", flush=True)

View File

@@ -82,7 +82,7 @@ def handle_docker_compose_services(
continue continue
dir_path = entry.path dir_path = entry.path
name = os.path.basename(dir_path) name = Path(dir_path).name
print(f"Checking directory: {dir_path}", flush=True) print(f"Checking directory: {dir_path}", flush=True)

View File

@@ -1,27 +1,50 @@
from __future__ import annotations from __future__ import annotations
import logging import logging
import os
import pathlib import pathlib
import re import re
from typing import TYPE_CHECKING
import pandas
from baudolo.databases import CLUSTER_ROW, validate_database from baudolo.databases import CLUSTER_ROW, validate_database
from baudolo.generation import CLUSTER_SUFFIX, DUMP_SUFFIX, SQL_DIR
from .docker import docker_exec_argv from .docker import docker_exec_argv
from .shell import BackupException, execute_to_file from .shell import BackupError, execute_to_file
if TYPE_CHECKING:
import pandas as pd
log = logging.getLogger(__name__) log = logging.getLogger(__name__)
ENGINE_NAMES = ("database", "postgres", "mariadb", "mysql", "db")
_SUFFIX_RE = re.compile(rf"(_|-)({'|'.join(ENGINE_NAMES)})")
def get_instance(container: str, database_containers: list[str]) -> str:
""" def get_instance(container: str, database_containers: list[str]) -> str | None:
Derive a stable instance name from the container name. """The databases.csv instance a container serves, or None for no database.
A declared container is its own instance. Every other name is read against
ENGINE_NAMES: carrying one as a suffix makes the rest the instance, which
maps `<app>-database` from compose and `<app>_database.1.<task>` from swarm
onto the same one; being one outright makes the container its own instance,
the shape a compose file writes as `container_name: postgres`.
Args:
container: the running container's name.
database_containers: names passed via --database-containers, taken as
declared engines whatever they are called.
Returns:
The instance name, or None when the name neither carries nor is an
engine name: an application container is not an engine, even when it
ships the client tools that would let a dump command start.
""" """
if container in database_containers: if container in database_containers:
return container return container
return re.split(r"(_|-)(database|db|postgres)", container)[0] parts = _SUFFIX_RE.split(container)
if len(parts) > 1:
return parts[0]
return container if container in ENGINE_NAMES else None
def fallback_pg_dumpall( def fallback_pg_dumpall(
@@ -35,6 +58,7 @@ def fallback_pg_dumpall(
container, container,
["pg_dumpall", "-U", username, "-h", "localhost"], ["pg_dumpall", "-U", username, "-h", "localhost"],
interactive=True, interactive=True,
forward_env=["PGPASSWORD"],
), ),
out_file, out_file,
env={"PGPASSWORD": password}, env={"PGPASSWORD": password},
@@ -47,7 +71,7 @@ def backup_database(
volume_dir: str, volume_dir: str,
db_type: str, db_type: str,
dump_tool: str, dump_tool: str,
databases_df: pandas.DataFrame, databases_df: pd.DataFrame,
database_containers: list[str], database_containers: list[str],
) -> bool: ) -> bool:
""" """
@@ -60,14 +84,17 @@ def backup_database(
Returns True if at least one dump was produced. Returns True if at least one dump was produced.
""" """
instance_name = get_instance(container, database_containers) instance_name = get_instance(container, database_containers)
if instance_name is None:
log.debug("Container '%s' carries no database token", container)
return False
entries = databases_df[databases_df["instance"] == instance_name] entries = databases_df[databases_df["instance"] == instance_name]
if entries.empty: if entries.empty:
log.debug("No database entries for instance '%s'", instance_name) log.debug("No database entries for instance '%s'", instance_name)
return False return False
out_dir = os.path.join(volume_dir, "sql") out_dir = pathlib.Path(volume_dir) / SQL_DIR
pathlib.Path(out_dir).mkdir(parents=True, exist_ok=True) out_dir.mkdir(parents=True, exist_ok=True)
produced = False produced = False
@@ -85,13 +112,13 @@ def backup_database(
f"'{CLUSTER_ROW}' is currently only supported for Postgres." f"'{CLUSTER_ROW}' is currently only supported for Postgres."
) )
cluster_file = os.path.join(out_dir, f"{instance_name}.cluster.backup.sql") cluster_file = str(out_dir / f"{instance_name}{CLUSTER_SUFFIX}")
fallback_pg_dumpall(container, user, password, cluster_file) fallback_pg_dumpall(container, user, password, cluster_file)
produced = True produced = True
continue continue
db_name = db_value db_name = db_value
dump_file = os.path.join(out_dir, f"{db_name}.backup.sql") dump_file = str(out_dir / f"{db_name}{DUMP_SUFFIX}")
if db_type == "mariadb": if db_type == "mariadb":
# Force TCP so auth matches '<user>'@'%' instead of socket -> 'localhost'. # Force TCP so auth matches '<user>'@'%' instead of socket -> 'localhost'.
@@ -131,18 +158,19 @@ def backup_database(
"--no-privileges", "--no-privileges",
], ],
interactive=True, interactive=True,
forward_env=["PGPASSWORD"],
), ),
dump_file, dump_file,
env={"PGPASSWORD": password}, env={"PGPASSWORD": password},
) )
produced = True produced = True
except BackupException as e: except BackupError as e:
raise BackupException( raise BackupError(
f"Postgres dump failed for instance '{instance_name}', " f"Postgres dump failed for instance '{instance_name}', "
f"database '{db_name}'. This database was explicitly configured " f"database '{db_name}'. This database was explicitly configured "
"and therefore must succeed.\n" "and therefore must succeed.\n"
f"{e}" f"{e}"
) ) from e
continue continue
return produced return produced

View File

@@ -1,15 +1,43 @@
from __future__ import annotations from __future__ import annotations
from collections.abc import Sequence from typing import TYPE_CHECKING
from .shell import BackupException, execute_shell_command from .shell import BackupError, execute_shell_command
if TYPE_CHECKING:
from collections.abc import Sequence
def docker_exec_argv( def docker_exec_argv(
container: str, argv: Sequence[str], *, interactive: bool = False container: str,
argv: Sequence[str],
*,
interactive: bool = False,
forward_env: Sequence[str] = (),
) -> list[str]: ) -> list[str]:
"""The argv that runs *argv* inside *container*.""" """The argv that runs *argv* inside *container*.
return ["docker", "exec", *(["-i"] if interactive else []), container, *argv]
Args:
container: the container to run in.
argv: the command, already split.
interactive: keep stdin open, for a command that is fed a dump.
forward_env: names of environment variables to hand to the container.
Passed as bare ``-e NAME``, so docker copies the value out of this
process's own environment; spelling ``-e NAME=value`` instead would
publish a secret in the host's process list.
Returns:
The argv list.
"""
forwarded = [arg for name in forward_env for arg in ("-e", name)]
return [
"docker",
"exec",
*(["-i"] if interactive else []),
*forwarded,
container,
*argv,
]
def get_image_info(container: str) -> str: def get_image_info(container: str) -> str:
@@ -34,7 +62,7 @@ def has_tool(container: str, tool: str) -> bool:
""" """
try: try:
execute_shell_command(docker_exec_argv(container, [tool, "--version"])) execute_shell_command(docker_exec_argv(container, [tool, "--version"]))
except BackupException: except BackupError:
return False return False
return True return True
@@ -74,7 +102,7 @@ def is_swarm_task(container: str) -> bool:
container, container,
] ]
) )
except BackupException: except BackupError:
still_listed = execute_shell_command( still_listed = execute_shell_command(
[ [
"docker", "docker",

View File

@@ -3,13 +3,14 @@
from __future__ import annotations from __future__ import annotations
import sys import sys
from typing import NamedTuple
import pandas import pandas as pd
from pandas.errors import EmptyDataError from pandas.errors import EmptyDataError
from baudolo.databases import COLUMNS, DELIMITER from baudolo.databases import COLUMNS, DELIMITER
from .db import backup_database from .db import backup_database, get_instance
from .docker import has_tool, image_id from .docker import has_tool, image_id
DUMP_TOOLS: tuple[tuple[str, str], ...] = ( DUMP_TOOLS: tuple[tuple[str, str], ...] = (
@@ -21,6 +22,19 @@ DUMP_TOOLS: tuple[tuple[str, str], ...] = (
_ENGINE_BY_IMAGE: dict[str, tuple[str, str] | None] = {} _ENGINE_BY_IMAGE: dict[str, tuple[str, str] | None] = {}
class VolumeOutcome(NamedTuple):
"""What a dump attempt established about one volume.
``database`` says a container serving the volume speaks an engine this
tool can dump; ``dumped`` says a dump was actually written. ``engine`` is
the engine that was detected, or None when none was.
"""
database: bool
dumped: bool
engine: str | None = None
def container_engine(container: str) -> tuple[str, str] | None: def container_engine(container: str) -> tuple[str, str] | None:
"""The (engine, dump tool) a container can serve, or None for neither. """The (engine, dump tool) a container can serve, or None for neither.
@@ -55,15 +69,15 @@ def backup_mariadb_or_postgres(
*, *,
container: str, container: str,
volume_dir: str, volume_dir: str,
databases_df: pandas.DataFrame, databases_df: pd.DataFrame,
database_containers: list[str], database_containers: list[str],
) -> tuple[bool, bool]: ) -> VolumeOutcome:
""" """What this container contributes to its volume's outcome."""
Returns (is_db_container, dumped_any)
"""
engine = container_engine(container) engine = container_engine(container)
if engine is None: if engine is None:
return False, False return VolumeOutcome(database=False, dumped=False)
if get_instance(container, database_containers) is None:
return VolumeOutcome(database=False, dumped=False)
db_type, dump_tool = engine db_type, dump_tool = engine
dumped = backup_database( dumped = backup_database(
container=container, container=container,
@@ -73,20 +87,20 @@ def backup_mariadb_or_postgres(
databases_df=databases_df, databases_df=databases_df,
database_containers=database_containers, database_containers=database_containers,
) )
return True, dumped return VolumeOutcome(database=True, dumped=dumped, engine=db_type)
def _empty_databases_df() -> pandas.DataFrame: def _empty_databases_df() -> pd.DataFrame:
""" """
Create an empty DataFrame with the expected schema for databases.csv. Create an empty DataFrame with the expected schema for databases.csv.
This allows the backup to continue without DB dumps when the CSV is missing This allows the backup to continue without DB dumps when the CSV is missing
or empty (pandas EmptyDataError). or empty (pandas EmptyDataError).
""" """
return pandas.DataFrame(columns=list(COLUMNS)) return pd.DataFrame(columns=list(COLUMNS))
def load_databases_df(csv_path: str) -> pandas.DataFrame: def load_databases_df(csv_path: str) -> pd.DataFrame:
""" """
Load databases.csv robustly. Load databases.csv robustly.
@@ -95,9 +109,7 @@ def load_databases_df(csv_path: str) -> pandas.DataFrame:
- Valid CSV -> return dataframe - Valid CSV -> return dataframe
""" """
try: try:
return pandas.read_csv( return pd.read_csv(csv_path, sep=DELIMITER, keep_default_na=False, dtype=str)
csv_path, sep=DELIMITER, keep_default_na=False, dtype=str
)
except FileNotFoundError: except FileNotFoundError:
print( print(
f"WARNING: databases.csv not found: {csv_path}. Continuing without database dumps.", f"WARNING: databases.csv not found: {csv_path}. Continuing without database dumps.",
@@ -118,25 +130,26 @@ def backup_dumps_for_volume(
*, *,
containers: list[str], containers: list[str],
vol_dir: str, vol_dir: str,
databases_df: pandas.DataFrame, databases_df: pd.DataFrame,
database_containers: list[str], database_containers: list[str],
) -> tuple[bool, bool]: ) -> VolumeOutcome:
""" """The volume's outcome across every container that mounts it."""
Returns (found_db_container, dumped_any)
"""
found_db = False found_db = False
dumped_any = False dumped_any = False
engine: str | None = None
for c in containers: for c in containers:
is_db, dumped = backup_mariadb_or_postgres( outcome = backup_mariadb_or_postgres(
container=c, container=c,
volume_dir=vol_dir, volume_dir=vol_dir,
databases_df=databases_df, databases_df=databases_df,
database_containers=database_containers, database_containers=database_containers,
) )
if is_db: if outcome.database:
found_db = True found_db = True
if dumped: if outcome.dumped:
dumped_any = True dumped_any = True
if engine is None:
engine = outcome.engine
return found_db, dumped_any return VolumeOutcome(database=found_db, dumped=dumped_any, engine=engine)

View File

@@ -2,12 +2,14 @@
from __future__ import annotations from __future__ import annotations
import os import json
import pathlib import pathlib
from dirval import create_stamp_file from dirval import create_stamp_file
from .shell import BackupException, execute_shell_command from baudolo.generation import MANIFEST_FILE, manifest_document
from .shell import BackupError, execute_shell_command
def get_machine_id() -> str: def get_machine_id() -> str:
@@ -22,11 +24,11 @@ def stamp_directory(version_dir: str) -> None:
def create_version_directory(versions_dir: str, backup_time: str) -> str: def create_version_directory(versions_dir: str, backup_time: str) -> str:
version_dir = os.path.join(versions_dir, backup_time) version_dir = str(pathlib.Path(versions_dir) / backup_time)
try: try:
pathlib.Path(version_dir).mkdir(parents=True) pathlib.Path(version_dir).mkdir(parents=True)
except FileExistsError: except FileExistsError:
raise BackupException( raise BackupError(
f"generation {backup_time} already exists at {version_dir}; " f"generation {backup_time} already exists at {version_dir}; "
"another run claimed this second - refusing to write into it, " "another run claimed this second - refusing to write into it, "
"since rsync --delete would overwrite that generation" "since rsync --delete would overwrite that generation"
@@ -35,6 +37,25 @@ def create_version_directory(versions_dir: str, backup_time: str) -> str:
def create_volume_directory(version_dir: str, volume_name: str) -> str: def create_volume_directory(version_dir: str, volume_name: str) -> str:
path = os.path.join(version_dir, volume_name) path = pathlib.Path(version_dir) / volume_name
pathlib.Path(path).mkdir(parents=True, exist_ok=True) path.mkdir(parents=True, exist_ok=True)
return path return str(path)
def write_manifest(version_dir: str, volumes: dict[str, dict[str, bool]]) -> str:
"""Record the generation's layout and per-volume outcome.
Written before the directory is stamped, so the stamp covers it.
Args:
version_dir: the generation directory.
volumes: per volume name, ``database`` and ``dumped``.
Returns:
The path written.
"""
path = pathlib.Path(version_dir) / MANIFEST_FILE
with path.open("w", encoding="utf-8") as handle:
json.dump(manifest_document(volumes), handle, indent=2, sort_keys=True)
handle.write("\n")
return str(path)

View File

@@ -9,10 +9,14 @@ from __future__ import annotations
import os import os
import subprocess import subprocess
from collections.abc import Mapping, Sequence from pathlib import Path
from typing import TYPE_CHECKING
if TYPE_CHECKING:
from collections.abc import Mapping, Sequence
class BackupException(Exception): class BackupError(Exception):
"""Generic exception for backup errors.""" """Generic exception for backup errors."""
@@ -21,7 +25,7 @@ def _child_env(env: Mapping[str, str] | None) -> dict[str, str] | None:
def _fail(command: Sequence[str], returncode: int, out: bytes, err: bytes) -> None: def _fail(command: Sequence[str], returncode: int, out: bytes, err: bytes) -> None:
raise BackupException( raise BackupError(
f"Error in command: {' '.join(command)}\n" f"Error in command: {' '.join(command)}\n"
f"Output: {out}\nError: {err}\n" f"Output: {out}\nError: {err}\n"
f"Exit code: {returncode}" f"Exit code: {returncode}"
@@ -59,13 +63,13 @@ def execute_to_file(
""" """
command = list(command) command = list(command)
print(" ".join(command), flush=True) print(" ".join(command), flush=True)
tmp = f"{out_file}.tmp" tmp = Path(f"{out_file}.tmp")
with open(tmp, "wb") as handle: with tmp.open("wb") as handle:
process = subprocess.Popen( process = subprocess.Popen(
command, stdout=handle, stderr=subprocess.PIPE, env=_child_env(env) command, stdout=handle, stderr=subprocess.PIPE, env=_child_env(env)
) )
_, err = process.communicate() _, err = process.communicate()
if process.returncode != 0: if process.returncode != 0:
os.unlink(tmp) tmp.unlink()
_fail(command, process.returncode, b"", err) _fail(command, process.returncode, b"", err)
os.replace(tmp, out_file) tmp.replace(out_file)

View File

@@ -21,11 +21,16 @@ keeps its snapshot.
from __future__ import annotations from __future__ import annotations
import os import os
from collections.abc import Callable, Iterator
from contextlib import contextmanager from contextlib import contextmanager
from pathlib import Path
from typing import TYPE_CHECKING
from .shell import BackupException, execute_shell_command from .shell import BackupError, execute_shell_command
from .volume import Backing
if TYPE_CHECKING:
from collections.abc import Callable, Iterator
from .volume import Backing
KINDS = ("btrfs", "zfs") KINDS = ("btrfs", "zfs")
@@ -36,10 +41,15 @@ class SnapshotError(RuntimeError):
def _resolver(subject: str, root: str) -> Callable[[str], str]: def _resolver(subject: str, root: str) -> Callable[[str], str]:
def resolve(path: str) -> str: def resolve(path: str) -> str:
relative = os.path.relpath(os.path.abspath(path), os.path.abspath(subject)) # Exception: abspath, not Path.resolve() - resolve() follows symlinks,
# which would let a symlinked volume test as inside the subject.
relative = os.path.relpath(
os.path.abspath(path), # noqa: PTH100
os.path.abspath(subject), # noqa: PTH100
)
if relative.startswith(".."): if relative.startswith(".."):
raise SnapshotError(f"{path} lies outside the snapshot subject {subject}") raise SnapshotError(f"{path} lies outside the snapshot subject {subject}")
resolved = root if relative == "." else os.path.join(root, relative) resolved = root if relative == "." else str(Path(root) / relative)
# abspath drops a trailing separator, and rsync reads "dir/" as its # abspath drops a trailing separator, and rsync reads "dir/" as its
# contents where "dir" means the directory itself. # contents where "dir" means the directory itself.
@@ -54,7 +64,7 @@ def _btrfs(
# The snapshot goes inside the subject, never beside it: the kernel rejects # The snapshot goes inside the subject, never beside it: the kernel rejects
# a snapshot whose destination is on another filesystem, which is exactly # a snapshot whose destination is on another filesystem, which is exactly
# what the parent directory is when the subject is a mountpoint of its own. # what the parent directory is when the subject is a mountpoint of its own.
target = os.path.join(os.path.abspath(subject), f".{name}") target = str(Path(os.path.abspath(subject)) / f".{name}") # noqa: PTH100 - see _resolver
run(["btrfs", "subvolume", "snapshot", "-r", subject, target]) run(["btrfs", "subvolume", "snapshot", "-r", subject, target])
return target, ["btrfs", "subvolume", "delete", target] return target, ["btrfs", "subvolume", "delete", target]
@@ -67,7 +77,7 @@ def _zfs(
if not dataset: if not dataset:
raise SnapshotError(f"no zfs dataset is mounted at {subject}") raise SnapshotError(f"no zfs dataset is mounted at {subject}")
run(["zfs", "snapshot", f"{dataset}@{name}"]) run(["zfs", "snapshot", f"{dataset}@{name}"])
root = os.path.join(subject, ".zfs", "snapshot", name) root = str(Path(subject) / ".zfs" / "snapshot" / name)
return root, ["zfs", "destroy", f"{dataset}@{name}"] return root, ["zfs", "destroy", f"{dataset}@{name}"]
@@ -99,7 +109,9 @@ def unsnapshotted(backing: Backing, subject: str) -> str | None:
if os.path.ismount(real): if os.path.ismount(real):
return f"its mountpoint {backing.mountpoint} sits on its own mount" return f"its mountpoint {backing.mountpoint} sits on its own mount"
try: try:
crosses = os.stat(real).st_dev != os.stat(os.path.realpath(subject)).st_dev crosses = (
Path(real).stat().st_dev != Path(os.path.realpath(subject)).stat().st_dev
)
except OSError as error: except OSError as error:
return f"its mountpoint {backing.mountpoint} could not be read: {error}" return f"its mountpoint {backing.mountpoint} could not be read: {error}"
if crosses: if crosses:
@@ -123,7 +135,7 @@ def snapshot_source(
source = resolve(backing.source) source = resolve(backing.source)
except SnapshotError as error: except SnapshotError as error:
return None, str(error) return None, str(error)
if not os.path.isdir(source): if not Path(source).is_dir():
return None, "it was created after the snapshot was taken" return None, "it was created after the snapshot was taken"
return source, "" return source, ""
@@ -159,6 +171,6 @@ def volume_snapshot(
finally: finally:
try: try:
run(remove) run(remove)
except BackupException as error: except BackupError as error:
# Raising here would also mask whatever the body raised. # Raising here would also mask whatever the body raised.
print(f"WARNING: {root} could not be removed: {error}", flush=True) print(f"WARNING: {root} could not be removed: {error}", flush=True)

View File

@@ -5,7 +5,9 @@ import os
import pathlib import pathlib
from dataclasses import dataclass, field from dataclasses import dataclass, field
from .shell import BackupException, execute_shell_command from baudolo.generation import FILES_DIR
from .shell import BackupError, execute_shell_command
@dataclass(frozen=True) @dataclass(frozen=True)
@@ -45,8 +47,8 @@ def get_last_backup_dir(
) -> str | None: ) -> str | None:
versions = sorted(os.listdir(versions_dir), reverse=True) versions = sorted(os.listdir(versions_dir), reverse=True)
for version in versions: for version in versions:
candidate = os.path.join(versions_dir, version, volume_name, "files", "") candidate = f"{pathlib.Path(versions_dir) / version / volume_name / FILES_DIR}/"
if candidate != current_backup_dir and os.path.isdir(candidate): if candidate != current_backup_dir and pathlib.Path(candidate).is_dir():
return candidate return candidate
return None return None
@@ -69,7 +71,7 @@ def backup_volume(
source: directory to read from - the volume's mountpoint, or its path source: directory to read from - the volume's mountpoint, or its path
inside a snapshot. inside a snapshot.
""" """
dest = os.path.join(volume_dir, "files") + "/" dest = f"{pathlib.Path(volume_dir) / FILES_DIR}/"
pathlib.Path(dest).mkdir(parents=True, exist_ok=True) pathlib.Path(dest).mkdir(parents=True, exist_ok=True)
last = get_last_backup_dir(versions_dir, volume_name, dest) last = get_last_backup_dir(versions_dir, volume_name, dest)
@@ -82,7 +84,7 @@ def backup_volume(
try: try:
execute_shell_command(cmd) execute_shell_command(cmd)
except BackupException as e: except BackupError as e:
if "file has vanished" in str(e): if "file has vanished" in str(e):
print( print(
"Warning: Some files vanished before transfer. Continuing.", flush=True "Warning: Some files vanished before transfer. Continuing.", flush=True

View File

@@ -14,6 +14,7 @@ from __future__ import annotations
import csv import csv
import re import re
from pathlib import Path
from typing import NamedTuple from typing import NamedTuple
COLUMNS = ("instance", "database", "username", "password") COLUMNS = ("instance", "database", "username", "password")
@@ -87,7 +88,7 @@ def read_rows(csv_path: str) -> list[Row]:
DatabasesCsvError: a row holds fewer columns than :data:`COLUMNS`. DatabasesCsvError: a row holds fewer columns than :data:`COLUMNS`.
""" """
rows: list[Row] = [] rows: list[Row] = []
with open(csv_path, newline="", encoding="utf-8") as handle: with Path(csv_path).open(newline="", encoding="utf-8") as handle:
reader = csv.reader(handle, delimiter=DELIMITER) reader = csv.reader(handle, delimiter=DELIMITER)
next(reader, None) next(reader, None)
for raw in reader: for raw in reader:

54
src/baudolo/generation.py Normal file
View File

@@ -0,0 +1,54 @@
"""The on-disk shape of a generation, and the manifest that states it.
Every name a reader needs to find payload in a generation is declared here
once, and written into each generation's own manifest. A consumer therefore
never has to hardcode the layout or match this package's version: it reads
what the run that produced the tree recorded.
The manifest also carries what only the run itself can know: per volume,
``database`` (it held one), ``dumped`` (a dump was produced for it) and
``engine`` (which one was detected). Both flags true is a replayable dump;
``database`` without ``dumped`` is a raw copy of live engine files.
Kept import-free: consumers read the manifest with nothing but ``json``, on
hosts that do not have this package installed.
"""
from __future__ import annotations
FILES_DIR = "files"
SQL_DIR = "sql"
DUMP_SUFFIX = ".backup.sql"
CLUSTER_SUFFIX = ".cluster.backup.sql"
MANIFEST_FILE = "manifest.json"
MANIFEST_SCHEMA = 1
def manifest_document(volumes: dict[str, object]) -> dict[str, object]:
"""The manifest a finished run writes.
Args:
volumes: per volume name, an object carrying ``database``, ``dumped``
and ``engine`` -- a ``baudolo.backup.dumps.VolumeOutcome``.
Returns:
The document, ready for ``json.dump``.
"""
return {
"schema": MANIFEST_SCHEMA,
"layout": {
"files_dir": FILES_DIR,
"sql_dir": SQL_DIR,
"dump_suffix": DUMP_SUFFIX,
"cluster_suffix": CLUSTER_SUFFIX,
},
"volumes": {
name: {
"database": bool(outcome.database),
"dumped": bool(outcome.dumped),
"engine": outcome.engine,
}
for name, outcome in sorted(volumes.items())
},
}

View File

@@ -165,7 +165,7 @@ def main(argv: list[str] | None = None) -> int:
return 0 return 0
parser.error("Unhandled command") parser.error("Unhandled command")
return 2 return 2 # noqa: TRY300 - the try wraps the whole dispatch on purpose
except Exception as e: # noqa: BLE001 - CLI boundary: any failure becomes exit 1 except Exception as e: # noqa: BLE001 - CLI boundary: any failure becomes exit 1
print(f"ERROR: {e}", file=sys.stderr) print(f"ERROR: {e}", file=sys.stderr)

View File

@@ -19,16 +19,20 @@ the implementation:
from __future__ import annotations from __future__ import annotations
import os
import re import re
import tempfile import tempfile
from collections.abc import Iterable, Iterator from pathlib import Path
from typing import TYPE_CHECKING
from baudolo.restore.run import docker_exec
from ..run import docker_exec
from .version import guard from .version import guard
if TYPE_CHECKING:
from collections.abc import Iterable, Iterator
CONTROL_DB = "postgres" CONTROL_DB = "postgres"
_CLUSTER_PRECLEAN_SQL = os.path.join(os.path.dirname(__file__), "cluster_preclean.sql") _CLUSTER_PRECLEAN_SQL = Path(__file__).parent / "cluster_preclean.sql"
_CREATE_ROLE = re.compile(rb'^CREATE ROLE "?([^";]+)"?;\s*$') _CREATE_ROLE = re.compile(rb'^CREATE ROLE "?([^";]+)"?;\s*$')
_CREATE_DATABASE = re.compile(rb"^CREATE DATABASE\s+(.*)$") _CREATE_DATABASE = re.compile(rb"^CREATE DATABASE\s+(.*)$")
_CREATE_ROLE_LINE = re.compile(rb"^CREATE ROLE\s+(.*)$") _CREATE_ROLE_LINE = re.compile(rb"^CREATE ROLE\s+(.*)$")
@@ -92,7 +96,7 @@ def dump_inventory(sql_path: str) -> tuple[list[str], list[str]]:
""" """
databases: list[str] = [] databases: list[str] = []
roles: list[str] = [] roles: list[str] = []
with open(sql_path, "rb") as handle: with Path(sql_path).open("rb") as handle:
for raw in handle: for raw in handle:
line = raw.decode("utf-8", "replace") line = raw.decode("utf-8", "replace")
for pattern, sink, read in ( for pattern, sink, read in (
@@ -111,7 +115,7 @@ def dump_inventory(sql_path: str) -> tuple[list[str], list[str]]:
def preclean_sql() -> str: def preclean_sql() -> str:
"""The catalog-wide pre-clean, safe only behind the instance check.""" """The catalog-wide pre-clean, safe only behind the instance check."""
with open(_CLUSTER_PRECLEAN_SQL, encoding="utf-8") as preclean: with _CLUSTER_PRECLEAN_SQL.open(encoding="utf-8") as preclean:
return preclean.read() return preclean.read()
@@ -158,7 +162,7 @@ def assert_instance_matches_dump(
if foreign: if foreign:
raise RuntimeError( raise RuntimeError(
f"{container} also holds {', '.join(foreign)}, which " f"{container} also holds {', '.join(foreign)}, which "
f"{os.path.basename(sql_path)} does not carry. --empty wipes the " f"{Path(sql_path).name} does not carry. --empty wipes the "
"instance, so those would be destroyed with nothing to restore " "instance, so those would be destroyed with nothing to restore "
"them from. Move them off this instance, or drop them yourself if " "them from. Move them off this instance, or drop them yourself if "
"they are disposable." "they are disposable."
@@ -216,7 +220,7 @@ def restore_cluster_sql(
check_version: refuse a dump from a newer major version than the check_version: refuse a dump from a newer major version than the
running engine before anything is dropped. running engine before anything is dropped.
""" """
if not os.path.isfile(sql_path): if not Path(sql_path).is_file():
raise FileNotFoundError(sql_path) raise FileNotFoundError(sql_path)
if check_version: if check_version:
@@ -239,10 +243,10 @@ def restore_cluster_sql(
docker_env=docker_env, docker_env=docker_env,
) )
with open(sql_path, "rb") as src, tempfile.TemporaryFile() as filtered: with Path(sql_path).open("rb") as src, tempfile.TemporaryFile() as filtered:
for line in filter_own_role_creation(src, user): for line in filter_own_role_creation(src, user):
filtered.write(line) filtered.write(line)
filtered.seek(0) filtered.seek(0)
docker_exec(container, _psql(user), stdin=filtered, docker_env=docker_env) docker_exec(container, _psql(user), stdin=filtered, docker_env=docker_env)
print(f"PostgreSQL cluster restore complete from '{os.path.basename(sql_path)}'.") print(f"PostgreSQL cluster restore complete from '{Path(sql_path).name}'.")

View File

@@ -1,11 +1,14 @@
from __future__ import annotations from __future__ import annotations
import os
import sys import sys
from pathlib import Path
from baudolo.restore.run import docker_exec, docker_exec_sh
from ..run import docker_exec, docker_exec_sh
from .version import guard from .version import guard
_NO_CLIENT = "ERROR: neither 'mariadb' nor 'mysql' found in container."
def _pick_client(container: str) -> str: def _pick_client(container: str) -> str:
""" """
@@ -20,14 +23,13 @@ exit 42
""" """
try: try:
out = docker_exec_sh(container, script, capture=True).stdout.decode().strip() out = docker_exec_sh(container, script, capture=True).stdout.decode().strip()
if not out:
raise RuntimeError("empty client detection output")
return out
except Exception: except Exception:
print( print(_NO_CLIENT, file=sys.stderr)
"ERROR: neither 'mariadb' nor 'mysql' found in container.", file=sys.stderr
)
raise raise
if not out:
print(_NO_CLIENT, file=sys.stderr)
raise RuntimeError("empty client detection output")
return out
def restore_mariadb_sql( def restore_mariadb_sql(
@@ -42,7 +44,7 @@ def restore_mariadb_sql(
) -> None: ) -> None:
client = _pick_client(container) client = _pick_client(container)
if not os.path.isfile(sql_path): if not Path(sql_path).is_file():
raise FileNotFoundError(sql_path) raise FileNotFoundError(sql_path)
if check_version: if check_version:
@@ -66,7 +68,7 @@ def restore_mariadb_sql(
f"--password={password}", f"--password={password}",
"-N", "-N",
"-e", "-e",
f"SELECT table_name FROM information_schema.tables WHERE table_schema = '{db_name}';", f"SELECT table_name FROM information_schema.tables WHERE table_schema = '{db_name}';", # noqa: S608 - validate_database() constrains the name to ^[a-zA-Z0-9_][a-zA-Z0-9_-]*$
], ],
capture=True, capture=True,
) )
@@ -94,7 +96,7 @@ def restore_mariadb_sql(
], ],
) )
with open(sql_path, "rb") as f: with Path(sql_path).open("rb") as f:
docker_exec( docker_exec(
container, [client, "-u", user, f"--password={password}", db_name], stdin=f container, [client, "-u", user, f"--password={password}", db_name], stdin=f
) )

View File

@@ -1,14 +1,18 @@
from __future__ import annotations from __future__ import annotations
import os
import tempfile import tempfile
from collections.abc import Iterable, Iterator from pathlib import Path
from typing import TYPE_CHECKING
from baudolo.restore.run import docker_exec
from ..run import docker_exec
from .version import guard from .version import guard
if TYPE_CHECKING:
from collections.abc import Iterable, Iterator
_SUPERUSER_ONLY_PREFIXES = (b"COMMENT ON EXTENSION", b"ALTER DEFAULT PRIVILEGES") _SUPERUSER_ONLY_PREFIXES = (b"COMMENT ON EXTENSION", b"ALTER DEFAULT PRIVILEGES")
_EMPTY_PRECLEAN_SQL = os.path.join(os.path.dirname(__file__), "empty_preclean.sql") _EMPTY_PRECLEAN_SQL = Path(__file__).parent / "empty_preclean.sql"
def filter_superuser_only_lines(lines: Iterable[bytes]) -> Iterator[bytes]: def filter_superuser_only_lines(lines: Iterable[bytes]) -> Iterator[bytes]:
@@ -49,7 +53,7 @@ def restore_postgres_sql(
empty: bool, empty: bool,
check_version: bool = True, check_version: bool = True,
) -> None: ) -> None:
if not os.path.isfile(sql_path): if not Path(sql_path).is_file():
raise FileNotFoundError(sql_path) raise FileNotFoundError(sql_path)
if check_version: if check_version:
@@ -64,7 +68,7 @@ def restore_postgres_sql(
docker_env = {"PGPASSWORD": password} docker_env = {"PGPASSWORD": password}
if empty: if empty:
with open(_EMPTY_PRECLEAN_SQL, encoding="utf-8") as preclean: with _EMPTY_PRECLEAN_SQL.open(encoding="utf-8") as preclean:
drop_sql = preclean.read() drop_sql = preclean.read()
docker_exec( docker_exec(
container, container,
@@ -76,7 +80,7 @@ def restore_postgres_sql(
# Filter into a spooled temp file instead of building the whole dump in # Filter into a spooled temp file instead of building the whole dump in
# memory: production dumps reach many GB and the previous read/splitlines/ # memory: production dumps reach many GB and the previous read/splitlines/
# join needed roughly three times the dump size in RSS. # join needed roughly three times the dump size in RSS.
with open(sql_path, "rb") as src, tempfile.TemporaryFile() as filtered: with Path(sql_path).open("rb") as src, tempfile.TemporaryFile() as filtered:
for line in filter_superuser_only_lines(src): for line in filter_superuser_only_lines(src):
filtered.write(line) filtered.write(line)
filtered.seek(0) filtered.seek(0)

View File

@@ -24,8 +24,9 @@ with the cluster banner and the roles section, and the first
from __future__ import annotations from __future__ import annotations
import re import re
from pathlib import Path
from ..run import docker_exec, stdout_of from baudolo.restore.run import docker_exec, stdout_of
SCAN_LINES = 2000 SCAN_LINES = 2000
DUMP_VERSION = { DUMP_VERSION = {
@@ -34,7 +35,7 @@ DUMP_VERSION = {
} }
class VersionMismatch(Exception): class VersionMismatchError(Exception):
"""The dump cannot be replayed into this engine.""" """The dump cannot be replayed into this engine."""
@@ -46,11 +47,11 @@ def major_of(version: str) -> int:
``11.8.8-MariaDB-ubu2404``. ``11.8.8-MariaDB-ubu2404``.
Raises: Raises:
VersionMismatch: the string does not start with a number. VersionMismatchError: the string does not start with a number.
""" """
leading = re.match(r"(\d+)", version) leading = re.match(r"(\d+)", version)
if not leading: if not leading:
raise VersionMismatch(f"cannot read a major version from '{version}'") raise VersionMismatchError(f"cannot read a major version from '{version}'")
return int(leading.group(1)) return int(leading.group(1))
@@ -65,10 +66,10 @@ def dump_version(sql_path: str, engine: str) -> str:
The version string as the dump spells it. The version string as the dump spells it.
Raises: Raises:
VersionMismatch: no version line within the first ``SCAN_LINES``. VersionMismatchError: no version line within the first ``SCAN_LINES``.
""" """
pattern = DUMP_VERSION[engine] pattern = DUMP_VERSION[engine]
with open(sql_path, encoding="utf-8", errors="replace") as handle: with Path(sql_path).open(encoding="utf-8", errors="replace") as handle:
for _ in range(SCAN_LINES): for _ in range(SCAN_LINES):
line = handle.readline() line = handle.readline()
if not line: if not line:
@@ -76,7 +77,7 @@ def dump_version(sql_path: str, engine: str) -> str:
found = pattern.search(line) found = pattern.search(line)
if found: if found:
return found.group(1) return found.group(1)
raise VersionMismatch( raise VersionMismatchError(
f"{sql_path} carries no {engine} version header in its first {SCAN_LINES} lines" f"{sql_path} carries no {engine} version header in its first {SCAN_LINES} lines"
) )
@@ -120,10 +121,10 @@ def assert_replayable(sql_path: str, engine: str, dumped: str, serving: str) ->
server rejects and the pre-clean would already have dropped the schema. server rejects and the pre-clean would already have dropped the schema.
Raises: Raises:
VersionMismatch: the dump is newer than the engine. VersionMismatchError: the dump is newer than the engine.
""" """
if major_of(dumped) > major_of(serving): if major_of(dumped) > major_of(serving):
raise VersionMismatch( raise VersionMismatchError(
f"{sql_path} came from {engine} {dumped} but {serving} is running; " f"{sql_path} came from {engine} {dumped} but {serving} is running; "
"a newer dump does not replay into an older engine, and --empty " "a newer dump does not replay into an older engine, and --empty "
"would drop the schema before finding out" "would drop the schema before finding out"

View File

@@ -12,6 +12,7 @@ from __future__ import annotations
import os import os
import sys import sys
from pathlib import Path
from .run import docker_volume_exists, run, stdout_of from .run import docker_volume_exists, run, stdout_of
@@ -21,7 +22,7 @@ INSPECT_FORMAT = (
def restore_volume_files(volume_name: str, backup_files_dir: str) -> int: def restore_volume_files(volume_name: str, backup_files_dir: str) -> int:
if not os.path.isdir(backup_files_dir): if not Path(backup_files_dir).is_dir():
print(f"ERROR: backup files dir not found: {backup_files_dir}", file=sys.stderr) print(f"ERROR: backup files dir not found: {backup_files_dir}", file=sys.stderr)
return 2 return 2
@@ -44,7 +45,7 @@ def restore_volume_files(volume_name: str, backup_files_dir: str) -> int:
) )
return 2 return 2
driver, options = (fields + ["local", "plain"])[1:3] driver, options = ([*fields, "local", "plain"])[1:3]
if (driver != "local" or options == "opts") and not os.path.ismount(mountpoint): if (driver != "local" or options == "opts") and not os.path.ismount(mountpoint):
print( print(
f"ERROR: volume {volume_name} has a backing store of its own " f"ERROR: volume {volume_name} has a backing store of its own "
@@ -55,8 +56,9 @@ def restore_volume_files(volume_name: str, backup_files_dir: str) -> int:
) )
return 2 return 2
src = os.path.join(backup_files_dir, "") # rsync reads "dir/" as its contents and "dir" as the directory itself.
dest = os.path.join(mountpoint, "") src = f"{Path(backup_files_dir)}{os.sep}"
dest = f"{Path(mountpoint)}{os.sep}"
run(["rsync", "-avv", "--delete", src, dest]) run(["rsync", "-avv", "--delete", src, dest])
print("File restore complete.") print("File restore complete.")
return 0 return 0

View File

@@ -1,7 +1,9 @@
from __future__ import annotations from __future__ import annotations
import os
from dataclasses import dataclass from dataclasses import dataclass
from pathlib import Path
from baudolo.generation import CLUSTER_SUFFIX, DUMP_SUFFIX, FILES_DIR, SQL_DIR
@dataclass(frozen=True) @dataclass(frozen=True)
@@ -14,20 +16,20 @@ class BackupPaths:
def root(self) -> str: def root(self) -> str:
# Always build an absolute path under backups_dir # Always build an absolute path under backups_dir
return os.path.join( return str(
self.backups_dir, Path(self.backups_dir)
self.backup_hash, / self.backup_hash
self.repo_name, / self.repo_name
self.version, / self.version
self.volume_name, / self.volume_name
) )
def files_dir(self) -> str: def files_dir(self) -> str:
return os.path.join(self.root(), "files") return str(Path(self.root()) / FILES_DIR)
def sql_file(self, db_name: str) -> str: def sql_file(self, db_name: str) -> str:
return os.path.join(self.root(), "sql", f"{db_name}.backup.sql") return str(Path(self.root()) / SQL_DIR / f"{db_name}{DUMP_SUFFIX}")
def cluster_file(self, instance: str) -> str: def cluster_file(self, instance: str) -> str:
"""The pg_dumpall stream a `database = '*'` row produces.""" """The pg_dumpall stream a `database = '*'` row produces."""
return os.path.join(self.root(), "sql", f"{instance}.cluster.backup.sql") return str(Path(self.root()) / SQL_DIR / f"{instance}{CLUSTER_SUFFIX}")

View File

@@ -1,8 +1,8 @@
from __future__ import annotations from __future__ import annotations
import argparse import argparse
import os
import sys import sys
from pathlib import Path
import pandas as pd import pandas as pd
from pandas.errors import EmptyDataError from pandas.errors import EmptyDataError
@@ -30,7 +30,7 @@ def check_and_add_entry(
""" """
database = validate_database(database, instance=instance) database = validate_database(database, instance=instance)
if os.path.exists(file_path): if Path(file_path).exists():
try: try:
df = pd.read_csv( df = pd.read_csv(
file_path, file_path,

View File

@@ -95,7 +95,7 @@ def write_databases_csv(path: str, rows: list[tuple[str, str, str, str]]) -> Non
database may be '' (empty) to trigger pg_dumpall behavior if you want, but here we use db name. database may be '' (empty) to trigger pg_dumpall behavior if you want, but here we use db name.
""" """
Path(path).parent.mkdir(parents=True, exist_ok=True) Path(path).parent.mkdir(parents=True, exist_ok=True)
with open(path, "w", encoding="utf-8") as f: with Path(path).open("w", encoding="utf-8") as f:
f.write("instance;database;username;password\n") f.write("instance;database;username;password\n")
f.writelines(f"{inst};{db};{user};{pw}\n" for inst, db, user, pw in rows) f.writelines(f"{inst};{db};{user};{pw}\n" for inst, db, user, pw in rows)

View File

@@ -0,0 +1,184 @@
"""An application container that ships the engine's client tools.
This is the shape a dedicated database deploys in: the engine runs as
`<app>-database` while the application itself runs as `<app>`, and neither is
declared through --database-containers, so both names go through the instance
regex. `<app>-database` loses its suffix and lands on the instance `<app>` -
and `<app>` carries no database token at all, so a fallback that returns the
name unchanged lands on that same instance and offers the application container
as a second engine for the same row.
Discourse is the live example: its application container is named `discourse`
by its own launcher and ships pg_dumpall, so a dump command starts there and
writes a file that looks like a backup and holds none of the data.
"""
import json
import unittest
from pathlib import Path
from baudolo.generation import DUMP_SUFFIX, FILES_DIR, MANIFEST_FILE, SQL_DIR
from .helpers import (
POSTGRES_DATA_DIR,
POSTGRES_IMAGE,
backup_path,
backup_run,
cleanup_docker,
create_minimal_compose_dir,
ensure_empty_dir,
latest_version_dir,
require_docker,
run,
unique,
wait_for_postgres,
write_databases_csv,
)
MARKER = "the-application-volume-holds-files"
PAYLOAD = "shop-payload"
class TestE2EAppContainerShipsClientTools(unittest.TestCase):
@classmethod
def setUpClass(cls) -> None:
require_docker()
# uuid4 hex may begin with "db", which the instance regex would split
# on and turn the application container into a different instance,
# hiding exactly the collision this module is about.
cls.prefix = unique("baudolo-e2e-app-tools").replace("-db", "-xb")
cls.backups_dir = f"/tmp/{cls.prefix}/Backups"
ensure_empty_dir(cls.backups_dir)
cls.compose_dir = create_minimal_compose_dir(f"/tmp/{cls.prefix}")
cls.repo_name = cls.prefix
cls.engine = f"{cls.prefix}-shop-database"
cls.app = f"{cls.prefix}-shop"
cls.engine_volume = f"{cls.prefix}-shop-database-vol"
cls.app_volume = f"{cls.prefix}-shop-app-vol"
cls.containers = [cls.engine, cls.app]
cls.volumes = [cls.engine_volume, cls.app_volume]
run(["docker", "volume", "create", cls.engine_volume])
run(["docker", "volume", "create", cls.app_volume])
run(
[
"docker",
"run",
"-d",
"--name",
cls.engine,
"-e",
"POSTGRES_PASSWORD=shoppw",
"-e",
"POSTGRES_DB=shopdb",
"-e",
"POSTGRES_USER=postgres",
"-v",
f"{cls.engine_volume}:{POSTGRES_DATA_DIR}",
POSTGRES_IMAGE,
]
)
run(
[
"docker",
"run",
"-d",
"--name",
cls.app,
"--entrypoint",
"sh",
"-v",
f"{cls.app_volume}:/data",
POSTGRES_IMAGE,
"-c",
f"echo '{MARKER}' > /data/marker.txt && sleep 3600",
]
)
wait_for_postgres(cls.engine, user="postgres", timeout_s=90)
run(
[
"docker",
"exec",
cls.engine,
"sh",
"-lc",
(
'psql -U postgres -d shopdb -c "CREATE TABLE orders (id int, '
f"note text); INSERT INTO orders VALUES (1,'{PAYLOAD}');\""
),
],
check=True,
)
cls.databases_csv = f"/tmp/{cls.prefix}/databases.csv"
write_databases_csv(
cls.databases_csv,
[(cls.app, "shopdb", "postgres", "shoppw")],
)
backup_run(
backups_dir=cls.backups_dir,
repo_name=cls.repo_name,
compose_dir=cls.compose_dir,
databases_csv=cls.databases_csv,
database_containers=["dummy-db"],
images_no_stop_required=[POSTGRES_IMAGE],
)
cls.hash, cls.version = latest_version_dir(cls.backups_dir, cls.repo_name)
@classmethod
def tearDownClass(cls) -> None:
cleanup_docker(containers=cls.containers, volumes=cls.volumes)
def volume_dir(self, volume: str) -> Path:
return backup_path(self.backups_dir, self.repo_name, self.version, volume)
def test_the_engine_volume_was_dumped(self) -> None:
dump = self.volume_dir(self.engine_volume) / SQL_DIR / f"shopdb{DUMP_SUFFIX}"
self.assertTrue(dump.is_file(), f"expected a dump at {dump}")
self.assertIn(PAYLOAD, dump.read_text(encoding="utf-8"))
def test_the_application_volume_produced_no_dump(self) -> None:
"""The collision this module exists for: the application container
answers the same instance as the engine and starts a dump of its own."""
sql_dir = self.volume_dir(self.app_volume) / SQL_DIR
self.assertFalse(
sql_dir.exists(),
f"the application container was dumped: {sorted(sql_dir.iterdir())}"
if sql_dir.exists()
else "",
)
def test_the_application_volume_was_backed_up_as_files(self) -> None:
"""Refusing the dump must not cost the volume its backup."""
marker = self.volume_dir(self.app_volume) / FILES_DIR / "marker.txt"
self.assertTrue(marker.is_file(), f"expected a file backup at {marker}")
self.assertIn(MARKER, marker.read_text(encoding="utf-8"))
def test_the_manifest_does_not_call_the_application_volume_a_database(self) -> None:
manifest = json.loads(
(self.volume_dir(self.app_volume).parent / MANIFEST_FILE).read_text(
encoding="utf-8"
)
)
entry = manifest["volumes"][self.app_volume]
self.assertFalse(entry["database"], entry)
self.assertFalse(entry["dumped"], entry)
def test_the_manifest_records_the_engine_volume_as_dumped(self) -> None:
manifest = json.loads(
(self.volume_dir(self.engine_volume).parent / MANIFEST_FILE).read_text(
encoding="utf-8"
)
)
entry = manifest["volumes"][self.engine_volume]
self.assertTrue(entry["database"], entry)
self.assertTrue(entry["dumped"], entry)
if __name__ == "__main__":
unittest.main()

View File

@@ -21,13 +21,14 @@ are verifying is in the DB-dump stage, so testing backup_database() directly
keeps the assertion focused and the test runnable both on-host and in DinD. keeps the assertion focused and the test runnable both on-host and in DinD.
""" """
import os
import tempfile import tempfile
import unittest import unittest
from pathlib import Path
import pandas import pandas as pd
from baudolo.backup import db as db_mod from baudolo.backup import db as db_mod
from baudolo.generation import DUMP_SUFFIX, SQL_DIR
from .helpers import ( from .helpers import (
MARIADB_DATA_DIR, MARIADB_DATA_DIR,
@@ -140,7 +141,7 @@ class TestE2EMariaDBAnonymousPreemption(unittest.TestCase):
# paths — just the dump that the negative-control proved is failing # paths — just the dump that the negative-control proved is failing
# under the same preemption setup. # under the same preemption setup.
with tempfile.TemporaryDirectory() as volume_dir: with tempfile.TemporaryDirectory() as volume_dir:
df = pandas.DataFrame( df = pd.DataFrame(
[(self.db_container, self.db_name, self.db_user, self.db_password)], [(self.db_container, self.db_name, self.db_user, self.db_password)],
columns=["instance", "database", "username", "password"], columns=["instance", "database", "username", "password"],
) )
@@ -153,9 +154,9 @@ class TestE2EMariaDBAnonymousPreemption(unittest.TestCase):
database_containers=[self.db_container], database_containers=[self.db_container],
) )
self.assertTrue(produced, "backup_database did not produce a dump") self.assertTrue(produced, "backup_database did not produce a dump")
dump_path = os.path.join(volume_dir, "sql", f"{self.db_name}.backup.sql") dump_path = Path(volume_dir) / SQL_DIR / f"{self.db_name}{DUMP_SUFFIX}"
self.assertTrue(os.path.isfile(dump_path), f"expected dump at {dump_path}") self.assertTrue(dump_path.is_file(), f"expected dump at {dump_path}")
with open(dump_path, "r", encoding="utf-8", errors="replace") as f: with dump_path.open(encoding="utf-8", errors="replace") as f:
content = f.read() content = f.read()
self.assertIn("INSERT INTO", content) self.assertIn("INSERT INTO", content)
self.assertIn("'ok'", content) self.assertIn("'ok'", content)

View File

@@ -1,5 +1,8 @@
import json
import unittest import unittest
from baudolo.generation import FILES_DIR, MANIFEST_FILE, MANIFEST_SCHEMA, SQL_DIR
from .helpers import ( from .helpers import (
POSTGRES_DATA_DIR, POSTGRES_DATA_DIR,
POSTGRES_IMAGE, POSTGRES_IMAGE,
@@ -159,6 +162,31 @@ class TestE2EOnlySqlFallbackToFiles(unittest.TestCase):
f"Did not expect SQL dump files, found: {dumps}", f"Did not expect SQL dump files, found: {dumps}",
) )
def manifest(self) -> dict:
generation = backup_path(
self.backups_dir, self.repo_name, self.version, self.pg_volume
).parent
return json.loads((generation / MANIFEST_FILE).read_text(encoding="utf-8"))
def test_the_manifest_records_the_volume_as_a_database_left_undumped(self) -> None:
"""The fallback is invisible in the tree: files/ looks like any copy."""
self.assertEqual(
self.manifest()["volumes"][self.pg_volume],
{"database": True, "dumped": False, "engine": "postgres"},
)
def test_the_manifest_layout_names_where_the_payload_really_landed(self) -> None:
layout = self.manifest()["layout"]
volume_dir = backup_path(
self.backups_dir, self.repo_name, self.version, self.pg_volume
)
self.assertTrue((volume_dir / layout["files_dir"]).is_dir())
self.assertEqual(layout["files_dir"], FILES_DIR)
self.assertEqual(layout["sql_dir"], SQL_DIR)
def test_the_manifest_states_a_schema_a_reader_can_check(self) -> None:
self.assertEqual(self.manifest()["schema"], MANIFEST_SCHEMA)
def test_restored_files_contain_marker(self) -> None: def test_restored_files_contain_marker(self) -> None:
p = run( p = run(
[ [

View File

@@ -1,5 +1,8 @@
import json
import unittest import unittest
from baudolo.generation import MANIFEST_FILE
from .helpers import ( from .helpers import (
POSTGRES_DATA_DIR, POSTGRES_DATA_DIR,
POSTGRES_IMAGE, POSTGRES_IMAGE,
@@ -179,3 +182,21 @@ class TestE2EOnlySqlMixedRun(unittest.TestCase):
(base / "files").exists(), (base / "files").exists(),
f"Expected non-DB volume files backup to exist at: {base / 'files'}", f"Expected non-DB volume files backup to exist at: {base / 'files'}",
) )
def manifest(self) -> dict:
generation = backup_path(
self.backups_dir, self.repo_name, self.version, self.db_volume
).parent
return json.loads((generation / MANIFEST_FILE).read_text(encoding="utf-8"))
def test_the_manifest_records_the_dumped_volume_as_dumped(self) -> None:
self.assertEqual(
self.manifest()["volumes"][self.db_volume],
{"database": True, "dumped": True, "engine": "postgres"},
)
def test_the_manifest_records_the_plain_volume_as_no_database(self) -> None:
self.assertEqual(
self.manifest()["volumes"][self.files_volume],
{"database": False, "dumped": False, "engine": None},
)

View File

@@ -0,0 +1,159 @@
"""An engine whose loopback auth really demands a password.
Every other Postgres scenario runs stock postgres:alpine, whose generated
pg_hba grants trust on 127.0.0.1 and ::1 - so `pg_dump -h localhost` never
needs the password and a dump succeeds whether or not baudolo hands one to the
container. This module makes the password mandatory, which is what a dedicated
engine on a real host does.
"""
import unittest
from pathlib import Path
from baudolo.generation import CLUSTER_SUFFIX, DUMP_SUFFIX, FILES_DIR, SQL_DIR
from .helpers import (
POSTGRES_DATA_DIR,
POSTGRES_IMAGE,
backup_path,
backup_run,
cleanup_docker,
create_minimal_compose_dir,
ensure_empty_dir,
latest_version_dir,
require_docker,
run,
unique,
wait_for_postgres,
write_databases_csv,
)
class TestE2EPostgresPasswordRequired(unittest.TestCase):
@classmethod
def setUpClass(cls) -> None:
require_docker()
cls.prefix = unique("baudolo-e2e-pg-password-required")
cls.backups_dir = f"/tmp/{cls.prefix}/Backups"
ensure_empty_dir(cls.backups_dir)
cls.compose_dir = create_minimal_compose_dir(f"/tmp/{cls.prefix}")
cls.repo_name = cls.prefix
cls.pg_container = f"{cls.prefix}-pg"
cls.pg_volume = f"{cls.prefix}-pg-vol"
cls.containers = [cls.pg_container]
cls.volumes = [cls.pg_volume]
run(["docker", "volume", "create", cls.pg_volume])
run(
[
"docker",
"run",
"-d",
"--name",
cls.pg_container,
"-e",
"POSTGRES_PASSWORD=pgpw",
"-e",
"POSTGRES_DB=appdb",
"-e",
"POSTGRES_USER=postgres",
# The entrypoint evals this into its initdb call, so the host
# lines of pg_hba demand scram while the local socket stays
# trust - the entrypoint's own init and the seeding below keep
# working, and only a TCP connection needs the password.
"-e",
"POSTGRES_INITDB_ARGS=--auth-host=scram-sha-256",
"-v",
f"{cls.pg_volume}:{POSTGRES_DATA_DIR}",
POSTGRES_IMAGE,
]
)
wait_for_postgres(cls.pg_container, user="postgres", timeout_s=90)
run(
[
"docker",
"exec",
cls.pg_container,
"sh",
"-lc",
(
'psql -U postgres -d appdb -c "CREATE TABLE t (id int primary '
"key, v text); INSERT INTO t VALUES (1,'ok');\""
),
],
check=True,
)
cls.unauthenticated = run(
[
"docker",
"exec",
cls.pg_container,
"sh",
"-lc",
"pg_dump -U postgres -d appdb -h localhost",
],
capture=True,
check=False,
)
cls.databases_csv = f"/tmp/{cls.prefix}/databases.csv"
write_databases_csv(
cls.databases_csv,
[
(cls.pg_container, "appdb", "postgres", "pgpw"),
(cls.pg_container, "*", "postgres", "pgpw"),
],
)
backup_run(
backups_dir=cls.backups_dir,
repo_name=cls.repo_name,
compose_dir=cls.compose_dir,
databases_csv=cls.databases_csv,
database_containers=[cls.pg_container],
images_no_stop_required=[POSTGRES_IMAGE],
only_sql=True,
)
cls.hash, cls.version = latest_version_dir(cls.backups_dir, cls.repo_name)
@classmethod
def tearDownClass(cls) -> None:
cleanup_docker(containers=cls.containers, volumes=cls.volumes)
def volume_dir(self) -> Path:
return backup_path(
self.backups_dir, self.repo_name, self.version, self.pg_volume
)
def test_a_dump_without_the_password_is_refused_by_the_server(self) -> None:
"""Without this the module is vacuous: a pg_hba still saying trust would
let a baudolo that forwards nothing pass just as well."""
self.assertNotEqual(self.unauthenticated.returncode, 0)
self.assertIn("no password supplied", self.unauthenticated.stderr or "")
def test_the_configured_database_was_dumped(self) -> None:
dump = self.volume_dir() / SQL_DIR / f"appdb{DUMP_SUFFIX}"
self.assertTrue(dump.is_file(), f"expected a dump at {dump}")
self.assertIn("Dumped by pg_dump", dump.read_text(encoding="utf-8"))
def test_the_dump_carries_the_payload(self) -> None:
"""pg_dump emits table data as COPY ... FROM stdin, so the row reads as
tab-separated values rather than as an INSERT literal."""
dump = self.volume_dir() / SQL_DIR / f"appdb{DUMP_SUFFIX}"
self.assertIn("COPY public.t (id, v) FROM stdin;", dump.read_text("utf-8"))
self.assertIn("1\tok", dump.read_text(encoding="utf-8"))
def test_the_cluster_row_was_dumped_too(self) -> None:
cluster = self.volume_dir() / SQL_DIR / f"{self.pg_container}{CLUSTER_SUFFIX}"
self.assertTrue(cluster.is_file(), f"expected a cluster dump at {cluster}")
self.assertIn("CREATE DATABASE", cluster.read_text(encoding="utf-8"))
def test_only_sql_left_no_file_copy_behind(self) -> None:
self.assertFalse((self.volume_dir() / FILES_DIR).exists())
if __name__ == "__main__":
unittest.main()

View File

@@ -3,7 +3,10 @@
from __future__ import annotations from __future__ import annotations
import shutil import shutil
from pathlib import Path from typing import TYPE_CHECKING
if TYPE_CHECKING:
from pathlib import Path
def touch(p: Path) -> None: def touch(p: Path) -> None:

View File

@@ -1,8 +1,8 @@
import io import io
import os
import tempfile import tempfile
import unittest import unittest
from contextlib import redirect_stderr from contextlib import redirect_stderr
from pathlib import Path
import pandas as pd import pandas as pd
@@ -15,7 +15,7 @@ EXPECTED_COLUMNS = ["instance", "database", "username", "password"]
class TestLoadDatabasesDf(unittest.TestCase): class TestLoadDatabasesDf(unittest.TestCase):
def test_missing_csv_is_handled_with_warning_and_empty_df(self) -> None: def test_missing_csv_is_handled_with_warning_and_empty_df(self) -> None:
with tempfile.TemporaryDirectory() as td: with tempfile.TemporaryDirectory() as td:
missing_path = os.path.join(td, "does-not-exist.csv") missing_path = str(Path(td) / "does-not-exist.csv")
buf = io.StringIO() buf = io.StringIO()
with redirect_stderr(buf): with redirect_stderr(buf):
@@ -31,8 +31,8 @@ class TestLoadDatabasesDf(unittest.TestCase):
def test_empty_csv_is_handled_with_warning_and_empty_df(self) -> None: def test_empty_csv_is_handled_with_warning_and_empty_df(self) -> None:
with tempfile.TemporaryDirectory() as td: with tempfile.TemporaryDirectory() as td:
empty_path = os.path.join(td, "databases.csv") empty_path = Path(td) / "databases.csv"
with open(empty_path, "w", encoding="utf-8") as f: with empty_path.open("w", encoding="utf-8") as f:
f.write("") f.write("")
buf = io.StringIO() buf = io.StringIO()
@@ -49,10 +49,10 @@ class TestLoadDatabasesDf(unittest.TestCase):
def test_valid_csv_loads_without_warning(self) -> None: def test_valid_csv_loads_without_warning(self) -> None:
with tempfile.TemporaryDirectory() as td: with tempfile.TemporaryDirectory() as td:
csv_path = os.path.join(td, "databases.csv") csv_path = Path(td) / "databases.csv"
content = "instance;database;username;password\nmyapp;*;dbuser;secret\n" content = "instance;database;username;password\nmyapp;*;dbuser;secret\n"
with open(csv_path, "w", encoding="utf-8") as f: with csv_path.open("w", encoding="utf-8") as f:
f.write(content) f.write(content)
buf = io.StringIO() buf = io.StringIO()

View File

@@ -0,0 +1,84 @@
"""What main() records in the manifest for each volume it touched."""
from __future__ import annotations
import unittest
from unittest import mock
from baudolo.backup import app
from baudolo.backup.dumps import VolumeOutcome
from baudolo.backup.volume import Backing
from . import REQUIRED_PAIRS
ARGV = ["baudolo", *[arg for pair in REQUIRED_PAIRS for arg in pair]]
def drive(argv: list[str], dump_result: VolumeOutcome) -> dict:
"""Run main() over one volume and return the manifest's volume section.
Args:
argv: the command line under test.
dump_result: what backup_dumps_for_volume reports.
"""
with (
mock.patch("sys.argv", argv),
mock.patch.object(app, "get_machine_id", return_value="machine"),
mock.patch.object(app, "create_version_directory", return_value="/gen"),
mock.patch.object(app, "create_volume_directory", return_value="/gen/vol"),
mock.patch.object(app, "load_databases_df", return_value=None),
mock.patch.object(app, "docker_volume_names", return_value=["pgdata"]),
mock.patch.object(app, "containers_using_volume", return_value=["db"]),
mock.patch.object(app, "volume_is_fully_ignored", return_value=False),
mock.patch.object(app, "backup_dumps_for_volume", return_value=dump_result),
mock.patch.object(app, "inspect_backing", return_value=Backing("/data")),
mock.patch.object(app, "write_manifest") as manifest,
mock.patch.object(app, "stamp_directory"),
mock.patch.object(app, "handle_docker_compose_services"),
mock.patch("os.path.isdir", return_value=True),
mock.patch.object(app, "backup_volume"),
mock.patch.object(app, "filter_stoppable", return_value=[]),
mock.patch.object(app, "requires_stop", return_value=False),
mock.patch.object(app, "change_containers_status"),
):
app.main()
return manifest.call_args.args[1]
class TestManifest(unittest.TestCase):
def test_a_database_volume_without_a_dump_is_recorded_as_undumped(self) -> None:
volumes = drive(
[*ARGV, "--only-sql"],
VolumeOutcome(database=True, dumped=False, engine="postgres"),
)
self.assertEqual(
volumes["pgdata"],
VolumeOutcome(database=True, dumped=False, engine="postgres"),
)
def test_a_dumped_database_volume_is_recorded_as_dumped(self) -> None:
volumes = drive(
[*ARGV, "--only-sql"],
VolumeOutcome(database=True, dumped=True, engine="mariadb"),
)
self.assertEqual(volumes["pgdata"].dumped, True)
self.assertEqual(volumes["pgdata"].engine, "mariadb")
def test_a_plain_volume_is_recorded_as_no_database(self) -> None:
volumes = drive(ARGV, VolumeOutcome(database=False, dumped=False))
self.assertEqual(volumes["pgdata"].database, False)
self.assertIsNone(volumes["pgdata"].engine)
def test_the_dumped_volume_is_recorded_even_though_the_copy_is_skipped(
self,
) -> None:
"""--only-sql returns to the loop head on success, before the copy."""
volumes = drive(
[*ARGV, "--only-sql"],
VolumeOutcome(database=True, dumped=True, engine="postgres"),
)
self.assertIn("pgdata", volumes)
if __name__ == "__main__":
unittest.main()

View File

@@ -34,9 +34,10 @@ def drive(argv: list[str]) -> tuple[list[str], list, list]:
mock.patch.object(app, "volume_is_fully_ignored", return_value=False), mock.patch.object(app, "volume_is_fully_ignored", return_value=False),
mock.patch.object(app, "backup_dumps_for_volume") as dumps, mock.patch.object(app, "backup_dumps_for_volume") as dumps,
mock.patch.object(app, "inspect_backing", return_value=Backing("/data")), mock.patch.object(app, "inspect_backing", return_value=Backing("/data")),
mock.patch.object(app, "write_manifest"),
mock.patch.object(app, "stamp_directory"), mock.patch.object(app, "stamp_directory"),
mock.patch.object(app, "handle_docker_compose_services"), mock.patch.object(app, "handle_docker_compose_services"),
mock.patch.object(app.os.path, "isdir", return_value=True), mock.patch("os.path.isdir", return_value=True),
mock.patch.object(app, "backup_volume", side_effect=record_backup), mock.patch.object(app, "backup_volume", side_effect=record_backup),
mock.patch.object(app, "filter_stoppable", return_value=[]), mock.patch.object(app, "filter_stoppable", return_value=[]),
mock.patch.object(app, "requires_stop", return_value=False), mock.patch.object(app, "requires_stop", return_value=False),

View File

@@ -50,9 +50,10 @@ def drive(*, present: bool = True, reason: str | None = None) -> list[dict]:
return_value=Backing("/var/lib/docker/volumes/vol/_data"), return_value=Backing("/var/lib/docker/volumes/vol/_data"),
), ),
mock.patch.object(snapshot_mod, "unsnapshotted", return_value=reason), mock.patch.object(snapshot_mod, "unsnapshotted", return_value=reason),
mock.patch.object(app, "write_manifest"),
mock.patch.object(app, "stamp_directory"), mock.patch.object(app, "stamp_directory"),
mock.patch.object(app, "handle_docker_compose_services"), mock.patch.object(app, "handle_docker_compose_services"),
mock.patch.object(app.os.path, "isdir", return_value=present), mock.patch("os.path.isdir", return_value=present),
mock.patch.object(app, "backup_volume", side_effect=record), mock.patch.object(app, "backup_volume", side_effect=record),
mock.patch.object(app, "volume_snapshot", stubbed_snapshot), mock.patch.object(app, "volume_snapshot", stubbed_snapshot),
): ):

View File

@@ -47,9 +47,10 @@ def drive() -> tuple[list[str], list[str], list[str]]:
mock.patch.object(app, "volume_is_fully_ignored", return_value=False), mock.patch.object(app, "volume_is_fully_ignored", return_value=False),
mock.patch.object(app, "backup_dumps_for_volume", return_value=(False, False)), mock.patch.object(app, "backup_dumps_for_volume", return_value=(False, False)),
mock.patch.object(app, "inspect_backing", return_value=Backing("/data")), mock.patch.object(app, "inspect_backing", return_value=Backing("/data")),
mock.patch.object(app, "write_manifest"),
mock.patch.object(app, "stamp_directory"), mock.patch.object(app, "stamp_directory"),
mock.patch.object(app, "handle_docker_compose_services"), mock.patch.object(app, "handle_docker_compose_services"),
mock.patch.object(app.os.path, "isdir", return_value=True), mock.patch("os.path.isdir", return_value=True),
mock.patch.object(app, "backup_volume", side_effect=record_backup), mock.patch.object(app, "backup_volume", side_effect=record_backup),
mock.patch.object(app, "filter_stoppable", return_value=[]), mock.patch.object(app, "filter_stoppable", return_value=[]),
mock.patch.object(app, "requires_stop", return_value=False), mock.patch.object(app, "requires_stop", return_value=False),

View File

@@ -2,15 +2,13 @@ import tempfile
import unittest import unittest
from unittest.mock import patch from unittest.mock import patch
import pandas import pandas as pd
from baudolo.backup import db as db_mod from baudolo.backup import db as db_mod
def _df(rows): def _df(rows):
return pandas.DataFrame( return pd.DataFrame(rows, columns=["instance", "database", "username", "password"])
rows, columns=["instance", "database", "username", "password"]
)
def _capture_dumps(*, db_type, rows, container, dump_tool="mariadb-dump"): def _capture_dumps(*, db_type, rows, container, dump_tool="mariadb-dump"):

View File

@@ -0,0 +1,61 @@
"""How a secret reaches the command running inside the container."""
from __future__ import annotations
import unittest
from baudolo.backup.db import fallback_pg_dumpall
from baudolo.backup.docker import docker_exec_argv
class TestForwardEnv(unittest.TestCase):
def test_nothing_is_added_when_no_variable_is_named(self) -> None:
self.assertEqual(
docker_exec_argv("c1", ["true"]),
["docker", "exec", "c1", "true"],
)
def test_a_named_variable_is_forwarded_without_its_value(self) -> None:
"""-e NAME=value would publish the secret in the host's process list."""
argv = docker_exec_argv("c1", ["true"], forward_env=["PGPASSWORD"])
self.assertEqual(argv, ["docker", "exec", "-e", "PGPASSWORD", "c1", "true"])
def test_the_flag_precedes_the_container(self) -> None:
"""docker reads options before the container name, arguments after it."""
argv = docker_exec_argv(
"c1", ["pg_dump", "-U", "u"], interactive=True, forward_env=["PGPASSWORD"]
)
self.assertLess(argv.index("-e"), argv.index("c1"))
self.assertLess(argv.index("-i"), argv.index("c1"))
self.assertGreater(argv.index("pg_dump"), argv.index("c1"))
def test_several_variables_each_get_their_own_flag(self) -> None:
argv = docker_exec_argv("c1", ["true"], forward_env=["A", "B"])
self.assertEqual(argv[:6], ["docker", "exec", "-e", "A", "-e", "B"])
class TestPostgresDumpCarriesThePassword(unittest.TestCase):
def test_the_cluster_dump_forwards_pgpassword(self) -> None:
seen: dict = {}
def fake(command, out_file, *, env=None):
seen["command"] = command
seen["env"] = env
import baudolo.backup.db as db
original = db.execute_to_file
db.execute_to_file = fake
try:
fallback_pg_dumpall("pg", "user", "secret", "/tmp/out.sql")
finally:
db.execute_to_file = original
self.assertIn("-e", seen["command"])
self.assertEqual(seen["command"][seen["command"].index("-e") + 1], "PGPASSWORD")
self.assertEqual(seen["env"], {"PGPASSWORD": "secret"})
self.assertNotIn("secret", seen["command"])
if __name__ == "__main__":
unittest.main()

View File

@@ -2,7 +2,7 @@ import unittest
from unittest.mock import patch from unittest.mock import patch
from baudolo.backup import docker as docker_mod from baudolo.backup import docker as docker_mod
from baudolo.backup.shell import BackupException from baudolo.backup.shell import BackupError
class TestIsSwarmTask(unittest.TestCase): class TestIsSwarmTask(unittest.TestCase):
@@ -21,7 +21,7 @@ class TestIsSwarmTask(unittest.TestCase):
@patch.object( @patch.object(
docker_mod, docker_mod,
"execute_shell_command", "execute_shell_command",
side_effect=[BackupException("gone"), []], side_effect=[BackupError("gone"), []],
) )
def test_vanished_container_counts_as_not_stoppable(self, _mock) -> None: def test_vanished_container_counts_as_not_stoppable(self, _mock) -> None:
# A container removed between listing and inspect must not abort the # A container removed between listing and inspect must not abort the
@@ -32,13 +32,13 @@ class TestIsSwarmTask(unittest.TestCase):
@patch.object( @patch.object(
docker_mod, docker_mod,
"execute_shell_command", "execute_shell_command",
side_effect=[BackupException("daemon hiccup"), ["still-here"]], side_effect=[BackupError("daemon hiccup"), ["still-here"]],
) )
def test_inspect_failure_on_existing_container_still_fails(self, _mock) -> None: def test_inspect_failure_on_existing_container_still_fails(self, _mock) -> None:
# If the container still exists, an inspect failure must keep failing # If the container still exists, an inspect failure must keep failing
# the run: silently skipping the stop would back up a hot volume and # the run: silently skipping the stop would back up a hot volume and
# report green without the stop guarantee. # report green without the stop guarantee.
with self.assertRaises(BackupException): with self.assertRaises(BackupError):
docker_mod.is_swarm_task("still-here") docker_mod.is_swarm_task("still-here")

View File

@@ -2,7 +2,7 @@ import unittest
from unittest.mock import patch from unittest.mock import patch
from baudolo.backup import docker as docker_mod from baudolo.backup import docker as docker_mod
from baudolo.backup.shell import BackupException from baudolo.backup.shell import BackupError
class TestImageId(unittest.TestCase): class TestImageId(unittest.TestCase):
@@ -20,7 +20,7 @@ class TestHasTool(unittest.TestCase):
def test_a_tool_that_exits_non_zero_is_absent(self) -> None: def test_a_tool_that_exits_non_zero_is_absent(self) -> None:
with patch.object( with patch.object(
docker_mod, "execute_shell_command", side_effect=BackupException("127") docker_mod, "execute_shell_command", side_effect=BackupError("127")
): ):
self.assertFalse(docker_mod.has_tool("c1", "mariadb-dump")) self.assertFalse(docker_mod.has_tool("c1", "mariadb-dump"))

View File

@@ -1,15 +1,13 @@
import unittest import unittest
from unittest.mock import patch from unittest.mock import patch
import pandas import pandas as pd
from baudolo.backup import dumps as dumps_mod from baudolo.backup import dumps as dumps_mod
def _df(rows): def _df(rows):
return pandas.DataFrame( return pd.DataFrame(rows, columns=["instance", "database", "username", "password"])
rows, columns=["instance", "database", "username", "password"]
)
class _Probe: class _Probe:
@@ -104,15 +102,16 @@ class TestBackupDispatch(unittest.TestCase):
patch.object(dumps_mod, "image_id", probe.image_id), patch.object(dumps_mod, "image_id", probe.image_id),
patch.object(dumps_mod, "backup_database", _fake_backup_database), patch.object(dumps_mod, "backup_database", _fake_backup_database),
): ):
is_db, dumped = dumps_mod.backup_mariadb_or_postgres( outcome = dumps_mod.backup_mariadb_or_postgres(
container="c1", container="c1",
volume_dir="/tmp", volume_dir="/tmp",
databases_df=_df([("c1", "appdb", "u", "p")]), databases_df=_df([("c1", "appdb", "u", "p")]),
database_containers=["c1"], database_containers=["c1"],
) )
self.assertTrue(is_db) self.assertTrue(outcome.database)
self.assertTrue(dumped) self.assertTrue(outcome.dumped)
self.assertEqual(outcome.engine, "mariadb")
self.assertEqual(seen["db_type"], "mariadb") self.assertEqual(seen["db_type"], "mariadb")
self.assertEqual(seen["dump_tool"], "mysqldump") self.assertEqual(seen["dump_tool"], "mysqldump")
@@ -130,7 +129,7 @@ class TestBackupDispatch(unittest.TestCase):
databases_df=_df([]), databases_df=_df([]),
database_containers=[], database_containers=[],
), ),
(False, False), dumps_mod.VolumeOutcome(database=False, dumped=False, engine=None),
) )

View File

@@ -0,0 +1,80 @@
"""Which databases.csv instance a container name resolves to.
The cases are the container names real deployments produce, in both compose
and swarm, so a change to the regex has to state which shape it gives up.
"""
from __future__ import annotations
import unittest
from baudolo.backup.db import get_instance
class TestDeclaredContainers(unittest.TestCase):
def test_a_declared_container_is_its_own_instance(self) -> None:
self.assertEqual(
get_instance("postgres-central", ["postgres-central"]), "postgres-central"
)
def test_a_declaration_beats_the_regex(self) -> None:
"""A declared name is taken whole even when it carries a token the
fallback would otherwise strip."""
self.assertEqual(
get_instance("shop-database", ["shop-database"]), "shop-database"
)
def test_a_qualified_central_name_still_has_to_be_declared(self) -> None:
self.assertIsNone(get_instance("postgres-central", []))
class TestContainersNamedAfterTheirEngine(unittest.TestCase):
def test_a_bare_engine_name_is_its_own_instance(self) -> None:
for name in ("postgres", "mariadb", "mysql", "db", "database"):
with self.subTest(container=name):
self.assertEqual(get_instance(name, []), name)
def test_a_swarm_task_of_such_a_container_keeps_the_instance(self) -> None:
self.assertEqual(get_instance("postgres_postgres.1.k3f9x2", []), "postgres")
self.assertEqual(get_instance("mariadb_mariadb.1.k3f9x2", []), "mariadb")
class TestDedicatedEngines(unittest.TestCase):
def test_compose_names_the_container_with_a_hyphen(self) -> None:
self.assertEqual(get_instance("discourse-database", []), "discourse")
def test_swarm_names_the_task_with_an_underscore_and_a_slot(self) -> None:
"""Swarm suppresses container_name and names the task
<stack>_<service>.<slot>.<id>, which must land on the same instance as
the compose name so one databases.csv serves both modes."""
self.assertEqual(get_instance("discourse_database.1.k3f9x2", []), "discourse")
def test_an_explicitly_named_engine_keeps_its_entity(self) -> None:
self.assertEqual(get_instance("bigbluebutton-postgres-1", []), "bigbluebutton")
def test_the_short_token_is_stripped_too(self) -> None:
self.assertEqual(get_instance("matomo-db", []), "matomo")
def test_mariadb_uses_the_same_suffix(self) -> None:
self.assertEqual(get_instance("matomo-database", []), "matomo")
def test_an_engine_named_suffix_is_stripped_too(self) -> None:
self.assertEqual(get_instance("shop-mariadb", []), "shop")
self.assertEqual(get_instance("shop-mysql", []), "shop")
class TestApplicationContainers(unittest.TestCase):
def test_a_bare_application_name_is_not_a_database(self) -> None:
"""Returning the name unchanged here would offer the application as a
second engine for its own dedicated database's instance."""
self.assertIsNone(get_instance("discourse", []))
def test_a_swarm_application_task_is_not_a_database(self) -> None:
self.assertIsNone(get_instance("discourse_discourse.1.k3f9x2", []))
def test_an_application_that_merely_starts_with_a_token_is_not_split(self) -> None:
self.assertIsNone(get_instance("dbeaver", []))
if __name__ == "__main__":
unittest.main()

View File

@@ -8,7 +8,7 @@ from pathlib import Path
from unittest import mock from unittest import mock
from baudolo.backup import layout as mod from baudolo.backup import layout as mod
from baudolo.backup.shell import BackupException from baudolo.backup.shell import BackupError
class TestVersionDirectory(unittest.TestCase): class TestVersionDirectory(unittest.TestCase):
@@ -21,7 +21,7 @@ class TestVersionDirectory(unittest.TestCase):
def test_it_refuses_a_generation_another_run_already_claimed(self) -> None: def test_it_refuses_a_generation_another_run_already_claimed(self) -> None:
with tempfile.TemporaryDirectory() as tmp: with tempfile.TemporaryDirectory() as tmp:
mod.create_version_directory(tmp, "20260731") mod.create_version_directory(tmp, "20260731")
with self.assertRaises(BackupException) as caught: with self.assertRaises(BackupError) as caught:
mod.create_version_directory(tmp, "20260731") mod.create_version_directory(tmp, "20260731")
self.assertIn("20260731", str(caught.exception)) self.assertIn("20260731", str(caught.exception))

View File

@@ -4,7 +4,7 @@ from __future__ import annotations
import unittest import unittest
from baudolo.backup.shell import BackupException from baudolo.backup.shell import BackupError
from baudolo.backup.snapshot import SnapshotError, volume_snapshot from baudolo.backup.snapshot import SnapshotError, volume_snapshot
@@ -143,7 +143,7 @@ class TestRejections(unittest.TestCase):
class Busy(Runner): class Busy(Runner):
def __call__(self, command: list[str]) -> list[str]: def __call__(self, command: list[str]) -> list[str]:
if command[:3] == ["btrfs", "subvolume", "delete"]: if command[:3] == ["btrfs", "subvolume", "delete"]:
raise BackupException("target is busy") raise BackupError("target is busy")
return super().__call__(command) return super().__call__(command)

View File

@@ -10,6 +10,7 @@ from __future__ import annotations
import os import os
import tempfile import tempfile
import unittest import unittest
from pathlib import Path
from unittest import mock from unittest import mock
from baudolo.backup.snapshot import SnapshotError, snapshot_source, unsnapshotted from baudolo.backup.snapshot import SnapshotError, snapshot_source, unsnapshotted
@@ -19,8 +20,8 @@ from baudolo.backup.volume import Backing
class TestUnsnapshotted(unittest.TestCase): class TestUnsnapshotted(unittest.TestCase):
def setUp(self) -> None: def setUp(self) -> None:
self.subject = tempfile.mkdtemp() self.subject = tempfile.mkdtemp()
self.mountpoint = os.path.join(self.subject, "volumes", "app", "_data") self.mountpoint = str(Path(self.subject) / "volumes" / "app" / "_data")
os.makedirs(self.mountpoint) Path(self.mountpoint).mkdir(parents=True)
def backing(self, **kwargs) -> Backing: def backing(self, **kwargs) -> Backing:
return Backing(kwargs.pop("mountpoint", self.mountpoint), **kwargs) return Backing(kwargs.pop("mountpoint", self.mountpoint), **kwargs)
@@ -71,7 +72,7 @@ class TestUnsnapshotted(unittest.TestCase):
def test_an_unreadable_mountpoint_is_not(self) -> None: def test_an_unreadable_mountpoint_is_not(self) -> None:
reason = unsnapshotted( reason = unsnapshotted(
self.backing(mountpoint=os.path.join(self.subject, "gone")), self.subject self.backing(mountpoint=str(Path(self.subject) / "gone")), self.subject
) )
self.assertIn("could not be read", reason) self.assertIn("could not be read", reason)
@@ -79,12 +80,12 @@ class TestUnsnapshotted(unittest.TestCase):
class TestSnapshotSource(unittest.TestCase): class TestSnapshotSource(unittest.TestCase):
def setUp(self) -> None: def setUp(self) -> None:
self.subject = tempfile.mkdtemp() self.subject = tempfile.mkdtemp()
self.mountpoint = os.path.join(self.subject, "volumes", "app", "_data") self.mountpoint = str(Path(self.subject) / "volumes" / "app" / "_data")
os.makedirs(self.mountpoint) Path(self.mountpoint).mkdir(parents=True)
self.snapshot = os.path.join( self.snapshot = str(
self.subject, ".baudolo-tag", "volumes", "app", "_data" Path(self.subject) / ".baudolo-tag" / "volumes" / "app" / "_data"
) )
os.makedirs(self.snapshot) Path(self.snapshot).mkdir(parents=True)
self.backing = Backing(self.mountpoint) self.backing = Backing(self.mountpoint)
def test_a_captured_volume_reads_from_the_snapshot(self) -> None: def test_a_captured_volume_reads_from_the_snapshot(self) -> None:
@@ -114,7 +115,7 @@ class TestSnapshotSource(unittest.TestCase):
def test_a_volume_created_after_the_snapshot_degrades(self) -> None: def test_a_volume_created_after_the_snapshot_degrades(self) -> None:
source, reason = snapshot_source( source, reason = snapshot_source(
lambda path: os.path.join(self.subject, "absent") + "/", lambda path: str(Path(self.subject) / "absent") + "/",
self.backing, self.backing,
self.subject, self.subject,
) )

View File

@@ -1,6 +1,6 @@
import os
import tempfile import tempfile
import unittest import unittest
from pathlib import Path
from unittest.mock import MagicMock, patch from unittest.mock import MagicMock, patch
from baudolo.restore.db import cluster as cluster_mod from baudolo.restore.db import cluster as cluster_mod
@@ -53,10 +53,10 @@ def cluster_header(roles: int) -> str:
def dump_file(text: str) -> str: def dump_file(text: str) -> str:
path = os.path.join(tempfile.mkdtemp(), "app.backup.sql") path = Path(tempfile.mkdtemp()) / "app.backup.sql"
with open(path, "w", encoding="utf-8") as handle: with path.open("w", encoding="utf-8") as handle:
handle.write(text) handle.write(text)
return path return str(path)
class TestDumpVersion(unittest.TestCase): class TestDumpVersion(unittest.TestCase):
@@ -78,19 +78,19 @@ class TestDumpVersion(unittest.TestCase):
def test_cluster_dump_states_its_version_far_below_the_header(self) -> None: def test_cluster_dump_states_its_version_far_below_the_header(self) -> None:
path = dump_file(cluster_header(roles=200)) path = dump_file(cluster_header(roles=200))
with open(path, encoding="utf-8") as handle: with Path(path).open(encoding="utf-8") as handle:
offset = next(i for i, line in enumerate(handle) if "Dumped from" in line) offset = next(i for i, line in enumerate(handle) if "Dumped from" in line)
self.assertGreater(offset, 100, "fixture must exercise the deep scan") self.assertGreater(offset, 100, "fixture must exercise the deep scan")
self.assertEqual(ver.dump_version(path, "postgres"), "17.11") self.assertEqual(ver.dump_version(path, "postgres"), "17.11")
def test_a_version_beyond_the_scan_limit_is_refused_not_ignored(self) -> None: def test_a_version_beyond_the_scan_limit_is_refused_not_ignored(self) -> None:
path = dump_file(cluster_header(roles=ver.SCAN_LINES)) path = dump_file(cluster_header(roles=ver.SCAN_LINES))
with self.assertRaises(ver.VersionMismatch): with self.assertRaises(ver.VersionMismatchError):
ver.dump_version(path, "postgres") ver.dump_version(path, "postgres")
def test_a_dump_without_a_version_header_is_refused(self) -> None: def test_a_dump_without_a_version_header_is_refused(self) -> None:
path = dump_file("CREATE TABLE t (id int);\n") path = dump_file("CREATE TABLE t (id int);\n")
with self.assertRaises(ver.VersionMismatch): with self.assertRaises(ver.VersionMismatchError):
ver.dump_version(path, "postgres") ver.dump_version(path, "postgres")
@@ -102,13 +102,13 @@ class TestMajorOf(unittest.TestCase):
self.assertEqual(ver.major_of("18beta1"), 18) self.assertEqual(ver.major_of("18beta1"), 18)
def test_refuses_an_unreadable_version(self) -> None: def test_refuses_an_unreadable_version(self) -> None:
with self.assertRaises(ver.VersionMismatch): with self.assertRaises(ver.VersionMismatchError):
ver.major_of("unknown") ver.major_of("unknown")
class TestAssertReplayable(unittest.TestCase): class TestAssertReplayable(unittest.TestCase):
def test_newer_dump_into_older_engine_is_refused(self) -> None: def test_newer_dump_into_older_engine_is_refused(self) -> None:
with self.assertRaises(ver.VersionMismatch) as caught: with self.assertRaises(ver.VersionMismatchError) as caught:
ver.assert_replayable("/b/app.sql", "postgres", "17.11", "15.6") ver.assert_replayable("/b/app.sql", "postgres", "17.11", "15.6")
self.assertIn("17.11", str(caught.exception)) self.assertIn("17.11", str(caught.exception))
self.assertIn("15.6", str(caught.exception)) self.assertIn("15.6", str(caught.exception))
@@ -154,7 +154,7 @@ class TestGateStopsBeforeDestroying(unittest.TestCase):
path = dump_file(POSTGRES_HEADER) path = dump_file(POSTGRES_HEADER)
with ( with (
patch.object(pg_mod, "docker_exec") as replay, patch.object(pg_mod, "docker_exec") as replay,
self.assertRaises(ver.VersionMismatch), self.assertRaises(ver.VersionMismatchError),
): ):
pg_mod.restore_postgres_sql( pg_mod.restore_postgres_sql(
container="db", container="db",
@@ -171,7 +171,7 @@ class TestGateStopsBeforeDestroying(unittest.TestCase):
path = dump_file(cluster_header(roles=3)) path = dump_file(cluster_header(roles=3))
with ( with (
patch.object(cluster_mod, "docker_exec") as replay, patch.object(cluster_mod, "docker_exec") as replay,
self.assertRaises(ver.VersionMismatch), self.assertRaises(ver.VersionMismatchError),
): ):
cluster_mod.restore_cluster_sql( cluster_mod.restore_cluster_sql(
container="db", container="db",
@@ -188,7 +188,7 @@ class TestGateStopsBeforeDestroying(unittest.TestCase):
with ( with (
patch.object(mdb_mod, "_pick_client", return_value="mariadb"), patch.object(mdb_mod, "_pick_client", return_value="mariadb"),
patch.object(mdb_mod, "docker_exec") as replay, patch.object(mdb_mod, "docker_exec") as replay,
self.assertRaises(ver.VersionMismatch), self.assertRaises(ver.VersionMismatchError),
): ):
mdb_mod.restore_mariadb_sql( mdb_mod.restore_mariadb_sql(
container="db", container="db",
@@ -235,7 +235,7 @@ class TestGateStopsBeforeDestroying(unittest.TestCase):
db_name="app", db_name="app",
user="app", user="app",
password="pw", password="pw",
sql_path=os.path.join(tempfile.mkdtemp(), "absent.sql"), sql_path=str(Path(tempfile.mkdtemp()) / "absent.sql"),
empty=True, empty=True,
) )

View File

@@ -51,7 +51,7 @@ class TestSeedMain(unittest.TestCase):
return df return df
@patch("baudolo.seed.__main__.os.path.exists", return_value=False) @patch("baudolo.seed.__main__.Path.exists", return_value=False)
@patch("baudolo.seed.__main__.pd.read_csv") @patch("baudolo.seed.__main__.pd.read_csv")
@patch("baudolo.seed.__main__._empty_df") @patch("baudolo.seed.__main__._empty_df")
@patch("baudolo.seed.__main__.pd.concat") @patch("baudolo.seed.__main__.pd.concat")
@@ -83,7 +83,7 @@ class TestSeedMain(unittest.TestCase):
"/tmp/databases.csv", sep=";", index=False "/tmp/databases.csv", sep=";", index=False
) )
@patch("baudolo.seed.__main__.os.path.exists", return_value=True) @patch("baudolo.seed.__main__.Path.exists", return_value=True)
@patch("baudolo.seed.__main__.pd.read_csv", side_effect=EmptyDataError("empty")) @patch("baudolo.seed.__main__.pd.read_csv", side_effect=EmptyDataError("empty"))
@patch("baudolo.seed.__main__._empty_df") @patch("baudolo.seed.__main__._empty_df")
@patch("baudolo.seed.__main__.pd.concat") @patch("baudolo.seed.__main__.pd.concat")
@@ -110,8 +110,8 @@ class TestSeedMain(unittest.TestCase):
password="pass", password="pass",
) )
exists.assert_called_once_with("/tmp/databases.csv") exists.assert_called_once_with()
read_csv.assert_called_once() self.assertEqual(read_csv.call_args.args, ("/tmp/databases.csv",))
empty_df.assert_called_once() empty_df.assert_called_once()
concat.assert_called_once() concat.assert_called_once()
@@ -133,7 +133,7 @@ class TestSeedMain(unittest.TestCase):
"/tmp/databases.csv", sep=";", index=False "/tmp/databases.csv", sep=";", index=False
) )
@patch("baudolo.seed.__main__.os.path.exists", return_value=True) @patch("baudolo.seed.__main__.Path.exists", return_value=True)
@patch("baudolo.seed.__main__.pd.read_csv") @patch("baudolo.seed.__main__.pd.read_csv")
def test_check_and_add_entry_updates_existing_row( def test_check_and_add_entry_updates_existing_row(
self, self,

View File

@@ -0,0 +1,65 @@
"""Contract of the generation manifest document and the file it lands in."""
from __future__ import annotations
import json
import tempfile
import unittest
from pathlib import Path
from baudolo.backup.dumps import VolumeOutcome
from baudolo.backup.layout import write_manifest
from baudolo.generation import (
CLUSTER_SUFFIX,
DUMP_SUFFIX,
FILES_DIR,
MANIFEST_FILE,
MANIFEST_SCHEMA,
SQL_DIR,
manifest_document,
)
class TestManifestDocument(unittest.TestCase):
def test_it_states_the_layout_a_reader_needs(self) -> None:
document = manifest_document({})
self.assertEqual(
document["layout"],
{
"files_dir": FILES_DIR,
"sql_dir": SQL_DIR,
"dump_suffix": DUMP_SUFFIX,
"cluster_suffix": CLUSTER_SUFFIX,
},
)
def test_it_carries_a_schema_so_a_reader_can_refuse_a_newer_one(self) -> None:
self.assertEqual(manifest_document({})["schema"], MANIFEST_SCHEMA)
def test_it_sorts_volumes_so_two_runs_produce_the_same_bytes(self) -> None:
state = VolumeOutcome(database=False, dumped=False)
document = manifest_document({"b": state, "a": state})
self.assertEqual(list(document["volumes"]), ["a", "b"])
class TestWriteManifest(unittest.TestCase):
def test_it_writes_readable_json_next_to_the_volumes(self) -> None:
with tempfile.TemporaryDirectory() as version_dir:
path = write_manifest(
version_dir,
{
"pgdata": VolumeOutcome(
database=True, dumped=False, engine="postgres"
)
},
)
self.assertEqual(Path(path).name, MANIFEST_FILE)
document = json.loads(Path(path).read_text(encoding="utf-8"))
self.assertEqual(
document["volumes"]["pgdata"],
{"database": True, "dumped": False, "engine": "postgres"},
)
if __name__ == "__main__":
unittest.main()