Files
docker-volume-backup/tests/e2e/snapshot_driver.py
Kevin Veen-Birkenbach d4317827bd feat(backup): capture volumes from a filesystem snapshot
Backing up a live volume with rsync copies a moving target: a database
written to mid-copy lands on disk in a state no engine ever committed.
Stopping the container avoids that at the cost of downtime.

A snapshot removes both. `--snapshot {btrfs,zfs}` with `--snapshot-subject`
freezes the docker root once per run, and every volume copy is then read
from that frozen tree while the containers keep serving. A restore of such
a copy is an ordinary crash recovery, which every supported engine performs
on its own at startup.

An unsupported filesystem or an unknown snapshot kind fails loudly rather
than degrading to a live copy, since a silent fallback would return exactly
the torn backup the mode exists to prevent. `--shutdown` is rejected
alongside `--snapshot` instead of being ignored: under a snapshot no
container is ever stopped, so accepting the flag would promise downtime
semantics the run does not deliver.

Copies out of a snapshot skip rsync's --checksum verification. The source
is immutable for the lifetime of the copy, so size-and-mtime cannot race,
and dropping the second full read roughly halves the I/O per volume.

backup/app.py grew past what one module could carry and is split into
layout, policy and dumps along the lines it already had internally.

Tests: unit coverage for the new snapshot, layout, policy, volume and cli
units; e2e cases drive real btrfs, zfs and ext4 filesystems on loop devices
in a privileged container, including a MariaDB that is written to across
the snapshot and must recover from the restored copy without losing a
committed row. CI installs zfs and sets E2E_REQUIRE_FILESYSTEMS so a
missing kernel module fails the build instead of silently skipping a
filesystem.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 13:01:43 +02:00

60 lines
2.0 KiB
Python

"""Exercise volume_snapshot against a real filesystem, from inside a container.
Runs where loop devices exist. Prints one PASS/FAIL line per assertion and exits
non-zero on the first failure, so the calling test can surface the reason.
"""
from __future__ import annotations
import subprocess
import sys
from pathlib import Path
sys.path.insert(0, "/src")
from baudolo.backup.snapshot import SnapshotError, volume_snapshot # noqa: E402
KIND = sys.argv[1]
SUBJECT = sys.argv[2]
EXPECT = sys.argv[3]
def shell(command: str) -> list[str]:
proc = subprocess.run(command, shell=True, capture_output=True, text=True)
if proc.returncode != 0:
raise SnapshotError(f"{command} exited {proc.returncode}: {proc.stderr.strip()}")
return proc.stdout.splitlines()
def check(label: str, condition: bool) -> None:
print(f"{'PASS' if condition else 'FAIL'} {label}", flush=True)
if not condition:
sys.exit(1)
volume = Path(SUBJECT) / "volumes" / "demo" / "_data"
volume.mkdir(parents=True, exist_ok=True)
(volume / "state").write_text("before\n")
if EXPECT == "unsupported":
try:
with volume_snapshot(KIND, SUBJECT, "e2e", run=shell):
check("snapshot on an unsupported filesystem must not succeed", False)
except SnapshotError as exc:
check(f"refused loudly: {str(exc)[:60]}", True)
sys.exit(0)
with volume_snapshot(KIND, SUBJECT, "e2e", run=shell) as resolve:
frozen = Path(resolve(str(volume))) / "state"
check("the snapshot exposes the volume", frozen.is_file())
check("the snapshot carries the content", frozen.read_text() == "before\n")
(volume / "state").write_text("after\n")
check("a later write does not reach the snapshot", frozen.read_text() == "before\n")
check("the live tree did change", (volume / "state").read_text() == "after\n")
root = Path(resolve(SUBJECT))
check("the snapshot is removed afterwards", not root.exists() or not (root / "volumes").exists())
print("ALL OK", flush=True)