Files
bookstore/deploy/backup.sh
T
twooey 1020a496ab Fix seven medium-severity bugs from the full-session code review
- SupplierOffer::insert(): now checks $wpdb->insert()'s return value and
  throws, matching Work/Edition/Isbn (the one Catalog class that hadn't
  been hardened this way). Work::set_product_id() gets the same treatment
  (was flagged low-severity but same fix, bundled here) — a real update
  failure now counts as a per-work sync failure instead of silently
  leaving a stale wc_product_id pointer.

- SupplierOffer: fetched_at was written via current_time('mysql') (site-
  local) while expires_at (SyntheticOfferGenerator) is written in UTC, and
  best_offer_for_isbns() compared against site-local time too — a mismatch
  masked today only because dev's gmt_offset is 0. Switched both writer and
  reader to current_time('mysql', true) (UTC). Verified the read path
  still finds all active offers correctly under a simulated -5 (US
  Eastern) offset, not just at offset 0.

- HardcoverAdapter: rate-limit throttling moved from "once per work" (in
  Commands.php) to "once per actual HTTP request" (inside query() itself).
  find_book()'s ISBN-then-title/author fallback can fire two real requests
  per work — under the old scheme both shared one throttle sleep, roughly
  doubling the real request rate against a beta API. query() also now
  retries network errors/5xx/429 up to 3x (mirroring
  OpenLibraryAdapter::get_with_retry()), while a GraphQL-level `errors` field
  or other 4xx throws immediately (retrying a rejected query can't fix it).
  Commands.php adds a 5-consecutive-failure circuit breaker so a bad token
  or a wrong field in the still-unverified schema can't silently burn
  through the whole catalog with zero progress. Verified all of this
  directly against Hardcover's real API with a deliberately invalid token:
  5 fast (non-retried) 401s, correct abort message, and confirmed the
  failed works were NOT marked synced (so a real token can retry them).

- restore.sh: now drops and recreates the target database before restoring
  the dump, and clears the uploads directory before extracting the
  archive — previously both restored on top of existing state, so a stray
  table or file NOT in the backup would silently survive a restore drill.
  Verified end-to-end against a real local staging stack: planted a stray
  table and a stray upload file after taking a backup, ran restore.sh, and
  confirmed both were gone afterward while the actual backed-up data (20
  works, a known upload file) came back correctly. Also fixed a real
  permission gap hit during that same test: a fresh volume's uploads dir
  is root-owned until something chowns it, which broke the new www-data
  clear step — now clears as root and chowns to www-data afterward, which
  also means restore self-heals the exact root-owned-uploads class of bug
  fixed earlier this session for the cron sidecar.

- poll-deploy.sh: added a non-blocking flock so a deploy that runs longer
  than the cron interval can't have a second poll fire mid-deploy and race
  its git checkout/reset against the same live working tree. Verified: a
  concurrent run correctly skips instantly while the lock is held, and
  proceeds normally once it's released. (Full atomicity of the live PHP
  file swap under real traffic is a bigger architectural question —
  blue-green or symlinked releases — flagged to the user rather than
  attempted here.)

- backup.sh: now also archives .env.<environment> itself (chmod 600) and
  includes it in the off-host rclone sync alongside the DB dump and
  uploads archive. Every API key and both DB passwords previously lived
  only on the host in this one gitignored file — losing the host lost all
  of it even with DB/uploads backups intact.
2026-08-27 15:50:49 -04:00

69 lines
3.3 KiB
Bash
Executable File

#!/usr/bin/env bash
# Off-host database + uploads backup (design doc §01, launch gate: "A full
# backup has been restored successfully into staging" — see restore.sh).
set -euo pipefail
ENVIRONMENT="${1:?Usage: backup.sh <staging|production>}"
export ENVIRONMENT
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
cd "$REPO_ROOT"
ENV_FILE=".env.${ENVIRONMENT}"
# shellcheck source=lib/env.sh
source "$REPO_ROOT/deploy/lib/env.sh"
TIMESTAMP="$(date +%Y%m%d-%H%M%S)"
BACKUP_DIR="$(env_get "$ENV_FILE" BACKUP_DIR)"
BACKUP_REMOTE="$(env_get "$ENV_FILE" BACKUP_REMOTE)"
BACKUP_RETENTION_DAYS="$(env_get "$ENV_FILE" BACKUP_RETENTION_DAYS)"
BACKUP_DIR="${BACKUP_DIR:-./backups}/${ENVIRONMENT}"
mkdir -p "$BACKUP_DIR"
# See deploy.sh for why: Compose warns on any $-shaped value in --env-file
# even when unused, so DB_PASSWORD/DB_ROOT_PASSWORD are filtered out of the
# copy Compose actually sees.
COMPOSE_ENV_FILE="$(mktemp)"
trap 'rm -f "$COMPOSE_ENV_FILE"' EXIT
grep -Ev '^(DB_PASSWORD|DB_ROOT_PASSWORD)=' "$ENV_FILE" > "$COMPOSE_ENV_FILE"
COMPOSE="docker compose -p bookstore-${ENVIRONMENT} -f docker-compose.yml -f docker-compose.${ENVIRONMENT}.yml --env-file ${COMPOSE_ENV_FILE}"
echo "==> dumping database"
# --single-transaction: without it, a dump against a live site either
# table-locks for its duration (blocking writes) or, if MARIADB_USER lacks
# LOCK TABLES privilege, produces a non-atomic dump — rows written after the
# dump starts but before it reaches their table can be captured
# inconsistently with rows it already passed. InnoDB (this project's engine
# throughout) supports a consistent snapshot via a single transaction instead.
$COMPOSE exec -T db sh -c "exec mariadb-dump --single-transaction -u\"\$MARIADB_USER\" -p\"\$(cat /run/secrets/db_password)\" \"\$MARIADB_DATABASE\"" \
| gzip > "$BACKUP_DIR/db-${TIMESTAMP}.sql.gz"
echo "==> archiving uploads"
$COMPOSE run --rm -T -u www-data wordpress tar -czf - -C /var/www/html/wp-content uploads \
> "$BACKUP_DIR/uploads-${TIMESTAMP}.tar.gz"
echo "==> archiving environment config"
# Every API key (Booksrun/Ingram/Helcim/MailerLite/Hardcover), the WP admin
# bootstrap credentials, and both DB passwords live ONLY in this one
# gitignored host file — losing the host without this backed up loses all of
# it, even with the DB dump and uploads intact. Same trust model as the DB
# dump above (also plaintext, also only as protected as $BACKUP_DIR/
# $BACKUP_REMOTE are) — chmod 600 since, unlike the DB/uploads archives, this
# one is directly the credentials themselves, not data that merely contains some.
cp "$ENV_FILE" "$BACKUP_DIR/env-${TIMESTAMP}"
chmod 600 "$BACKUP_DIR/env-${TIMESTAMP}"
if [[ -n "${BACKUP_REMOTE:-}" ]]; then
echo "==> syncing to off-host storage ($BACKUP_REMOTE)"
rclone copy "$BACKUP_DIR/db-${TIMESTAMP}.sql.gz" "$BACKUP_REMOTE/${ENVIRONMENT}/"
rclone copy "$BACKUP_DIR/uploads-${TIMESTAMP}.tar.gz" "$BACKUP_REMOTE/${ENVIRONMENT}/"
rclone copy "$BACKUP_DIR/env-${TIMESTAMP}" "$BACKUP_REMOTE/${ENVIRONMENT}/"
else
echo "==> BACKUP_REMOTE not set — backup stayed local only; configure rclone before launch"
fi
echo "==> pruning local backups older than ${BACKUP_RETENTION_DAYS:-14} days"
find "$BACKUP_DIR" -type f -mtime "+${BACKUP_RETENTION_DAYS:-14}" -delete
echo "==> backup complete: $BACKUP_DIR"