From 0f91a79ba6bf1e3cc6b3d7ca3898967d80e51dbc Mon Sep 17 00:00:00 2001 From: Felitendo Date: Sat, 8 Aug 2026 15:40:07 +0200 Subject: [PATCH] Recover from a pacman lock left behind by a power cut A machine switched off mid-update leaves /var/lib/pacman/db.lck behind. Nothing removed it, so every subsequent run deferred on it - one power cut would have stopped updates permanently and silently, which on an unattended machine is the worst outcome there is. A lock older than the current boot is provably abandoned: no process that could hold it still exists. Those are now removed and the interrupted upgrade is repeated, with pacman reinstalling anything caught half-written. A lock that is merely unheld within the same boot stays untouched and is only reported, since removing it could corrupt a live transaction; the boot-time test is what makes the difference between a proof and a guess. A fuser check is kept alongside it so a backwards clock jump cannot make a live lock look abandoned. Documented what each layer can actually promise: suspend and normal shutdown are blocked by the existing inhibitor, a hard power-off cannot be prevented by anything, and snap-pac's pre/post snapshots remain the backstop. --- Makefile | 2 +- README.md | 30 +++++++++++++++++++++ doc/cachy-auto-update.1.scd | 20 ++++++++++++++ po/cachy-auto-update.pot | 6 +++++ po/de.po | 6 +++++ src/cachy-auto-update-run | 9 +++++++ src/lib/locks.sh | 52 ++++++++++++++++++++++++++++++++++--- 7 files changed, 120 insertions(+), 5 deletions(-) diff --git a/Makefile b/Makefile index 7620aa9..cda2fd4 100644 --- a/Makefile +++ b/Makefile @@ -8,7 +8,7 @@ # Overridable so a packager can pass the version it is actually building # (`make VERSION=$pkgver`). The literal below is the fallback for builds # straight from a checkout, and is what a release tag has to carry. -VERSION ?= 1.0.7 +VERSION ?= 1.0.8 PREFIX ?= /usr DESTDIR ?= diff --git a/README.md b/README.md index 3c56370..7ff33f7 100644 --- a/README.md +++ b/README.md @@ -123,6 +123,36 @@ A leftover `db.lck` from a crashed transaction is never deleted automatically guessing wrong there corrupts a live transaction. After it has been seen unheld on several consecutive runs, you get a notification instead. +## What if the machine is switched off mid-update + +Three layers, in order of how much they can actually promise: + +**Suspend and a normal shutdown are blocked.** The run holds a +`systemd-inhibit --what=sleep:shutdown --mode=block` lock, so closing the lid, +picking "Shut down" from the menu or a short press of the power button will not +interrupt a transaction — the desktop says something is still busy instead. + +**A hard power-off cannot be prevented by anything.** Holding the power button +or pulling the plug cuts power in firmware. What limits the damage is that +pacman's commit phase is short (about a minute even for a 200-package upgrade) +and that most of a run is downloading, where an interruption costs nothing but +a partial file. + +**The next run repairs it.** A `db.lck` left behind is detected and removed — +but only when it is *provably* dead, meaning it is older than the current boot, +so no process that could hold it still exists. The interrupted upgrade is then +simply run again; pacman reinstalls anything that was caught half-written. A +lock that is merely unheld within the same boot is never removed, only +reported, because there the guess could be wrong. + +This last part matters more than it sounds: without it, a single power cut +during an update would leave a lock file that makes every future run defer, +and the machine would stop updating silently and permanently. + +On a Btrfs system with `snapper` and `snap-pac` — the CachyOS default — every +pacman transaction is bracketed by a pre and post snapshot, so a genuinely +broken upgrade can still be rolled back with `snapper rollback`. + ## Configuration `/etc/cachy-auto-update/cachy-auto-update.conf`, one `Key=Value` per line, every diff --git a/doc/cachy-auto-update.1.scd b/doc/cachy-auto-update.1.scd index c39ccf2..1a78092 100644 --- a/doc/cachy-auto-update.1.scd +++ b/doc/cachy-auto-update.1.scd @@ -100,6 +100,26 @@ running are skipped. The machine is never restarted automatically. When a kernel update makes a restart necessary, a notification says so. +# INTERRUPTED UPDATES + +While a transaction is running, *cachy-auto-update* holds a +*systemd-inhibit*(1) lock on _sleep_ and _shutdown_ in blocking mode, so a +suspend, a lid close or a normal shutdown request cannot cut it short. + +A hard power-off - holding the power button, or losing mains power - is not +preventable. On the next run a leftover _/var/lib/pacman/db.lck_ is removed if +it is older than the current boot, since no process able to hold it can still +exist; the upgrade is then repeated and pacman reinstalls whatever was caught +half-written. A lock file that is unheld but was created during the current +boot is reported rather than removed, because there is no way to prove it is +abandoned. + +Without that recovery a single power cut would leave a lock that makes every +subsequent run defer, silently stopping updates for good. + +On Btrfs with *snapper*(8) and *snap-pac*, each pacman transaction is bracketed +by a pre and post snapshot, so a broken upgrade remains rollbackable. + # FILES _/etc/cachy-auto-update/cachy-auto-update.conf_ diff --git a/po/cachy-auto-update.pot b/po/cachy-auto-update.pot index 1c13b92..0c4666d 100644 --- a/po/cachy-auto-update.pot +++ b/po/cachy-auto-update.pot @@ -257,3 +257,9 @@ msgstr "" msgid "The last run was stopped before it finished." msgstr "" + +msgid "Finishing an interrupted update" +msgstr "" + +msgid "The last update was cut short, most likely because the machine was switched off. It is being finished now." +msgstr "" diff --git a/po/de.po b/po/de.po index e6c80d0..f57fcf7 100644 --- a/po/de.po +++ b/po/de.po @@ -258,3 +258,9 @@ msgstr "Zurückgehalten: %s" msgid "The last run was stopped before it finished." msgstr "Der letzte Lauf wurde abgebrochen, bevor er fertig war." + +msgid "Finishing an interrupted update" +msgstr "Abgebrochenes Update wird beendet" + +msgid "The last update was cut short, most likely because the machine was switched off. It is being finished now." +msgstr "Das letzte Update wurde unterbrochen, vermutlich weil der Rechner ausgeschaltet wurde. Es wird jetzt zu Ende geführt." diff --git a/src/cachy-auto-update-run b/src/cachy-auto-update-run index 9c5e8b9..5feb0ff 100644 --- a/src/cachy-auto-update-run +++ b/src/cachy-auto-update-run @@ -117,6 +117,15 @@ if (( ! FORCE )); then fi fi +# A lock left behind by a power cut during a previous update. It is cleared +# before the busy check, because otherwise every future run would defer on it +# forever and the machine would quietly stop updating. +if cau_recover_stale_lock; then + cau_notify normal \ + "Finishing an interrupted update" \ + "The last update was cut short, most likely because the machine was switched off. It is being finished now." +fi + # This one is checked even with --force: proceeding anyway would just hand the # user a lock error instead of doing anything useful. CAU_SKIP_REASON='' diff --git a/src/lib/locks.sh b/src/lib/locks.sh index 35649f4..4203a8f 100644 --- a/src/lib/locks.sh +++ b/src/lib/locks.sh @@ -60,11 +60,55 @@ cau_package_manager_busy() { return 1 } +# cau_pacman_lock_is_stale +# True only when the lock provably cannot belong to anything alive. +# +# The rigorous test is the boot time: no process that existed before the +# current boot can still be running, so a db.lck older than boot is abandoned +# by definition - which is exactly what a power cut during an update leaves +# behind. A lock that is merely unheld *within* this boot is not provable in +# the same way, so it is only reported (see cau_track_stale_lock) and never +# removed; guessing wrong there would corrupt a live transaction. +# +# The fuser check is kept as a second condition purely to survive a backwards +# clock jump making a live lock look pre-boot. +cau_pacman_lock_is_stale() { + local boot lock + + [[ -e $CAU_PACMAN_LOCK ]] || return 1 + + boot="$(awk '/^btime /{print $2}' /proc/stat 2>/dev/null)" + [[ $boot =~ ^[0-9]+$ ]] || return 1 + + lock="$(stat -c %Y "$CAU_PACMAN_LOCK" 2>/dev/null)" || return 1 + [[ $lock =~ ^[0-9]+$ ]] || return 1 + + (( lock < boot )) || return 1 + [[ -z "$(cau_pacman_lock_holder)" ]] +} + +# cau_recover_stale_lock +# Clears a provably abandoned lock so an interrupted update can be finished on +# the next run. Without this, one power cut during an update stops every future +# update permanently and silently - the worst possible outcome for a machine +# nobody is watching. +cau_recover_stale_lock() { + cau_pacman_lock_is_stale || return 1 + + cau_warn "Found a pacman lock older than this boot - an update was cut short" + rm -f "$CAU_PACMAN_LOCK" 2>/dev/null || { + cau_error "Could not remove the stale pacman lock" + return 1 + } + cau_state_clear stale_lock_count + cau_info "Stale lock removed; the interrupted update will be finished now" + return 0 +} + # cau_track_stale_lock -# A db.lck with no process behind it is left over from a crashed transaction. -# Removing it automatically would be reckless - if the guess is wrong it -# corrupts a live transaction - so instead it is counted, and after enough -# consecutive sightings the user is told to clean it up. +# A db.lck with no process behind it but created during this boot: a crashed +# pacman rather than a power cut. Not provable, so it is counted, and after +# enough consecutive sightings the user is told to clean it up. CAU_STALE_LOCK_RUNS=3 cau_track_stale_lock() {