Make the drive migration safe against lost data and interrupted runs #3

Merged
SkyfaR merged 4 commits from migration-safety into main 2026-09-30 16:55:41 +02:00
Owner

Four defects of the drive migration, each of which could lose data or leave a drive that the tool could no longer finish. Three commits.

List what lies next to a library in the same folder

A top-level folder that contains a Steam library was skipped as a whole, and so was the library's own folder. With a library in Games/SteamLibrary, Games/Heroic next to it was neither offered for backup nor listed as unprotected, and the re-download method formatted it away without a warning. The same held for the userdata of a whole Steam installation and for folders in steamapps that Steam does not restore, such as sourcemods. Such folders are now entered, and what is not part of the library becomes an item of its own.

Delete only the backup, not the folder it is in

ogc migrate plan --backup-dir /mnt/backup stored the folder as given, and delete-backup then removed the whole folder with everything else in it. The command line now gives the backup a folder of its own (opengamecompressor-backup/<drive id>), as the desktop app already did. Deleting removes only what the backup wrote, so plans written by earlier versions are safe too.

Check fstab before the drive is touched, and resume after it was

Two halves of one problem in the privileged helper:

  • /etc/fstab was validated only after the drive had been converted or formatted. An entry by LABEL=, PARTUUID= or a /dev/disk/by-* path made the helper fail at that point, with a btrfs drive and an fstab that still described the old filesystem.
  • Once the helper had changed the filesystem without reporting success, the migration could not go on. After a re-download the data then existed only in the backup folder.

What changes:

  • The fstab rewrite is tried as a dry run while the drive is still mounted. Entries by label, partition UUID or label and links to the device are recognised; entries by UUID or label name the new UUID afterwards.
  • Before it unmounts, the helper writes a record of the drive to /var/lib/opengamecompressor/migrate/, where only root can write, including the UUID the new filesystem gets (chosen beforehand). A later run that finds the drive changed goes on only with such a record, and then acts on the record instead of on its arguments.
  • The drive must hold the old filesystem, nothing (a format was interrupted; it is formatted only if it is known by more than its name), or the new btrfs, which is never converted or formatted again. Everything after the change can be repeated.
  • The client finds the drive in each of these states, takes over a finished helper run without calling the helper again, and refuses to discard a plan once the drive was touched. ogc migrate abort --force is the way out.
  • A drive that is not mounted is refused on a first run. Its mount point used to be whatever the caller named; the drive was mounted there and its root directory given to the named owner.

Behaviour users can notice

  • A first run needs the drive to be mounted.
  • LABEL= entries in fstab become UUID=.
  • Without an fstab entry the drive is mounted with nosuid,nodev.
  • A second fstab entry for the drive, or a symlinked /etc/fstab, stops the migration until that is fixed.
  • One small record file per migrated drive stays in /var/lib/opengamecompressor/migrate/.

Accepted limits are listed in docs/ROADMAP.md: a short window in which fstab still describes the old filesystem, an empty drive without any ID after a reboot, and migrations that version 0.1.0 left interrupted after formatting NTFS or FAT.

Testing

  • cargo fmt --check, clippy with -D warnings (with and without the GUI), cargo test --locked: green. 97 library tests, 4 image tests with the real mkfs, wipefs and btrfs-convert.
  • scripts/vm-e2e.sh: all four scenarios pass, including the two new ones. label has fstab name the drive by LABEL= and checks that a duplicate entry and an unmounted drive are refused before anything is touched. resume interrupts a re-download from FAT32 after the wipe, after the format and after the helper finished, and checks that nothing is formatted twice.
  • Safeguards were removed one at a time to see a test fail: 34 in the library and 2 in the image tests did; one (the mode of the record under another umask) cannot be tested in-process.
  • The record store test passes under umask 000, 002, 022, 027 and 077.
  • Not tested: a migration of a real drive.

The VM test is not part of CI, so CI does not exercise the helper as root.

Four defects of the drive migration, each of which could lose data or leave a drive that the tool could no longer finish. Three commits. ## List what lies next to a library in the same folder A top-level folder that contains a Steam library was skipped as a whole, and so was the library's own folder. With a library in `Games/SteamLibrary`, `Games/Heroic` next to it was neither offered for backup nor listed as unprotected, and the re-download method formatted it away without a warning. The same held for the `userdata` of a whole Steam installation and for folders in `steamapps` that Steam does not restore, such as `sourcemods`. Such folders are now entered, and what is not part of the library becomes an item of its own. ## Delete only the backup, not the folder it is in `ogc migrate plan --backup-dir /mnt/backup` stored the folder as given, and `delete-backup` then removed the whole folder with everything else in it. The command line now gives the backup a folder of its own (`opengamecompressor-backup/<drive id>`), as the desktop app already did. Deleting removes only what the backup wrote, so plans written by earlier versions are safe too. ## Check fstab before the drive is touched, and resume after it was Two halves of one problem in the privileged helper: - `/etc/fstab` was validated only after the drive had been converted or formatted. An entry by `LABEL=`, `PARTUUID=` or a `/dev/disk/by-*` path made the helper fail at that point, with a btrfs drive and an fstab that still described the old filesystem. - Once the helper had changed the filesystem without reporting success, the migration could not go on. After a re-download the data then existed only in the backup folder. What changes: - The fstab rewrite is tried as a dry run while the drive is still mounted. Entries by label, partition UUID or label and links to the device are recognised; entries by UUID or label name the new UUID afterwards. - Before it unmounts, the helper writes a record of the drive to `/var/lib/opengamecompressor/migrate/`, where only root can write, including the UUID the new filesystem gets (chosen beforehand). A later run that finds the drive changed goes on only with such a record, and then acts on the record instead of on its arguments. - The drive must hold the old filesystem, nothing (a format was interrupted; it is formatted only if it is known by more than its name), or the new btrfs, which is never converted or formatted again. Everything after the change can be repeated. - The client finds the drive in each of these states, takes over a finished helper run without calling the helper again, and refuses to discard a plan once the drive was touched. `ogc migrate abort --force` is the way out. - A drive that is not mounted is refused on a first run. Its mount point used to be whatever the caller named; the drive was mounted there and its root directory given to the named owner. ## Behaviour users can notice - A first run needs the drive to be mounted. - `LABEL=` entries in fstab become `UUID=`. - Without an fstab entry the drive is mounted with `nosuid,nodev`. - A second fstab entry for the drive, or a symlinked `/etc/fstab`, stops the migration until that is fixed. - One small record file per migrated drive stays in `/var/lib/opengamecompressor/migrate/`. Accepted limits are listed in `docs/ROADMAP.md`: a short window in which fstab still describes the old filesystem, an empty drive without any ID after a reboot, and migrations that version 0.1.0 left interrupted after formatting NTFS or FAT. ## Testing - `cargo fmt --check`, clippy with `-D warnings` (with and without the GUI), `cargo test --locked`: green. 97 library tests, 4 image tests with the real `mkfs`, `wipefs` and `btrfs-convert`. - `scripts/vm-e2e.sh`: all four scenarios pass, including the two new ones. `label` has fstab name the drive by `LABEL=` and checks that a duplicate entry and an unmounted drive are refused before anything is touched. `resume` interrupts a re-download from FAT32 after the wipe, after the format and after the helper finished, and checks that nothing is formatted twice. - Safeguards were removed one at a time to see a test fail: 34 in the library and 2 in the image tests did; one (the mode of the record under another umask) cannot be tested in-process. - The record store test passes under umask 000, 002, 022, 027 and 077. - Not tested: a migration of a real drive. The VM test is not part of CI, so CI does not exercise the helper as root.
A top-level folder that contains a Steam library was skipped as a
whole, and so was the library's own folder. With a library in
Games/SteamLibrary, Games/Heroic next to it was neither offered for
backup nor listed as unprotected data, and the re-download method
formatted it away without a warning. The same held for the userdata
of a whole Steam installation on the drive, and for folders in
steamapps that Steam does not restore, such as sourcemods.

Such folders are now entered, and what is not part of the library
becomes an item of its own.
`ogc migrate plan --backup-dir /mnt/backup` stored the folder as given,
the backup was written directly into it, and `delete-backup` then
removed the whole folder, including everything else the user kept
there.

The command line now does what the desktop app does and gives the
backup a folder of its own inside the one the user names. Deleting
removes only what the backup wrote (data, steam, COMPLETE) and the
folders it created if they are empty afterwards, so plans written by
earlier versions are safe as well. Read-only folders in the backup no
longer stop it from being deleted.
Check fstab before the drive is touched, and resume after it was
All checks were successful
CI / Format, lint and test (pull_request) Successful in 1m57s
CI / Arch package (pull_request) Successful in 2m9s
da8c124be7
Two defects of the privileged helper, which are two halves of one
problem.

The fstab entry was validated only after the drive had been converted
or formatted. An entry by LABEL=, PARTUUID= or a /dev/disk/by-* path,
or a mount point listed twice, made the helper fail at that point: the
drive was btrfs and unmounted, and fstab still described the old
filesystem, which stops the next boot without nofail.

And once the helper had changed the filesystem without reporting
success, the migration could not go on: the runner and the helper both
required the old filesystem. After a re-download the data then existed
only in the backup folder.

The helper now tries the fstab rewrite as a dry run while the drive is
still mounted, and recognises entries by label, partition UUID or
label and links to the device. Entries by UUID or label name the new
UUID afterwards. Before it unmounts, it writes a record of the drive
where only root can write, including the UUID the new filesystem gets,
which is chosen beforehand. A later run that finds the drive changed
goes on only with such a record and then acts on the record, not on
its arguments. The drive must hold the old filesystem, nothing (a
format was interrupted; it is formatted only if it is known by more
than its name), or the new btrfs, which is never converted or
formatted again. Everything after the change is repeated without harm.

The client finds the drive in each of these states, takes over a
finished helper run without calling the helper again, and refuses to
discard a plan once the drive was touched; `ogc migrate abort --force`
is the way out.

A drive that is not mounted is refused on a first run. Its mount point
was whatever the caller named, and the drive was mounted there and its
root directory given to the named owner.

The helper sets its own umask, takes the lock before anything else,
and verifies a mount by device and mount point, which no longer fails
on x-systemd.automount entries.

The VM test gains two scenarios: `label` (fstab by LABEL=, a duplicate
entry and an unmounted drive refused before anything is touched) and
`resume` (a re-download from FAT32 interrupted after the wipe, after
the format and after the helper finished).
Let the helper's record decide whether a plan may be discarded
All checks were successful
CI / Format, lint and test (pull_request) Successful in 1m56s
CI / Arch package (pull_request) Successful in 2m9s
CI / Format, lint and test (push) Successful in 1m58s
CI / Arch package (push) Successful in 2m10s
df1fa3aaf5
btrfs-convert changes the filesystem signature only at its very end,
so a drive under conversion looks untouched for the whole run. Whether
its plan could be discarded then depended on the file that shows a
helper at work, which everybody can lock, and which the helper does
not wait for.

A drive with a record that is not finished and that is not mounted at
its mount point now counts as touched. The helper writes the record
right before it unmounts the drive and removes it when it has mounted
the unchanged drive again, so this is what root itself has written
down. The file that shows a helper at work only chooses the message.

The store test now also locks that file from a second descriptor to
see that the helper is not kept out, and changes the modes of both
lock files to see them set again.
SkyfaR merged commit df1fa3aaf5 into main 2026-09-30 16:55:41 +02:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
LevelXStudios/OpenGameCompressor!3
No description provided.