Add a synchronization point after recreating missing index files to
ensure all uploads are settled and committed to the database before
the backup process continues. This prevents a race condition where
the backup could start while index files are still being uploaded.
In this case, it would be possible for the backup to reference the file that is currently being deleted. This would be very rare in production as the delays have to align and the new data must be changed and processed very quickly to trigger the issue.
In unittests we have observed the issue for a while but only showing up very rarely.
Instead of requesting all scopes up front, each sub-section of the backup/restore will request the permissions that are relevant for the operation.
The effect is that it is not required to grant all permissions to get a backup running. If permissions are failing, that part of the backup will generate warnings for items that are excluded.
When performing queries with large inputs, there is a possibility that we reach the maximum number of possible parameters. If the query exceeds the maximum number of parameters the query fails.
This change uses a temporary table that will be created in the case the inputs exceed the total number available parameters, making sure the calls proceed as before with small inputs, but uses temporary tables on larger inputs.
Some tests were added to ensure the update works as expected, even with large inputs.
This PR fixes a reported issue with strings that are less than 3 characters.
This is unlikely, but happens if the code tries to get the scheme of a path, like `/a` or `/`.
The updated logic returns the scheme if one exists, and it is less than 15 characters (to avoid leaking sensitive information).
This PR adds checks for free temporary space when starting the server and when restoring.
On systems that have limited space in the temp folder a warning will now be shown.
This is intended to capture issues on some Docker systems where /tmp is mounted in memory instead of being disk backed.
It will also detect the issue on other systems that are space constrained.
This PR adds guards to prevent creating filesets with multiple files that have the same path.
While it should technically be impossible to have multiple entries that have the same path, it could happend either due to glitches or because the source data (manual lists, remote sources, etc) returns duplicates.
This PR adds a simple check for each folder to ensure that on a folder-level, duplicate paths cannot be introduced.
There is also a post-backup check to evict any duplicates, and the recreate process will reject duplicate paths.
Finally, the repair command will remove duplicates if they somehow manage to get into the database anyway.
A database-level prevention is not currently feasible as it needs a cross-table check for uniqueness, which requires more work from the database to check for each added file. Since this is expected to be a very rare event, the added processing was not justified.