Cloud storage (Google Drive, OneDrive, Dropbox)
Paid add-on. Attachments that are explicitly allowed are copied into a folder of your Google Drive, OneDrive or Dropbox, after the submission.
This document describes what the module guarantees and what it refuses. The user guide gives the short version, under “Copying attachments to a storage space”.
Copying a file to a third party is not sending a value
The whole module rests on that difference.
A text field goes to a CRM because an administrator connected that CRM. An identity document goes to a storage space because somebody ticked that particular field, by name.
Hence a whitelist that is empty by default, and empty means “nothing”. That is the opposite of the Zapier trigger, where an empty list lets everything out. A module that copied everything by default would, the day a form gains a file field, deposit documents nobody decided to send.
The path comes from the database, and is therefore untrusted
A file field’s value is a JSON document stored in the database, containing a
path. A form import, a partial restore, an injection: a single forged path would
be enough for this module to read wp-config.php and drop it on somebody’s
Drive.
Every path is therefore checked against the plugin’s upload folder before any
read, by the check that already guards the purge — it resolves the real path,
symbolic links and .. included. Two different checks for the same question
would have ended up diverging, and it is the one we forgot that would be reading
files.
Three providers, and what really distinguishes them
| Google Drive | OneDrive | Dropbox | |
|---|---|---|---|
| Size in one request | 5 MB | 4 MB | 150 MB |
| Creates folders | no | yes | yes |
| Returns a link | yes | yes | no |
Size. Aligning all three on the lowest would have forbidden at Dropbox files it accepts without difficulty; taking the highest would have produced incomprehensible refusals at Microsoft. Each provider announces its bound, and the screen shows it. A file above it is refused without being attempted: sending it to be refused would cost the server’s time and bandwidth for a certain failure.
Folders. Google designates a folder by id and creates nothing: the module looks up and then creates each level. The other two create the tree from the path.
The link. Google and Microsoft return an address that opens the file in their interface, for whoever already has access. Dropbox returns none without creating a share — an address that opens the file to whoever knows it. The plan rules that out, and for good reason: a supporting document uploaded by a visitor would become a public document the moment an address leaked. The screen therefore shows the path.
No share is ever created, at any of the three.
The scopes are minimal, and it shows
Google: drive.file. It gives access only to the files this application
created — the module can neither read, modify nor delete anything else in the
client’s Drive, including if it were hijacked. Plain drive would have given
complete access to the space, in order to drop files into it.
A consequence worth knowing: the root folder must be created by the module, and not picked from the existing folders, which are invisible to it.
Microsoft: Files.ReadWrite, on the OneDrive of the person authorising — and
not Files.ReadWrite.All, which covers every file they have access to, including
the organisation’s shared spaces. offline_access is not a convenience: without
it, no refresh token is issued.
The folder is built, and no value decides its structure
A template — {form}/{year}/{entry} — files things away as they come. A thousand
submissions in one folder give a thousand files that no interface can browse.
Every value is sanitised before entering the template. Without that, a form
title containing One/Two/Three would create three folder levels: the structure
would be decided by the title, and not by the template that exists for it. And a
title containing .. would climb out of the root folder — at Dropbox and
OneDrive, where the path is authoritative, “out” can lead anywhere in the client’s
space.
Depth is bounded to six levels: a thirty-level template would have Google walk thirty folders, two requests each.
The file’s name comes from the browser of the person who uploaded it. It loses its path, its control characters and its excess length — keeping its extension, because a file deposited without one does not open on a double-click.
One file, one copy
UNIQUE (submission, field): an item goes out once. A second copy would cost the
client storage and leave two files with nobody knowing which is authoritative.
The field is part of the key because a submission legitimately carries several items — an identity card and a proof of address — copied separately: one can fail and the other not.
The row claim decides who copies when two tasks cross, which happens as soon as an upload takes longer than the retry interval — and an upload of several megabytes easily does.
“Skipped” and “failed” do not say the same thing
Skipped: nothing was attempted and nothing will be — file absent, size above the provider’s bound, path refused, field not ticked. There is a setting to correct, or nothing to do.
Failed: the provider refused or did not answer. Definitive refusals — 403
for a withdrawn right or an exhausted quota, 404 for a folder that has gone —
are not replayed: every retry consumes the client’s API quota to obtain the same
refusal. 429 and 5xx are, at 1, 5 then 30 minutes. A 401 earns an immediate
renewal, then a second upload.
Remote deletion never commands the local purge
It is optional and off by default: erasing a submission here does not have to erase the document over there, which the client may have filed, annotated or added to a case.
When it is asked for, the remote ids are read before the erasure — afterwards there is nothing left to target — and the deletion is then scheduled. Its failure leaves one file too many at the client’s; holding it back would have left the data here, which is worse.
It is idempotent: a file already absent is a success, not an error. Without that, a second purge would fail forever on what the first had properly erased.
Copying by hand can create a second file
The submission’s detail carries a button per item. All three providers rename rather than overwrite, and that is deliberate: an upload that overwrote would replace a document somebody may have annotated. The screen says so rather than letting you find out.
The checks are not bypassed: the row goes back through the path verification, the size and the whitelist.
What the log does not contain
Neither the local path nor the contents: the second is the file itself, and recording it would make the log a second copy that would go off into backups and exports.
Schema
slf_cloud_copies: one row per submission and per file field, with the provider,
the remote id, the path uploaded to, the link where applicable, the state and the
reason for the last problem.
The connections live in an option, one entry per provider, tokens encrypted with a key of their own.
What V1 does not do
No public folder, no automatic sharing, no two-way synchronisation, no file transformation, and no resumable session upload — that is what bounds the size, and it is the first job to take on if larger files have to get through.
