
For five years I filmed works in progress on my phone: a few seconds of a track playing in Ableton, or a piano sketch, shot for social media. The album holds 893 of those videos, 2021 to 2026, and almost none of them say which project they show. This is a set of scripts that finds out, by laying two clocks that never knew about each other side by side. The phone stamps every video with the moment recording started. Ableton stamps every autosave and every recorded clip with the time, in the filename. A video shot at 08:15:25 and an autosave written at 08:15:47 are the same evening at the same desk.
The album arrived as a single 131 GB zip on a drive with 40 GB free, so
snippets_scan.py never extracts it. It pulls one video at a time into a
temporary folder, reads the capture date from its metadata and the tempo from
its soundtrack, saves one mid-clip frame and a 64 kbps mono copy of the audio,
deletes the video, and moves to the next. Peak extra disk use is the largest
single clip; the kept frames and audio come to 389 MB, and the scan resumes
from its own manifest.
The other side of the join needs no drive to be mounted. The indexes built for
the archive recovery already list every project file and every audio file on
every drive with its name, and Ableton writes a [YYYY-MM-DD HHMMSS] stamp
into the filename of each autosave in a project’s Backup/ folder and each
recording in Samples/Recorded. build_activity_timeline.py extracts those,
adds each project file’s last-saved time, and drops bulk-copy artifacts,
giving a timeline of 166,000 Ableton events across the indexed drives,
including the dead one.
snippets_match.py scores each clip by the nearest event and assigns a tier:
a recording started within ten minutes of the video, an autosave or recording
within thirty, anything within three hours, or nothing. Two further signals
settle ties. The tempo heard in the video is compared with the BPM stored in
each candidate set, which separates two projects open on the same evening but
needs the drive holding the set. And when Ableton’s window is in shot, the
saved frame’s title bar names the set outright; those readings are kept as
ground truth against which the clocks are checked.
The tiers are calibrated rather than assumed. Shifting every clip by weeks, so that no true link can exist, still produces a match within thirty minutes for about fifteen percent of clips, because on most evenings some project was active. Against that chance level, a match within ten minutes is right about four times in five, a match within thirty minutes about three times in four, and the three-hour tier is mostly coincidence and is presented as a lead to check. The offset histogram has a single peak at zero, so the phone and the Mac agree on the time and no clock correction is applied.
The ledger page, build_ledger_page.py, groups videos shot within an hour of
each other into one filming session, 355 sessions for the 893 clips, and inside
each session groups consecutive clips with agreeing tempo into take clusters,
since one evening often covered several songs. Each cluster carries its own
thumbnails, candidate sets, trust badge and verdict, in date order, so that
confirming or correcting a match is done by eye against the frame. Match quality
follows the year: 2023 clips mostly resolve within thirty minutes, while 2021
clips miss more often, because Live keeps only about ten autosaves per set and
early sessions were rotated out by later work on the same project. Adding each
project file’s last-saved time cut the unmatched clips from 310 to 191.