
A cataloguing pipeline and an editing engine over my own footage archive: about 3,400 clips, 56 hours, shot between 2015 and 2026. The pipeline reads every clip for affect in Brian Massumi’s sense, the intensity of feeling over time and the direction it is moving, rather than for what is in the frame. The engine then selects, trims and orders clips so that the assembled piece’s measured intensity tracks a target curve. The curve can be composed by hand or taken from a piece of music.
Every source clip is reduced to a silent 360p proxy, which is all a low-resolution
vision pass ever sees; 327 GB of originals becomes about 10 GB of proxies, and the
originals are never written to. The proxies go through Gemini’s batch API in
three modes per clip. affect returns a time series: an intensity from 1 to 10,
a vector naming the direction of the charge (expanding, forward, toward and
upward on the rising side; contracting, receding, away and inward on the falling
side; static and stable in between), and a free-text atmosphere tag. sizzle
reads the footage as an editor on deadline would: overall energy, dominant tone,
what the shot is best for. The third mode is one of philm’s critical-theory
lenses, rotated across the archive. Batch responses are matched to their requests
by metadata key rather than by position, parsed leniently so that a truncated
reply yields its complete items rather than nothing, and kept raw on disk.
Takes longer than five minutes are sliced into chunks of about two and a half minutes and analysed as clips in their own right, 52 takes into 318 chunks. Short inputs keep the model’s running timecode honest, and each chunk gets its own description, so a 74-minute roll becomes thirty searchable entries instead of one. Coverage stands at 3,380 of 3,384 unique clips; the remaining four are unreadable source files. The catalogue imports into philm’s SQLite archive, with full-text search over the descriptions and structured facets for intensity, energy, tone and atmosphere.
The engine, affect_score.py, takes a target as a function from normalised time
to intensity. Four curves are built in: a swell, an accelerating heartbeat, a
flatline broken by one spike, and a double climax. Alternatively a music track is
the score: its smoothed loudness envelope becomes the curve and its beat grid
becomes the set of cut points. Walking the timeline, the engine chooses for each
moment the clip sub-window whose mean intensity is nearest the target, whose
internal slope matches the curve’s local slope, and whose dominant vector agrees
with the direction of that slope, with a soft continuity term on atmosphere and a
diversity penalty that may only reorder candidates whose fit lies within a fixed
tolerance of the best. Hold length shrinks as intensity rises, so the cut rate
tracks the curve as well as the content. A structure-aware variant labels each
span of a track as breakdown, build, drop or outro from the heavily smoothed
envelope and changes the editing grammar per section. Every run writes a cut
sheet naming each shot, its in-point, its hold and its function, and renders a
rough cut over the proxies with the track muxed underneath. On the four built-in
curves the mean distance between target and achieved intensity is 0.05 to 0.14
points.
Pieces so far include The Room Without Me, a 23-shot film cast by hand from catalogue searches against a written beat sheet; The Modern Subject, six concept films sequenced by their own affect signatures; the four curve studies; and music-driven cuts for my own tracks. All exist at proxy resolution with their cut sheets.
One result came out of the music-driven mode. The same track, cut twice with the same seed, catalogue and settings, once from the pre-master and once from the mastered version, shared 16 to 19 percent of its footage. Mastering flattens the loudness envelope, and a flatter envelope changes which footage matches at each beat and how quickly the cuts come. A mastering decision is also an editing decision.