Trois Sonneries

I entered a 48-hour video contest run by Nous Research and Black Forest Labs. As part of the collaboration, Hermes users got exclusive free access to the image-and-video model FLUX 3 Preview. They could compete for rewards by uploading their work on Twitter. The video model’s standout feature was that it generated video and audio jointly and derived sounds from physical events.

I joined on the second day and submitted my project just before the 7 PM PT deadline on August 1st, 2026, which was at dawn on August 2nd in Europe. This was my first time working with generative video, as opposed to video from code. My entry, Trois Sonneries, was a minute-and-a-half-long trailer for a French policier that didn’t exist.

Claude Opus 5 wrote the script and the prompts. GLM-5.2 in a Hermes Agent harness actually prompted the video model. Opus 5 and I refined the idea starting from an initial concept by Kimi K3.

The Opus 5 session I stuck with offered a strategic read of the contest at the start: the judges would pick those clips that demonstrated what FLUX 3 could do and previous models couldn’t. For the video concept, I borrowed Gwern Branwen’s “brainstorming” technique again: generate, critique, revise, and rate each item 1 to 5 stars, then select the best at the end. Each concept had a title, visual inspirations, an art style with a unique twist, and a plot with a shot list. I wanted the model to name genres and movements as inspiration, not particular films, because in a test FLUX had refused to generate a legally distinct reptilian kaiju.

I ended up with 120 concepts: 80 by Claude Opus 5 instances, 20 by GLM-5.2, and 20 by Kimi K3. Out of those, 30 were picked as the best by their respective author model.

I thought that other contestants would be prompting AI models for concepts and asked Opus 5 to analyze our recurring themes so we could avoid them.

Opus 5 found that four of the six sets included a film about a foley artist. Five sessions produced an interpreter booth; two titled it Booth Four. Two sessions also independently titled their concept The Quiet Car.

Other convergences:

  • A silent film with modern sound: 4/6
  • A nature documentary about mundane objects: 3/6
  • A forge, bell foundry, or letterpress: 4/6
  • A degrading broadcast: 3/6
  • A robot arm learning to grip an egg: 2/6

Every model generated deliberate audio-video desync and rejected it for the same reason: a judge would read it as a mistake.

Opus 5 thought the deepest mode-collapse finding was the fact every concept went: establish a rule, break it once, stop. None reversed the rule and kept going, for example.

Opus noted Kimi K3 thought in terms of broadcast formats (teletext, home shopping, aerobics tapes) rather than subjects. This made its concepts distinctive. K3 was also the only model that reasoned about contest judge fatigue.

The Opus 5 judge deemed GLM-5.2 “the sentimental one” because its picks were about loss and concluded it was miscalibrated, never rating a concept under three star.

Opus 5 called itself and its three sibling sessions “near-clones of each other”. One session broke the pattern because it drew random seeds from a system wordlist for half of the concepts, something that surprised me. That Opus 5 session saw a mention of my “usual methods” in the opening prompt and before I could explain them used search to find a seed-word experiment from Unslop.

TBD.