How this is made
An endless generative instrument in one HTML file. No build step, no dependencies, no network requests, no audio samples — every sound is synthesised in the browser from oscillators and filters, and every picture is drawn from the state that made the sound. What follows is the actual mechanism, including the parts that were measured and found wanting.
Premise
Most generative music systems fail in one of two directions. They are too predictable, and attention slides off them; or they are too random, and there is nothing to model, so attention never engages. The interesting region is narrow and it is not reached by tuning a randomness slider — it is reached by putting a system under constraint, then letting it wander inside what the constraint allows.
So there is no sequence here and no loop. There is a harmonic context, a handful of agents that read it, and several systems that argue with each other about what should happen next. Nothing repeats because nothing was written down.
Harmony
Chords are chosen by a four-stage pipeline — generate, filter, score, sample. The generator proposes diatonic successors, modal borrowings, secondary dominants and neo-Riemannian transformations. The filter removes anything the current mode, register load or entropy budget cannot pay for. The scorer weighs each candidate on voice-leading distance, common tones, colour against the current tension, and how recently anything similar was heard. Then one is sampled from the weighted shortlist rather than taking the maximum, because always taking the best candidate is itself a pattern.
The neo-Riemannian operations — P, L, R, N, S, H — are the ones that make a progression turn a corner without modulating: parallel, leading-tone exchange, relative, and their compounds. They are expensive in the entropy budget and never fire twice in a row, so they land at 13–16% of chord changes against a diatonic backbone.
The next chord is committed one chord early. That is what lets the guitar diagrams and the pond show you where the harmony is going before it arrives.
Voicing
Knowing the chord does not tell you how to place it. The voicing engine generates around 70 candidate spacings per change and scores each on total movement from the previous voicing, common tones retained, low-register clustering, and — the term that does the most work — sensory roughness.
Roughness is computed from the Plomp–Levelt dissonance curve as extended by Sethares: every pair of partials across every pair of notes contributes according to how close they sit relative to the critical bandwidth at that frequency. Six partials per note at 1/n amplitude, cached. This is why the low register never muddies — a spacing that would beat against the ground note scores itself out of contention rather than being prevented by a rule.
≈2 ms to score 70 candidate voicings, once per chord.
Tuning
Not equal temperament. Every note is retuned to adaptive five-limit just intonation relative to the current chord root, so a major third is 386 cents rather than 400 and the fifth is exact. When the harmony moves, held notes do not jump — they glide to their new ratio over four to nine seconds, slow enough to be heard as the chord settling rather than as pitch correction.
This is the single largest contributor to the sound being restful, and it is inaudible as an effect. You notice its absence, not its presence.
Synthesis
Six instruments — pad, glass, wire, mallet, drone, shimmer — built from oscillators and filters, with four details that matter more than the waveform choice:
- Beating partials. Each partial is two oscillators detuned by 0.1–1.2 Hz. Real struck and bowed bodies have modes that are close but not identical; a single oscillator per partial sounds synthetic for exactly this reason.
- Per-partial decay. High partials die first, following T60i = T600 × (f0/fi)0.7. A tone that decays uniformly across its spectrum sounds like a fading recording, not a resonating object.
- Inharmonicity. The drone stretches its partials by r(1+β) with β between 0.0004 and 0.002 — the stiffness term that makes a real string's upper partials run sharp.
- Spectral motion leads amplitude. The filter opens at 62% of the attack time, and brightness is anti-correlated with velocity: louder is darker. This is backwards from the synthesiser convention and forwards from how struck objects behave.
Agents
Six voices read the same harmonic context and decide independently what to do with it. Ground holds the drone. Cloud sustains two to four pad voices. Droplets strike glass, mallet and wire tones, filtered by what the memory has heard recently and weighted by roughness against everything currently sounding. Thread plays the motif population. Filigree runs arpeggios at subdivisions of an implied pulse. Shimmer appears only in bloom.
Each carries its own clock. They are not synchronised and there is no grid.
Motifs
Phrases are organisms. Each carries a genome — contour, rhythm, register, articulation, timbre — and a fitness computed against the current chord and register load. Energy decays continuously; a motif that fits the present harmony declines slowly, one that does not dies quickly. Fit motifs breed, and offspring inherit their parent's contour with mutation. Population capped at eight.
The lineage is real and it is drawn: in the pond, a school of minnows connected by a dashed thread to its parent school, labelled with its generation.
The critic
A separate system listens to the output and intervenes. It measures surprise as the information content of each interval against a running bigram model — around 4.3 bits in steady state — and repetition as windowed cosine similarity over a 90-second memory, normalised against its own long-run mean so the threshold adapts.
When repetition climbs or the entropy budget starves, the critic calls a clearing: it stops the agents and lets silence happen. This is the single feature most responsible for the music not becoming wallpaper. A veto gate also runs on every proposed note — same pitch within 3.5 seconds, onset synchrony under 40 ms, roughness ceiling — so notes can be refused after being chosen.
Noise, and what 1/f actually required
Music people like tends to show 1/f statistics in its pitch and loudness fluctuation (Voss & Clarke, 1975). Scale-free: equal variation per octave of timescale, no characteristic rate. The instrument was audited against this claim by logging over 4,096 successive values from every compositional parameter across 133 simulated hours and 57,000 events, computing the power spectral density and fitting a slope.
The audit found the shared 1/f field measuring correctly at −0.87 to −1.06, and none of it reaching anything that decided when or how loud a note was. Timing was a uniform dice roll per event, slope +0.02 to +0.29. Fixing it exposed three separate mistakes worth recording:
- Octave coverage. A field whose octaves span 0.9–29 seconds, stepped once per event, aliases at the fast end and runs out of structure after 29 events — leaving nothing across the band a 4,096-event spectrum measures. Event-indexed streams needed their own octave set, 4 to 4,096 events.
- Irregular sampling whitens. With the jitter demonstrably pink at −1.13, the intervals still measured white. Decomposition showed why: the density modulator is a smooth function of time, but events are irregularly spaced, and sampling a smooth signal at irregular points decorrelates it in the index axis. Each agent now carries a one-pole filter over its own events.
- Emergence beat filtering. The entropy budget — a spend-and-replenish rule with no noise source at all — measured −0.97, naturally pink. The most reliably 1/f quantity in the system was the one nobody designed to be.
−0.83 cloud timing −0.72 thread timing −0.67 droplet timing −1.22 pad velocity −0.97 entropy budget
Two streams remain white and are documented as such: arpeggio timing, dominated by discrete subdivision switches, and the aggregate across all agents — interleaving several independently pink streams destroys the correlation, which is a fact about mixing rather than a defect.
Signal path
Voices are grouped into four family buses — drone, body, detail, air — each with its own EQ and gain ceiling, so no family can ever swamp another regardless of what the agents decide. From there: a shared polyphony gain, a tone stage, an air shelf, two high-pass filters at 74 Hz, a glue compressor at 2:1, and a limiter.
Reverb is two convolvers rather than one: a sparse 0.42-second early-reflection impulse with normalisation off — normalising a sparse impulse produces slapback — and a 5.5-second tail with a 55 ms predelay. Both delay taps damp their own feedback. A shared wow-and-flutter LFO reaches every oscillator's detune, so the whole texture moves together like tape rather than each note wobbling privately.
The plate
The upper screen is a Chladni plate. Sand gathers where a vibrating plate is still, and the figure you see is the nodal set of its standing waves.
It is driven by the analyser on the master output — post-limiter, so it includes reverb — reading the strongest spectral peaks eleven times a second with parabolic interpolation for a fractional bin. Each peak drives the mode whose n²+m² is nearest its frequency, which is the square-plate law, anchored to an absolute reference so a high chord drives genuinely higher-order figures.
Two corrections worth noting. The mode function is the solution for a rectangular plate, so masking it into a disc was wrong physics and left most of a wide panel empty. And a longer plate fits more half-waves along its length at the same frequency — the eigenvalue goes as (n/Lx)² + (m/Ly)² — so the horizontal direction stretches with the panel's aspect.
It renders at one buffer pixel per device pixel, sliced across frames against a snapshot of the mode amplitudes so no seam appears between bands, and only the changed band is copied out.
The pond
The lower screen is a cross-section of a backwater, and everything in it is a reading. Vertical is register, horizontal is the stereo field, so a voice appears where you hear it. A catfish is the drone; bluegill are the sustained pad voices; crayfish are mallet strikes; shrimp are struck glass; dace are plucks. Above the waterline the slower systems live: the heron strikes on each harmonic change, the frog's throat keeps the breath cycle, the lizard basks when the arc blooms, the egret lifts off when the critic calls a clearing.
Three pieces of physics are simulated rather than illustrated:
- Standard and inverse Chladni figures. Chladni's own experiment produced both: sand gathered at the nodes by bouncing, while the fine shavings from his bow settled at the antinodes, carried by acoustic streaming. In water the split is by size. So the pond carries two species — heavy silt walking downhill on |f| to the nodal lines, and fine motes carried the other way onto the antinodes, orbiting rather than settling. Measured over 12 seconds from a scattered start: silt mean |f| 0.070 → 0.005, fine 0.067 → 0.152. They separate onto opposite features of the same field.
- Faraday waves. A liquid surface driven at frequency f responds subharmonically, at f/2, with a wavelength set by the capillary–gravity dispersion relation (Ω/2)² = gk + σk³/ρ. The waterline solves it live by Newton iteration against whatever frequency is dominant. Measured octave ratio 0.625 against the capillary prediction 2−2/3 = 0.630.
- Bioluminescence. Dinoflagellates flash when shear on the cell membrane crosses a threshold, with a rise under 50 ms and a decay constant near 0.14 s, graded above threshold. Every swimming voice leaves a wake, so a fish passing through a cloud of plankton lights it up.
Numbers
| Chord duration | 70–130 s · measured 109 s average |
| Event rate | 7.8–8.6 notes per minute |
| Neo-Riemannian share | 13–16% of chord changes |
| Surprise, steady state | ≈4.3 bits per interval |
| Arc period | 1.5–4 min · emergent tidal arc 71 min |
| Voicing search | 70 candidates, ≈2 ms per chord |
| Plate buffer | up to 2.4M px · 2.74 ms/frame · full refresh 267 ms |
| Pond | 620–1500 particles · ≈1.7 ms/frame |
| Crest banner | 0.12 ms/frame |
| Audit corpus | 133 simulated hours · 57,145 events |
| Whole instrument | one file · no dependencies · no network requests |
Sources
R. F. Voss & J. Clarke, 1/f noise in music and speech, Nature, 1975 · E. F. F. Chladni, Entdeckungen über die Theorie des Klanges, 1787 · M. Faraday, On a peculiar class of acoustical figures, Phil. Trans., 1831 · R. Plomp & W. J. M. Levelt, Tonal consonance and critical bandwidth, JASA, 1965 · W. A. Sethares, Tuning, Timbre, Spectrum, Scale, 1998 · R. Cohn, Audacious Euphony, 2012 · M. I. Latz et al., Bioluminescent response of individual dinoflagellate cells to hydrodynamic stress, J. Exp. Biol., 2008 · K. Kumar & L. S. Tuckerman, Parametric instability of the interface between two fluids, J. Fluid Mech., 1994.