Spatial Audio Waves
Powered by Fireworks AI

Spatially Coherent Audio
for AI Video.

Upload your video and let our intelligent agents separate stems, analyze visual context, and maintain persistent spatial placements across scene cuts using cross-scene coherence memory.

View Architecture

The Missing Spatial Layer for Generative Video

AI video generators like Veo, Runway, and Kling produce stunning visuals, but leave you with flat, disjointed mono or stereo audio. Standard upmixers randomly pan stems on every camera cut, destroying immersion. Continuum solves this. By tracking stems natively across cuts, we deliver true cinematic spatial audio that anchors the narrative.

Cross-Scene Coherence Memory

The centerpiece of Continuum. We persist a JSON state of every object's spatial placement. When a scene cuts, recurring elements (ambient beds, character dialogue, musical motifs) are checked against this memory and intelligently hold their position in the 3D room, rather than jumping randomly between speakers.

Proven Consistency

In a complex 8-scene cinematic test sequence, 27 / 28 recurring elements perfectly retained placement across cuts, only re-panning when justified by on-screen visual shifts.

Pipeline Trace: "Ambient Hum"

S1
Scene 1 Cut

Vision model detects wide shot.
Placed at: surround-left (-110°)

...

Multiple scene transitions...

S6
Scene 6 Cut Memory Match

Continuum anchors the stem to its original placement.
Persisted at: surround-left (-110°)

The Architecture

An end-to-end multi-agent pipeline built on commodity tools, supercharged by our novel coherence logic.

Scene Segmentation

The pipeline begins by analyzing the structural flow of the video. Using advanced threshold detection algorithms, the engine scans the entire video file to identify hard cuts and transitions, slicing the media into discrete, logical scenes.

Tech Stack:PySceneDetect, OpenCV

Cross-Scene Coherence

The first AI upmixer to track objects across cuts. Our memory state anchors stems structurally over time, yielding a continuous, highly immersive experience.

Intelligent Separation

Automatically isolate dialogue, music, and sound effects from your flat master track using state-of-the-art Demucs architecture.

Context-Aware Placement

Vision-language models watch your video to understand where sounds are originating, placing them logically within the 3D space.

Four Selectable Output Formats

Continuum exports true, standards-compliant spatial audio. No proprietary black boxes.

Stereo Binaural

HRTF Convolution

For headphones. Convolved using spaudiopy for immersive 3D simulation on any standard headset.

5.1 Surround

Channel-Based

Standard cinematic surround bed with discrete front, rear, and LFE channels via FFmpeg.

5.1.4 Immersive

Object-Based ADM

Open standard spatial rendering via EBU EAR to a 10-speaker BS.2051 layout, including height.

7.1.4 Immersive

Object-Based ADM

The ultimate cinematic target. 12-speaker spatial render fully supporting overhead and rear panning.

Note: We exclusively use open, object-based spatial rendering via ADM. We do not require or use proprietary Dolby Atmos encoders.