Introducing Magic SonoX — VeyoLabs' AI Audio Studio
Magic SonoX is VeyoLabs' purpose-built audio generation studio. Five creation modes, up to 120 seconds per generation, multi-speaker reference conditioning, Canvas export, and a Discover publishing layer — all inside the dashboard. Here is every capability, explained.
Introducing Magic SonoX — VeyoLabs' AI Audio Studio
Magic SonoX is VeyoLabs' audio creation engine. It generates original music, voice-overs, sound effects, ambient atmospheres, and scripted conversations from a natural-language prompt — with optional image and audio reference conditioning for precision control over the output's character and style.
It is available now from the Magic SonoX panel inside your VeyoLabs dashboard.
Five Creation Modes
Magic SonoX is not a single-purpose generator. It covers the full audio production spectrum through five discrete modes, each with its own prompt context and output character.
Music
Generate original scored music from a description. Specify instrumentation, tempo, mood, and cinematic context. Examples: "A cinematic orchestral swell with rising brass, war drums, and a full choir — building to a climactic finale", "A Paris romance: jazz piano, brushed drums, accordion, warm acoustic guitar". Magic SonoX produces production-ready tracks suitable for VeyoStudio sequences, TAR music videos, and brand content.
Voice Over
Generate narrated speech from a written script or a description of the narration style. Specify tone, pacing, accent, and context — from documentary narration to product advertising to audiobook delivery. Output is clean, single-speaker narration ready for sync with video.
Sound Effect
Generate discrete sound effects from a description: a car door closing, a forest rainstorm, a sci-fi weapon discharge, crowd ambience at a football stadium. The output is a single focused audio event, not a music track or ambient loop.
Ambient
Generate continuous atmospheric audio for scenes, installations, or background layers: ocean surf, a busy café, deep space hum, a medieval marketplace, late-night urban rain. Ambient outputs are designed to loop or layer under visual content without calling attention to themselves.
Conversation
Generate multi-voice dialogue scenes. Describe the exchange — two characters, an interviewer and guest, a customer service call — and Magic SonoX renders the full conversation as a single audio file. Speaker conditioning (see below) lets you anchor each voice to a reference clip for consistent character.
Duration: 30 Seconds, 60 Seconds, 120 Seconds
Every generation has a target duration: 30 seconds, 60 seconds, or 120 seconds (default: 60 seconds). The target duration guides credit pre-check and billing — you are billed per actual second generated, and the model determines final output length within the target window.
| Duration | Estimated credit cost | Best for |
|---|---|---|
| 30 s | ~5 credits | Short spots, SFX, stingers, single-verse demos |
| 60 s | ~9 credits | Standard scenes, narration segments, music loops |
| 120 s | ~17 credits | Full scenes, feature music, extended ambient layers |
The credit rate is approximately 1 credit per second of generated audio. The balance shown in the panel header reflects your live credit total; a low balance surfaces a top-up prompt automatically.
Reference Conditioning: Image and Audio
Magic SonoX accepts two types of optional reference input that condition the generation beyond the prompt alone.
Image Reference (T2A with Visual Context)
Attach a PNG, JPG, or WebP image alongside your prompt. Magic SonoX uses the visual content — scene, mood, palette, subject — as additional context for the audio generation. A sunset photograph conditions the output differently than a battle painting or a neon-lit cityscape, even with an identical text prompt. Image reference is available when no audio references are attached.
Audio References (TA2A — Up to 3 Clips)
Attach up to three audio reference clips (MP3, WAV, M4A, OGG, or WEBM, each up to 30 seconds) to condition the output on specific vocal or instrumental characteristics. Each reference is tagged SPK1, SPK2, or SPK3 — insert the corresponding tag into your prompt to direct which speaker or instrument the reference applies to.
Example: attach a reference voice clip as SPK1 and write "Generate a voice over narrated by `SPK1` describing the product launch" — the output narration will carry the tonal and character qualities of the reference voice.
Audio references and image references are mutually exclusive in a single generation. Use audio references when speaker or instrument character matters; use image reference when scene and visual mood should drive the audio atmosphere.
Advanced Controls
Expand the Advanced section to access output format and acoustic controls:
| Control | Options | Effect |
|---|---|---|
| Output format | MP3, WAV, OGG | File format of the download and Canvas export |
| Sample rate | 22 kHz, 44.1 kHz | Audio fidelity — 44.1 kHz for professional deliverables |
| Speech Rate | −20 to +20 | Pacing of narration and conversation outputs |
| Pitch Control | −10 to +10 | Pitch shift for voice outputs |
For music and ambient generations, format and sample rate are the primary controls. For voice-over and conversation, all four controls affect the output character.
Language: English and Chinese
Magic SonoX supports English and Chinese (ZH) as generation languages — switchable from the panel header. The language selector affects the output language for voice, narration, and conversation modes. Music, SFX, and ambient modes are language-agnostic.
My Score: Your Audio Library
Every generation you produce is saved to My Score — your persistent audio library, accessible from the right panel of the Magic SonoX interface. My Score stores:
- The generated audio file and its playback URL
- The full prompt, mode, duration, and credit cost
- Any reference inputs used
- Timestamp and generation metadata
From My Score you can play any track, download the file, delete the generation, or Pin to Discover to publish it to the shared Discover feed (see below).
Discover: The Shared Audio Feed
Discover is Magic SonoX's public-facing audio showcase layer. Pinned tracks from creators appear in the Discover feed, presented as cards with a title, gradient colour, and optional cover image.
Pinning a track to Discover makes it visible to other VeyoLabs users as a featured audio reference — a portfolio of what Magic SonoX can produce across different styles, modes, and creative briefs. The Discover feed is editable: you can unpin your own tracks at any time.
Canvas Export: Audio into VeyoStudio
Every generated track includes a Canvas export button. Clicking it sends the audio — with its full metadata — directly to the VeyoStudio canvas as an audio layer.
In VeyoStudio, the Canvas audio asset sits alongside your video sequence and can be synced to specific shots, used as the TAR music video track, or layered as ambient bed under a narration. The export preserves the original format and sample rate; no re-encoding occurs on export.
The Canvas button changes to On Canvas once the export is confirmed, preventing duplicate imports of the same generation.
The Playback Interface
The active player in Magic SonoX supports full production monitoring:
- Play / Pause — standard playback control
- Seek — scrub to any position in the generated audio
- Previous / Next — step through your My Score library in session
- Repeat — loop the current track for review under picture
- Shuffle — randomise playback order from the library
- Favourite — mark tracks for quick retrieval
- Download — export the file in the selected output format
Practical Workflows
Music video soundtrack Open Magic SonoX, select Music, write a brief for the track's emotional arc (verse, build, drop, outro), set 120 seconds, generate. Export to Canvas. In VeyoStudio TAR, sync the track to your Seedance 2.5 shot sequence. Magic SonoX handles the music; TAR handles the cut.
Brand campaign voice-over Select Voice Over, paste your ad script or describe the narration style. Attach a reference clip as SPK1 if you have a preferred voice character. Generate at 60 seconds, 44.1 kHz WAV for broadcast delivery.
Cinematic sound design Use multiple Sound Effect generations — one per scene element — then layer them in VeyoStudio alongside your video sequence. A two-minute scene might require five or six separate Magic SonoX generations: environmental ambience (Ambient mode), specific events (Sound Effect mode), and a score bed (Music mode).
Multi-speaker dialogue Select Conversation, attach two reference clips as SPK1 and SPK2, write a scene description. Magic SonoX renders the full exchange as a single track with both voices conditioned to your references.
Magic SonoX is available now on all VeyoLabs plans. Open it from the Magic SonoX icon in your dashboard sidebar.