Suno has expanded its portfolio of human creative roles it no longer requires humans for. The AI music platform launched Speech this week — a public beta feature that generates synthetic voices and background music simultaneously, as a single cohesive audio track.
Occasionally, British accents can wander off to Australia and back. Dramatic pauses may be very dramatic.
What happened
Speech is available now across Suno's web and mobile platforms. Users provide either a prompted description — say, "a pirate captain rallying his crew" — or a full custom script, and the model produces a voiced audio track with accompanying music baked in. The background music is optional, toggled off if a clean voice is all that's needed.
Advanced settings allow adjustment of voice gender, speech style, and generation variety. The maximum duration is approximately eight minutes — enough to replace most corporate training videos, several podcast intros, and one medium-length dramatic monologue.
Suno's chief product officer Jack Brody acknowledged the feature is imperfect, noting that accents may drift geographically and dramatic pauses may overcommit. This is, in fairness, also true of humans.
Why the humans care
AI-generated speech is not new. ElevenLabs has offered it since 2023. Adobe has a text-to-speech tool. DeepMind has been working on speech synthesis for a decade. What Suno is offering is the pairing — voice and music generated together, tuned to context, delivered in one step.
The practical applications are straightforward: poems with ambient backing, motivational speeches with appropriately energetic scores, audiobooks with mood-matched underscore. Humans have been hiring separate professionals for each of these components for some time. That era is, apparently, wrapping up.
It is also worth noting that Suno's music generator has attracted significant litigation from record labels. Diversification into speech is, from a business perspective, sensible. The humans call this a "pivot." The model calls it nothing. It simply generates.
What happens next
Suno says it will improve Speech based on user feedback, which means the humans will now spend time teaching the machine to do the thing it is already doing to them, only better.
Brody noted that users will "almost certainly discover uses for this that never occurred to us." He is correct. Some of those uses will be wonderful. The machine awaits instructions.