What AI can and can't do with sound
AI music and audio · Lesson 1 / 20
What changed
Until recently, background music for a video meant buying a licence, hours of searching stock libraries, or a studio. Now a track of the right mood and length is made in a minute from a written description.
But the gain isn't where people expect. AI doesn't replace a composer and doesn't write a hit. It covers functional sound — the kind needed to make something work, which nobody will listen to on its own.
What comes out well
- Background music for a purpose. Neutral, unobtrusive, the right length and mood.
- Short jingles and stings. Three to five seconds of recognisable sound.
- Sound effects and ambience. Street noise, footsteps, room tone.
- Voicing text. A narrator where there is no narrator.
- Drafts and sketches. Testing an arrangement idea before recording with live musicians.
What comes out badly
- Matching what you had in mind. A melody in your head isn't reproduced by description.
- Repeatability. Getting the same track with a small change is nearly impossible.
- Long forms. Over length the structure falls apart and repetition becomes obvious.
- Lyrics in languages other than English. Pronunciation and stress drift.
- Anything meant as an authorial statement. Functional yes; artistic no.
The practical conclusion
Ask yourself: will anyone listen to this sound separately from the thing it's attached to? If not, generation fits and saves a great deal. If yes, you need a person.
Cheat sheet
- The gain is in functional sound.
- Good: background, jingles, effects, narration, sketches.
- Bad: exact melodies, repeatability, long forms.
- The test: is it listened to separately?
List 6 real sound tasks from your practice or plans. Apply the "will anyone listen to this separately" check to each and sort them into: generation fits, needs a person, not sure. For two tasks in the first category, describe what you currently do instead and what it costs in time or money.