AI Meditation Audio Maker — Script Written And Voiced
Describe the session in a sentence and the script is written for you, shown to you, and then spoken. A body scan, a metta practice, a yoga nidra or a sleep story — without writing a word or recording anything.
It turns a one-line brief into a written guided-meditation script and a spoken recording of it. The distinction worth drawing is against the text-to-speech tools that dominate this search: those are voice engines that assume you already have a finished script, so you write it somewhere else and paste it in. Here the writing and the voicing are the same step, and the script is visible and editable before any audio is made. The other thing this page does that most do not is tell you exactly how the pauses work, because a guided meditation is mostly silence and silence is the one thing a voice model cannot be asked for.
15 free credits on signup — enough for one complete track with cover art. No card required.
Eight Practices And How Each Script Is Built
These are techniques rather than moods, and each one has a different script shape. Naming the practice settles the structure, the pacing and where the pauses belong.
Body scan
SequentialAttention moved part by part in a fixed order. Long, even pauses between each region.
Breath awareness
AnchoredOne object throughout. Very few words, wide gaps, gentle returns when attention wanders.
Loving-kindness
Phrase-basedRepeated well-wishing phrases directed outward in widening circles. Repetition is the technique.
Yoga nidra
Lying downRotation of consciousness through the body with an intention set at the start and returned to.
Progressive relaxation
Tense and releaseMuscle group by muscle group, holding then letting go. Precise cueing and timing.
Visualisation
Imagery-ledA described place built in sensory detail. The most writing-heavy of these practices.
Noting
LabellingNaming what arises in one word and letting it pass. Sparse and mostly silent.
Sleep story
NarrativeA gentle plotless story that dissolves rather than concludes. Longer text, slower delivery.
How to Make a Guided Meditation in 3 Steps
Name the practice, read the script, then place the silences.
Name the technique and the length
Body scan, metta, yoga nidra, noting, sleep story. The technique decides the script structure far more than the mood does.
Read the script before voicing it
It is written first and shown to you. Fix a phrase that will not land before spending a generation on speech.
Generate in segments, add pauses after
Break where the silences belong. There is no break tag, so the pauses go in on your timeline, where you control them exactly.
A Practice Beats A Feeling
Relaxing and calming describe the intended outcome. A named technique, a length and an audience describe the script.
“A relaxing meditation.”
“Write and voice a ten-minute body scan meditation for someone lying down before sleep, moving from the feet upward in a fixed order, short simple sentences with a natural break after each body region, second person, unhurried and undramatic, ending without a call to action.”
“A calming morning session.”
“Write and voice a five-minute breath awareness meditation for the start of a working day, one anchor at the nostrils, very few words with long gaps implied between them, gentle non-judgemental language for returning when the mind wanders, finishing by widening attention back to the room.”
“Something for anxiety.”
“Write and voice an eight-minute loving-kindness practice, four repeated well-wishing phrases directed first to oneself, then to someone close, then to a stranger, keeping the same phrases each round so the repetition carries it, warm and plain language with no imagery.”

Where The Silence Comes From
Almost everything distinctive about guided meditation audio is timing, and timing is the one thing a text-to-speech engine will not give you. There is no SSML here, no break tag, no way to write a fifteen-second pause into a line — punctuation is the only pacing the voice responds to. So the honest method is to generate narration in segments that break where the silences belong and place the gaps yourself on a timeline. That is not a workaround dressed up as a feature; it happens to be how meditation scripts are already written, with an instruction to pause after each paragraph, and doing it in an editor gives you exact control a tag would not. Everything else worth stating up front: 4,000 characters per generation, around 650 words, roughly six minutes of speech before pauses. Nine voices, no cloning, no accent selection. Narration and music are separate generations that you combine yourself.
- The script is written for you, not assumed to exist already
- Editable text before a single credit is spent on speech
- Eight named practices, each with its own script shape
- A real method for pauses instead of a silent limitation
- Honest gaps: no SSML, no cloning, no accents, no auto-mixed music bed
Session Briefs Worth Copying
Each names the practice, the length and who it is for — and says where the script should break.
Body scan for sleep
“Write and voice a ten-minute body scan meditation for someone lying down before sleep, moving from the feet upward in a fixed order, short simple sentences with a natural break after each body region, second person, unhurried and undramatic, ending without a call to action”
Morning breath practice
“Write and voice a five-minute breath awareness meditation for the start of a working day, one anchor at the nostrils, very few words with long gaps implied between them, gentle non-judgemental language for returning when the mind wanders, finishing by widening attention back to the room”
Loving-kindness
“Write and voice an eight-minute loving-kindness practice, four repeated well-wishing phrases directed first to oneself, then to someone close, then to a stranger, keeping the same phrases each round so the repetition carries it, warm and plain language with no imagery”
Yoga nidra
“Write and voice a yoga nidra script for lying down, opening with an intention set once and returned to at the end, then a rotation of consciousness naming body parts in sequence on the right side then the left, flat even delivery with no rising emphasis, paragraph breaks between each stage”
Sleep story
“Write and voice a slow sleep story about walking through an empty coastal town at night, plotless and gently descriptive, sentences getting shorter and quieter toward the end, dissolving rather than concluding, no characters speaking and nothing exciting happening”
Progressive relaxation
“Write and voice a twelve-minute progressive muscle relaxation, working from the hands up through the arms, shoulders, face, torso and legs, cueing a tense held for five seconds then a release for each group, precise and clearly timed language, paragraph break between every muscle group”
Who Makes Guided Audio Here
Teachers & coaches
Session drafts to review and adapt, in your own structure.
Faceless channels
Script and narration without recording anything yourself.
App & course makers
Uncompressed WAV narration, cleared commercially on a paid plan.
Workplace wellbeing
Short practices written for a specific team and moment.
What You Get
The written script
Readable and editable before anything is voiced.
The spoken track
16-bit mono WAV at 24 kHz, in the voice you picked.
Download on paid plans
Full commercial rights, no attribution, no expiry.
Free share link
Stream it and send it before spending a download.
Frequently Asked Questions
Starting with length and pauses, because those two decide whether this fits your session.
How long a meditation can it produce in one go?
The input cap is 4,000 characters, which is roughly 650 words. Guided meditation is delivered slowly — commonly around 100 to 120 words a minute, often less — so that is about six minutes of continuous narration in a single generation. The reason that stretches much further in practice is that a guided session is mostly silence: the standard way scripts are written is to pause after each paragraph, and those pauses are where the session length actually comes from. One 650-word script with proper pauses laid in comfortably carries a ten to twenty minute sitting. For anything longer, write it in parts and generate each part separately.
Can it insert the pauses itself?
No, and this is the most important thing to understand before you plan a session. There is no SSML support and no break tag, so you cannot ask for a fifteen-second silence in the middle of a line — punctuation is the only pacing control the voice responds to. The workflow that actually works is to generate the narration in segments that break where the pauses belong, then lay them out on a timeline with the silences between them. That sounds like extra work but it is the same decision a meditation teacher makes anyway, and putting the pauses in an editor gives you exact control over timing that a tag would not.
Which voice should I use for meditation?
There are nine, and two of them are built for this. Sage is calm and measured and is the one intended for meditation and audiobooks. Shimmer is softer and gentler and suits sleep content and slow narration. Among the male voices, Echo is steady and neutral and Fable is warmer and more expressive, while Onyx is deep and authoritative and generally too declarative for this use. Audition rather than assume — a voice that reads well in a body scan can feel wrong in a loving-kindness script. Every generation costs the same, so trying two is cheap.
Can it use my own voice, or a custom accent?
No to both, and both come up constantly in this category so they deserve a plain answer. There is no voice cloning here, so a meditation in your own voice is not something this can produce — if that is the goal, a cloning tool is the right choice. There is also no accent selection: the nine voices are what exist, all English-native, and there is no setting that makes one British or Australian. The tool used to display a much longer list of accented voices that silently collapsed onto the same nine, and that was removed precisely because it was not honest.
Does it add the background music underneath automatically?
No — each generation returns one thing, so the narration and the ambient bed are made separately and combined by you. In practice that means generating the spoken track here, generating a bed of bowls, drone or rain as its own generation, and laying them together in any editor with the music sitting well under the voice. There is no automatic ducking and no mixing stage. One platform ranking on this search does assemble script, voice and music in a single pass, which is genuinely more convenient, and it is fair to say so rather than pretend the step does not exist.
What format is the spoken audio, and can I download it?
Speech comes out as a 16-bit mono WAV at 24 kHz, which is uncompressed and well suited to sitting under music and being edited — a point in its favour, since the music side of this tool only produces MP3. Downloading requires a paid plan, from $15 a month for 250 credits, which also carries a full commercial licence with no attribution and no expiry. On the free tier you get 15 credits with no card, enough for one complete generation, and you can stream it and share it by public link but not download it.
Can it write the script if I do not have one?
That is the intended way to use it. Describe the session in a sentence — the technique, roughly how long, who it is for, what it should move toward — and the script is written and then spoken in the same step, rather than you writing it elsewhere and pasting it in. You can read and edit the text before anything is voiced. This is the part that differs most from the text-to-speech tools on this search, which are voice engines that assume you arrive with a finished script already written.
Can I sell meditations made with this, or use them in an app?
On a paid plan, yes. The licence is full commercial with no attribution required and no expiry, and there is no third-party rights holder in the recording, so a paid app, a course, a class or a monetised channel is covered. The free tier is personal and non-commercial with attribution. Credits refund automatically if a generation fails. Two things worth saying plainly: this writes and voices scripts, it does not offer medical or therapeutic advice, and you should review anything you publish rather than shipping a generated script unread.
More Voice & Calm Tools
Same studio, same credits — narration or the bed underneath it.
Write It And Voice It.
Name the practice and the length. The script arrives readable, the narration follows, and the silences are yours to place. 15 credits, no card.
