An audio essay begins with a paradox.
The moment a written text is given a voice, it becomes something else. A sentence that once existed silently between page and reader is now spoken. It acquires a pace, a tone, a breath. Someone decides where to pause, which word to hold for a fraction longer, where a thought ends and where the next one begins.
Sound inevitably interprets.
The question, then, is how much it should.
At Schrifton, our approach to the academic audio essay begins with a principle that sounds almost restrictive:
The production should not compete with the text.
Its task is not to turn an essay into an audio spectacle. It is not to surround every idea with atmosphere, to dramatise every transition, or to tell the listener what a passage is supposed to feel like.
Its task is more difficult.
It has to create the conditions in which a text can still be read — even when someone else is reading it for you.
The Voice Is Not the Author
Choosing a narrator is therefore not simply a matter of finding a pleasant voice.
Nor, for us, is the ideal voice necessarily the most distinctive one.
A voice can be too present.
It can carry so much personality, theatricality or emotional intention that the listener gradually begins to hear the performer instead of the text.
For an academic essay, this distinction matters.
The narrator has to inhabit the language without occupying it.
We listen for precision, certainly, but also restraint. For a voice capable of carrying a complex sentence without simplifying it through performance. For someone who understands that an argument has its own rhythm and that punctuation is not merely grammatical: it is part of thought.
The narrator should be present enough to guide the listener through the text, but transparent enough for the listener to remain with the writing.
In this sense, narration is less an act of performance than an act of attention.
The voice does not stand in front of the essay.
It reads from within it.
Reading With Someone Else's Voice
There is a subtle difference between listening to something and being read to.
We are interested in the second.
When reading a book or an essay silently, the reader constructs an internal voice. Its tempo is private. A sentence may be repeated. A paragraph may slow us down. Sometimes we stop without quite knowing why.
An audio version inevitably takes some of that freedom away. The pace has already been chosen. The voice has already been given a body.
The production cannot undo this.
But it can resist replacing the reader's imagination with its own.
Our aim is therefore to preserve something of the intimacy of silent reading.
The ideal experience is not:
I am listening to a produced audio piece.
It is closer to:
Someone is reading this text for me.
That difference may appear small.
For us, it determines almost every production decision that follows.
The Discipline of Less
Contemporary audio production gives us an enormous vocabulary.
Music, ambience, sound design, field recordings, spatial effects, transitions, textures, archival material — almost any written passage can be given an acoustic environment.
The fact that we can do this does not mean that we should.
An essay already contains an environment.
It exists in its syntax, its imagery, its references, its rhythm and in the spaces the writer leaves between thoughts.
Adding sound can enrich those spaces.
It can also close them.
This is why we approach the soundscape of an academic audio essay with restraint.
Silence is not an absence waiting to be filled.
It is part of the composition.
A clean voice against an almost empty acoustic field gives the listener room to construct images rather than receive them ready-made.
If a text describes a landscape, we do not necessarily need to hear the landscape.
If it enters a moment of tension, we do not automatically need music to announce that tension.
If a paragraph is moving, it does not require sound to prove that it is moving.
Often the text has already done the work.
Our responsibility is to notice when it has.
A Note, Not a Score
Music presents the same problem in a more concentrated form.
A musical score can transform the meaning of spoken language almost instantly. A chord can introduce melancholy where the sentence itself remained ambiguous. A crescendo can manufacture significance. A recurring theme can impose a narrative architecture that did not exist in the writing.
For some forms of audio storytelling, these are powerful tools.
For the kind of essay we are interested in producing, they can easily become too powerful.
So we began thinking about music differently.
Not as accompaniment.
Not as atmosphere.
Not as a score.
Sometimes, only a note is needed.
A single piano key. A single classical-guitar string. A sound lasting perhaps two seconds.
Its purpose is not to tell the listener what to feel.
It simply marks that something has shifted.
A new section.
A different movement of thought.
A small threshold within the text.
Then it disappears.
The distinction is important.
Music is no longer placed underneath the words. It exists beside them.
The text remains acoustically free.
Sound as Punctuation
This led us to think of these brief musical interventions less as music than as a form of punctuation.
A printed essay already possesses a visual architecture.
There are paragraphs, spaces, headings, page turns, perhaps a blank line between two sections. The eye recognises these structures before the mind necessarily names them.
Audio loses much of that visual information.
A listener cannot see that a new section begins three lines below.
Sound can restore some of this architecture.
But it does not need to illustrate it.
A brief note can function almost like a paragraph break.
Silence can become a margin.
A pause can perform the work of white space.
Seen this way, minimal sound design is not an aesthetic preference added to the essay.
It is a way of translating some of the architecture of reading into time.
Resisting the Fear of Silence
Minimal production is not necessarily easier production.
Often the opposite is true.
When there are fewer elements, each decision becomes more exposed.
The exact duration of a pause matters.
The distance between the voice and the microphone matters.
A breath matters.
The entrance of a two-second note matters precisely because there may be nothing else around it.
There is also a practical temptation to keep adding.
An empty passage can feel unfinished in the studio. Silence can appear to be a problem. Another layer of sound seems to offer polish, movement, production value.
But production value is not measured by the number of audible decisions.
Sometimes the most important decision is the one the listener never notices:
not to add something.
Restraint requires trusting the text.
It also requires trusting the listener.
Literature Before Audio
This is perhaps the principle that matters most to us.
An audio essay is an audio work, but it remains an essay.
Its origin is language.
The sentence comes before the microphone.
The argument comes before the soundscape.
The rhythm of the writing comes before the rhythm of the edit.
This does not mean treating audio as a neutral delivery mechanism. Voice, recording, editing, silence and music inevitably shape the experience.
The aim is not neutrality.
It is fidelity — not merely to the words themselves, but to the literary and intellectual space the words create.
We want sound to make that space accessible without occupying it.
To accompany without explaining.
To structure without decorating.
To read without performing the text away from the listener.
Reading in Another Form
At Schrifton, we often speak about allowing an idea to find the form it requires.
The audio essay introduces an interesting variation on that principle.
Here, the work already has a form.
It has been written.
Our task is therefore not to reinvent it, but to ask what happens when reading moves from the eye to the ear.
What must change?
What should remain untouched?
And how little sound is necessary for a written work to inhabit another medium without losing the qualities that made it worth reading in the first place?
There may be no single answer.
But our starting point is simple:
Listen to the text before adding sound to it.
Listen to its sentences.
Its pauses.
Its internal rhythm.
The images it already creates.
Then decide what, if anything, is missing.
Sometimes it may need music.
Sometimes silence.
Sometimes a single note.
And sometimes all it needs is a voice that understands that it is reading in someone else's place.




