Earns.io
Earns.ioA Company Charges $120 to Transcribe an Hour of Audio. The Machine Does It for $0.22

A Company Charges $120 to Transcribe an Hour of Audio. The Machine Does It for $0.22

What this is for

For anyone who wants to sell transcripts, or has a client waiting for one, without doing the typing. A company charges $120 for the hour and the machine does it for $0.22; you do the ten minutes of editing in between and keep the difference.

What you make

A speaker-labelled transcript with the names spelled right, plus an SRT with timestamps, from a 60-minute recording, in about ten minutes of your own time.

CHECKED 2026-09-1414 SOURCES22 NUMBERSUPDATED 2026-09-14
  • 22 published
  • 0 not published
  • 1 given two ways
What this counts
Typical earnings
Your price; API $0.003 to $0.0065 a minute
Startup cost
$0; APIs bill per minute from file one
Time to first payout
TranscribeMe weekly at $10; Rev within 1 week
Difficulty
2/5

Why this pays

The same hour of audio has four published prices, and they are two orders of magnitude apart.

WhoWhat the hour is worthWhere the number comes from
ElevenLabs Productions, human review of a transcript$120, from $2.00 a minuteElevenLabs' own service page
Rev, what it pays the person typing$24 to $66, at $0.40 to $1.10 a minuteRev's freelancer page
TranscribeMe, what it pays a beginner$15 per audio hourTranscribeMe's pay page
Scribe v2 by API, the machine$0.22ElevenLabs API pricing

Read it from the bottom up. The machine now types the hour for 22 cents. The person who used to type it is paid $15 to $66. The company that sells the finished transcript charges $120. The service that fits in that gap is ten minutes long: take the machine's draft, put the right names on the right lines, export the two files the client actually opens, and quote per audio minute.

That is the whole play. The rest of this page is how to do it so the client never has to rewind.

The play, in three lines

  1. Upload the recording to ElevenLabs Speech to Text in the browser, or send it to OpenAI's transcription endpoint with one command.
  2. Spend ten minutes renaming the speakers and fixing the proper nouns, then export a text file and an SRT.
  3. Quote per audio minute, working up from the machine cost, not down from a plan price.

The path below walks every click with screenshots from a real account. What follows is the reasoning behind each step and the numbers you quote from.

Where the money comes from

Four buyers for the same ten minutes, in the order they pay:

  1. People who own recordings. Podcasters, interviewers, researchers, anyone who records meetings. They pay per audio minute for the document, and the ceiling is the $2.00 a minute a company charges for human review. List it as a Fiverr Gig or send Upwork proposals; the freelance page covers the fees and the first payout.
  2. Video editors who need captions. Same file, exported as SRT. If they also want other languages, the subtitles page adds translation by the token.
  3. The same clients, a second time. Show notes, a summary, a blog post from the transcript: a text model drafts them and the drafting page prices that job.
  4. The platforms, as a floor. TranscribeMe pays $15 per audio hour and Rev $0.40 to $1.10 a minute for typing. If a client offers less than that for the finished document, the client is the wrong buyer.

The pain you are selling the cure for

The client does not want "a transcript". They want a document with the right people's names on the right lines, the product names spelled the way the company spells them, timestamps if there is video to cut, and nothing that makes them stop and rewind. Machine output on its own fails all four. Human typing passes them at $2.00 a minute, and slowly.

So the thing you sell is not transcription. It is the ten minutes that turn machine output into that document, delivered the same day. Price it as that.

The path, step by step

Every figure below is read from the platform profile it belongs to, not written into this page. Follow a name to see every answer we have for that platform.

  1. 1
    Get the recording into a file the tool will take

    OpenAI's transcription endpoint accepts mp3, mp4, mpeg, mpga, m4a, wav and webm, up to 25 MB per file. The ElevenLabs upload window takes audio and video files up to 1000MB; the API behind it goes to 3 GB and 10 hours. Sixty minutes of MP3 at 128 kbps is roughly 58 MB, so for OpenAI either re-encode or split at a pause, and for ElevenLabs upload it as it is.

  2. 2
    Path A, in the browser: open Speech to Text and click Transcribe files

    Speech to Text sits in the left menu of the ElevenLabs dashboard. The page lists your transcriptions, has a Speakers tab for a library of known voices, and a black Transcribe files button top right. The button opens the upload window; nothing is charged until you confirm inside it.

    Screenshot of Speech to Text page in the ElevenLabs dashboard, top right, ElevenLabs, with Transcribe files outlined in red
    Outlined in red: Transcribe files. Speech to Text page in the ElevenLabs dashboard, top right, ElevenLabs, screenshot taken 2026-09-14. Account details are masked.
  3. 3
    Set the language, turn Tag audio events off for a meeting, and add the names as keyterms

    The window has four tabs, Upload, Record, YouTube and URL, and six controls: Primary language (leave Detect unless you know it), Tag audio events (on by default; off for a transcript nobody wants laughter marked in), Include subtitles, No verbatim, Assign speakers from library, and Keyterms. Type the people, company and product names into Keyterms; the docs say keyterms raise the cost by 20%.

    Screenshot of Transcribe files window, with a test file attached and the six controls, ElevenLabs, with Keyterms outlined in red
    Outlined in red: Keyterms. Transcribe files window, with a test file attached and the six controls, ElevenLabs, screenshot taken 2026-09-14. Account details are masked.
  4. 4
    Read the credit estimate next to Upload files, then upload

    Once a file is attached, the window shows what it will cost in credits before you commit. For our 31-second test file it said 439 credits, and the figure did not move when we toggled Tag audio events or Include subtitles. That is about 14 credits a second, which is more than the 330 a minute on the pricing page; the price section below works through both numbers. Click Upload files; a half-minute file was back in under ten seconds.

    Screenshot of Transcribe files window, the credit estimate line beside Upload files, ElevenLabs, with Upload files outlined in red
    Outlined in red: Upload files. Transcribe files window, the credit estimate line beside Upload files, ElevenLabs, screenshot taken 2026-09-14. Account details are masked.
  5. 5
    Path B, one command: send it to OpenAI's transcription endpoint

    Send the file to /v1/audio/transcriptions with model gpt-transcribe, the model OpenAI recommends for recorded speech. Pass a prompt describing the recording, keywords for the terms you expect to hear, and languages for the languages spoken. If you need speaker labels, switch to gpt-4o-transcribe-diarize with response_format diarized_json and chunking_strategy auto for anything over 30 seconds.

  6. 6
    Rename the speakers and fix the names in the Transcript Editor

    Open the transcript from the list. Speakers arrive as Speaker 0 and Speaker 1. Click the label beside any segment and it turns into a text field; type the real name and press Enter, and every segment of that speaker takes the name. Then click each proper noun on your list to hear the audio from that word and correct the spelling. Edits save on their own; the top bar shows Last saved.

    Screenshot of Transcript Editor, a speaker label opened as a text field for renaming, ElevenLabs, with Marta outlined in red
    Outlined in red: Marta. Transcript Editor, a speaker label opened as a text field for renaming, ElevenLabs, screenshot taken 2026-09-14. Account details are masked.
  7. 7
    Generate the subtitle cues before you ask for an SRT

    Export offers SRT and VTT, but the first time you pick one the editor opens Generate subtitles instead of a download: Max characters per line 42, Max lines per cue 2, Max seconds per cue 7, and a choice between cutting cues from the existing speaker segments or to best satisfy those rules. Click Generate. A Subtitles tab appears beside Transcript, with every line's character count against the 42.

    Screenshot of Generate subtitles window that opens the first time you pick SRT or VTT under Export, ElevenLabs, with Generate outlined in red
    Outlined in red: Generate. Generate subtitles window that opens the first time you pick SRT or VTT under Export, ElevenLabs, screenshot taken 2026-09-14. Account details are masked.
  8. 8
    Export Text and SRT, then hand over

    The Export button top right lists Text, PDF, DOCX, JSON, HTML, SRT and VTT. Take Text for the document and SRT for whoever is cutting the video; DOCX if the client edits in Word. They cost nothing extra once the transcript exists. Download the same day, for the reason under the terms below.

    Screenshot of Export menu of the Transcript Editor, seven formats, ElevenLabs, with SRT outlined in red
    Outlined in red: SRT. Export menu of the Transcript Editor, seven formats, ElevenLabs, screenshot taken 2026-09-14. Account details are masked.
  9. 9
    Price the hour before you quote it

    By API an hour is $0.18 on gpt-4o-mini-transcribe, $0.22 on Scribe v2, $0.27 on gpt-transcribe, $0.36 on Whisper and $0.39 on Scribe v2 Realtime. Inside the ElevenLabs app the pricing page's 330 credits a minute makes an hour 19,800 credits; the upload window's own estimate for our file implies about 50,000. Budget on the higher figure until ElevenLabs reconciles the two, and compare both with the $2.00 a minute its Productions team charges for human review.

    What it costs to keep alive
    ElevenLabsStarter $6 per month, Creator $22 per month (first month 50% off at $11) with professional voice cloning and 121k credits, Pro $99 per month for 600k credits, Scale $299 per monthsource · checked 2026-09-11
  10. 10
    Check the plan allows client work

    ElevenLabs' Free plan carries no commercial licence and cannot be used for any commercial purpose. A transcript you are paid for is commercial use. The API is billed in dollars, not plan credits, and every paid plan includes the licence.

    The gate before it pays anything
    ElevenLabsA paid plan: the free plan does not include a commercial license and cannot be used for any commercial purpose, and free-plan output must credit elevenlabs.io in the title; every paid plan includes a commercial license outside Beta Servicessource · checked 2026-09-11
  11. 11
    Download it the same day, and know what stays behind

    OpenAI assigns the Output to you and you keep ownership of the Input, so recording and transcript are both yours in the terms. When the agreement ends OpenAI deletes all Customer Content within thirty days and hands nothing back.

    Can you take it with you
    OpenAI APINothing is handed back on the way out: when the agreement terminates OpenAI deletes all Customer Content from its systems within thirty days.source · checked 2026-09-07
  12. 12
    If you are the person typing instead of the machine

    TranscribeMe pays beginning General Transcribers $15 USD per audio hour, so a 2-minute file pays $0.50. Rev pays $0.40 to $1.10 per audio minute, $24 to $66 an hour. Both pay through PayPal only; TranscribeMe needs a $10 balance before a Thursday payout and Rev pays within a week of approval.

    What you are paid in
    TranscribeMeCash in US dollars, per audio hour rather than per hour worked: beginning General Transcribers are paid $15 USD per audio hour, so a 2-minute file pays $0.50 and a 4-minute file pays $1.00source · checked 2026-09-12
    RevCash per audio/video minute: $0.40 - $1.10 for transcriptionists, $0.54 - $1.10 for captioners, with subtitle translators earning the highest per-minute ratessource · checked 2026-09-12

Check these five before you sign up

  • Check the file first: under 25 MB and one of seven formats for OpenAI, under 3 GB and 10 hours for ElevenLabs. If you must split, split at a pause, not mid-sentence.
  • Write the list of names, places and product terms before you upload. It goes in as keyterms on ElevenLabs or as prompt and keywords on OpenAI, and it is the difference between a transcript you edit for ten minutes and one you edit for an hour.
  • Rename the speakers and click through the numbers and proper nouns with the audio playing. That is the whole quality pass for a meeting transcript.
  • Export text and SRT. The SRT costs nothing and saves the client asking for it next week.
  • Quote from the machine cost, not the plan price: $0.18 to $0.39 an hour by API; in the ElevenLabs app 19,800 credits an hour by the pricing page or about 50,000 by the upload window's own estimate, against $2.00 a minute for ElevenLabs' own human review.
  • Use the ElevenLabs Free plan for your own recordings only; it has no commercial licence. OpenAI's API has no free tier for this, so the question does not arise there.

ELEVENLABS PUBLISHES THIS TWO WAYS

What a minute of Speech to Text costs in ElevenLabs credits

The two figures are two and a half times apart. On the pricing page an hour is 19,800 credits and fits inside a $6 Starter month; at the rate the upload window quoted for our file an hour is about 50,000 credits and does not fit in Starter at all. We uploaded 31 seconds, not an hour, so we have not seen which figure bills at length. Until ElevenLabs reconciles them, budget on the higher one and read the estimate line before every upload.

Both pages were live when we read them. This site has not asked ElevenLabs which figure is current, and does not know.

Before you upload: the file

  • OpenAI's endpoint takes mp3, mp4, mpeg, mpga, m4a, wav and webm, up to 25 MB per file. Sixty minutes of MP3 at 128 kbps is roughly 58 MB, so either re-encode at a lower bitrate or split at a pause, never mid-sentence.
  • ElevenLabs' upload window says audio and video files up to 1000MB; the API behind it takes up to 3 GB and 10 hours. A one-hour MP3 goes up as it is.
  • Write the list of names, places and product terms before you upload. It goes in as keyterms on ElevenLabs, or as prompt and keywords on OpenAI, and it is the difference between ten minutes of editing and an hour.

Path A: ElevenLabs in the browser

Everything here was walked through in a real ElevenLabs account in 2026; the screenshots in the path below are from that session.

  • Speech to Text is in the left menu. Click the black Transcribe files button, top right.
  • The window has four tabs (Upload, Record, YouTube, URL) and six controls: Primary language, Tag audio events, Include subtitles, No verbatim, Assign speakers from library, Keyterms.
  • Tag audio events is on by default and marks laughter and applause; switch it off for a meeting transcript. No verbatim drops filler words; decide that with the client, because it has to be set before the upload.
  • Attach the file and a line appears beside Upload files with the credit cost of this exact file. Read it before you click; it is the only place the app quotes the price in advance.
  • A half-minute file came back in under ten seconds. The docs say a file over 8 minutes is cut into four pieces and transcribed in parallel, so an hour does not take an hour.
  • Open the transcript from the list: every word plays the audio from that point, speakers are separated (up to 32, per the docs), and each segment carries its times.

Path B: one OpenAI call

openai audio:transcriptions create --model gpt-transcribe --file meeting.mp3 --raw-output --transform text
  • gpt-transcribe is the model OpenAI's guide recommends for recorded speech. Pass prompt (a sentence about the recording), keywords (the literal terms you expect to hear) and languages (the codes spoken). Keywords are hints, not required output, so include only terms that are really in the audio.
  • For speaker labels, switch to gpt-4o-transcribe-diarize with the diarized_json response format and, for anything over 30 seconds, chunking_strategy set to auto. Up to four reference clips of 2 to 10 seconds let you name the speakers.
  • For an SRT or VTT, use whisper-1 with timestamp_granularities; that parameter is whisper-1 only, and whisper-1 takes a prompt of at most 224 tokens. Translation into English is a separate endpoint, whisper-1 only.

The ten minutes that make it deliverable

Four edits, always the same four:

  • Rename the speakers. They arrive as Speaker 0 and Speaker 1. Click the label beside any segment, type the real name, press Enter; every segment of that speaker takes the name. The top bar shows Last saved; there is nothing to save by hand.
  • Fix the proper nouns. Click each name on your list to hear the audio from that word. With keyterms done properly there will be few. On the API route, run a second pass through a text model with your spelling list, then check its corrections against the audio.
  • Fix the breaks. Enter splits a segment; merge segments joins two from the same speaker. The Segment properties panel shows start, end and duration, with Delete and Align; run Align after edits so the SRT stays in sync. Run Spell Check is one click in the same panel.
  • Settle verbatim once. Readable summary transcript: No verbatim on. Legal or research transcript: off. Ask the client once and write it down.

Export, then hand over

  • Export lists Text, PDF, DOCX, JSON, HTML, SRT and VTT.
  • SRT and VTT have a step in front: the first time you pick one, Generate subtitles opens with Max characters per line 42, Max lines per cue 2, Max seconds per cue 7. Those are the usual broadcast conventions; change them only if the client's style guide says so. Click Generate, and a Subtitles tab shows every line's character count.
  • Deliver Text and SRT; DOCX if the client edits in Word. On the API route, gpt-transcribe returns the text and whisper-1 returns the SRT.
  • Download everything the same day, for the reason under the terms below.

FIG 1

One hour of audio, priced three ways

Every row is a published per-minute or per-hour rate multiplied to 60 minutes, drawn on one linear scale. The three groups are three different transactions, which is why the numbers are so far apart.

What a client pays for a human-reviewed transcript

ElevenLabs Productions, human review$12012

from $2.00 a minute

What a platform pays the person who types it

Rev, high end$668

$1.10 a minute

Rev, low end$248

$0.40 a minute

TranscribeMe, beginner$155

per audio hour

What a client pays the machine, by API

ElevenLabs Scribe v2 Realtime$0.393

$0.39 per hour

OpenAI Whisper$0.361

$0.006 a minute

OpenAI gpt-transcribe$0.271

$0.0045 a minute

ElevenLabs Scribe v2$0.223

$0.22 per hour

OpenAI gpt-4o-mini-transcribe$0.181

$0.003 a minute

0125

ElevenLabs docs, Transcripts and the Transcript Editor12
Show the numbers
GroupRowValueSource
What a client pays for a human-reviewed transcriptElevenLabs Productions, human review$12012
What a platform pays the person who types itRev, high end$668
What a platform pays the person who types itRev, low end$248
What a platform pays the person who types itTranscribeMe, beginner$155
What a client pays the machine, by APIElevenLabs Scribe v2 Realtime$0.393
What a client pays the machine, by APIOpenAI Whisper$0.361
What a client pays the machine, by APIOpenAI gpt-transcribe$0.271
What a client pays the machine, by APIElevenLabs Scribe v2$0.223
What a client pays the machine, by APIOpenAI gpt-4o-mini-transcribe$0.181

What the hour costs

The chart above puts all nine published rates on one linear scale, so the length of each bar is the price: the company's $120 fills the width, the platforms' $15 to $66 take a fifth to a half of it, and the five machine rates are hairlines. By API, an hour is $0.18 on gpt-4o-mini-transcribe, $0.22 on Scribe v2, $0.27 on gpt-transcribe, $0.36 on Whisper and $0.39 on Scribe v2 Realtime; gpt-4o-transcribe-diarize, the speaker-label model, is $0.006 a minute like Whisper. The AI job cost calculator runs these for any length.

The browser route is priced in credits, and ElevenLabs publishes that price two ways:

  • The pricing page says Speech to Text costs 330 credits a minute, so an hour is 19,800 credits. The Free plan's 10,000 credits a month would cover about 30 minutes; Starter, at $6 for 30,000 credits, about 90 minutes; Creator, $22 for 121,000, about six hours; Pro, $99 for 600,000, about thirty.
  • The upload window quoted 439 credits for our 31-second file, about 14 credits a second. At that rate an hour is about 50,000 credits: more than a Starter month, nearly half of Creator.

We uploaded 31 seconds, not an hour, so we have not seen which figure bills at length. The conflict is logged above; until it is resolved, budget on the higher number, and read the estimate line before every upload. The same hour through the ElevenLabs API is $0.22 in dollars, outside the credit pool. Past a couple of hours a month, the API is the cheaper door into the same model and the app is the editor you use afterwards.

Against all of that stands the human price: ElevenLabs' Productions team will have a native speaker review a transcript from $2.00 per minute of audio, $120 for the hour. That is what a company charges a client for human work on this file, and it is the ceiling you quote under.

If you would rather be paid by a platform

The same ten minutes can be sold to a transcription platform instead of a client, at the platform's rate:

  • TranscribeMe pays beginning General Transcribers $15 USD per audio hour, so a 2-minute file pays $0.50. PayPal only; at least $10 in the balance, withdrawn before 9:00 AM PST on a Thursday, paid once a week.
  • Rev pays $0.40 to $1.10 per audio minute, $24 to $66 an hour, weekly via PayPal within a week of approval, after a waitlist.

Both pay for the typing, not for the client relationship. That is why the client route pays more: the $120 belongs to whoever hands the client the finished document.

Who owns the transcript, and what is deleted

  • OpenAI assigns the Output to you and you keep ownership of the Input, so the recording and the transcript are both yours in the terms.
  • When the agreement ends OpenAI deletes all Customer Content within thirty days and hands nothing back. Download the same day.
  • The ElevenLabs Free plan carries no commercial licence and cannot be used for any commercial purpose. A transcript you are paid for is commercial use; every paid plan includes the licence, and the API is billed in dollars outside the plan.

What to do today

  1. Take one recording you already have. Check the file: under 25 MB for OpenAI, under 1000MB for the ElevenLabs upload window.
  2. Write the name list. Upload with keyterms. Read the credit line before you click.
  3. Rename the speakers, click through the names, export Text and SRT.
  4. Time yourself. That number, and $0.22, are the two you quote from.
Cite this page

Earns.io (2026). A Company Charges $120 to Transcribe an Hour of Audio. The Machine Does It for $0.22. Figures checked 2026-09-14. Retrieved from https://earns.io/en/methods/an-hour-of-audio

https://earns.io/en/methods/an-hour-of-audio

OpenAI API

openai.com

GPT models billed per token.

Official site

This link pays this site nothing today. It goes to openai.com.

  • 22 published
  • 0 not published
  • 1 given two ways

ElevenLabs

elevenlabs.io

Voice and audio generation on a credit subscription; the free plan carries no commercial license, and a shared voice clone can earn Voice Actor Payouts.

Official site

This link pays this site nothing today. It goes to elevenlabs.io.

  • 22 published
  • 0 not published
  • 1 given two ways

TranscribeMe

transcribeme.com

Short-clip transcription crowd paying $15 per audio hour to beginners, weekly via PayPal once $10 has built up.

Official site

This link pays this site nothing today. It goes to transcribeme.com.

  • 22 published
  • 0 not published
  • 1 given two ways

Rev

rev.com

Transcription, captioning and subtitle work paid per audio minute, $0.40 to $1.10 for transcripts, weekly via PayPal; applications currently sit on a waitlist.

Official site

This link pays this site nothing today. It goes to rev.com.

  • 22 published
  • 0 not published
  • 1 given two ways