OpenAI API
openai.com
GPT models billed per token.
This link pays this site nothing today. It goes to openai.com.
- 19 published
- 0 not published
- 0 given two ways
For anyone paid to caption video, or with a video that needs captions in three languages. You get three SRT files inside the client's line and timing rules, made for under a cent a minute by machine, and a clear idea of what a captioner is paid for the same job.
Three SRT files for a ten-minute video, the source language plus two translations, every cue inside the client's line and duration rules.
A client has a ten-minute video and wants captions in its own language plus two more. The same minute of video has three prices, and they are two orders of magnitude apart.
| One minute of video | |
|---|---|
| Machine transcription with cues, gpt-4o-mini-transcribe or Whisper | $0.003 to $0.006 |
| Machine translation of the cue file, per language | a fraction of a cent, by the token |
| ElevenLabs Productions, human subtitling | from $2.20 |
| What Rev pays a captioner for the same minute | $0.54 to $1.10 |
Ten minutes in three languages is under a dollar of machine time. What the client is paying for, at anything up to $2.20 a minute, is the fix pass: cues that fit the screen, start when the speaker starts, and spell the names right in every language. That pass is yours, and it takes the same twenty minutes whether the client pays you a captioner's wage or a service's rate.
response_format=srt with whisper-1 on OpenAI.The path above walks every click. What follows is the reasoning behind each step and the numbers you quote from.
The client's video is finished and the platform it goes on wants captions, and the audience in two other countries wants them in their own language. Auto-captions spell the product name three ways and hold a cue for two sentences. The client wants files that pass the platform's line rules on the first upload. That is not a transcription problem; it is a cue problem, and it is why step one is a toggle and not a transcript.
Every figure below is read from the platform profile it belongs to, not written into this page. Follow a name to see every answer we have for that platform.
A subtitle file is not a transcript with timestamps. ElevenLabs' docs list the differences: cues carry only start and end times, no speaker labels, and must respect characters per line, lines on screen and cue duration. So ask the tool for subtitles from the start: the Include subtitles toggle on ElevenLabs, or response_format srt on OpenAI's whisper-1. Asking for a transcript and cutting it up by hand is the slow way.
On the Speech to Text page click Transcribe files, upload the video (audio and video both accepted, up to 3 GB), add the names and terms as keyterms, switch on Include subtitles and upload. Open the result and switch to subtitling mode with the tab at the top of the editor. Every cue is coloured red or green against the formatting rules, and Edit rules, behind the three dots next to Subtitles, sets characters per line, lines on screen and cue length to the client's spec.
Send the audio to /v1/audio/transcriptions with model whisper-1 and response_format srt (or vtt) and the file that comes back is the cue file. The file limit is 25 MB, so extract the audio track from the video first and keep it compressed. The newer gpt-4o-transcribe and gpt-4o-mini-transcribe return json only, so for subtitles whisper-1 is the model. Its prompt takes up to 224 tokens: put the names there.
Play through once in the subtitle editor. Fix names and numbers by typing; press Enter to split a cue that runs long, merge cues to join two short ones, drag the handles or type exact timestamps where a cue starts late. The docs warn that transcript and subtitles are separate objects in the editor: fixing a word in one does not fix it in the other. Export SRT when every cue is green.
Paste the SRT into a text model with the instruction to translate only the text lines, keep every cue number and timecode exactly as given, and stay under the client's characters per line. Ten minutes of speech is roughly 1,500 words, about 2,000 tokens in and 2,000 out per language, so at $0.60 per million output tokens a language costs a fraction of a cent and at $3.75 about a cent. Do not do it on Gemini's Free tier: the pricing page says that content is used to improve Google's products.
Open each translated file next to the video and read for length: the rule on characters per line that the source passed may fail in the target language, and the fix is a shorter phrasing or a split cue. Name the files by language code, deliver all three with the source file, and keep copies; nothing on this page is archived for you.
On the Dubbing page upload the video or paste its URL, choose the languages, adjust Speaker similarity and click generate; the cost appears before you confirm, charged per minute of source for each language. Free-plan dubs are watermarked with no way to remove it. At 3,000 credits a minute without watermark, dubbing ten minutes into two languages is 60,000 credits, which is more than a Starter month and half of a Creator month.
OpenAI assigns the Output to the customer and has not trained on API data since March 1, 2023 unless you opt in. Gemini's Paid tier does not use your content to improve Google's products; the Free tier does. The client's video is Input; read the line for it before you upload.
Rev pays captioners $0.54 to $1.10 per audio or video minute, and says subtitle translators earn its highest per-minute rates without printing them. ElevenLabs' Productions team charges a client from $2.20 a minute for human subtitling. The machine steps above cost about a cent a minute in total.
Weekly via PayPal, within 1 week of approval, after a waitlist that Rev's page says is due to high application volume.
response_format on /v1/audio/transcriptions is one of json, text, srt, verbose_json, vtt or diarized_json. For gpt-4o-transcribe and gpt-4o-mini-transcribe the only supported format is json; gpt-4o-transcribe-diarize returns json, text or diarized_json. The subtitle model is whisper-1: model=whisper-1, response_format=srt, and the response is the cue file.prompt takes up to 224 tokens, which is where the names and acronyms go.00:01:23,400 --> 00:01:26,900 timecodes.FIG 1LOG SCALE
Published per-minute rates. The first three bars are what a client pays an API to transcribe or live-translate a minute. The last two are what Rev says it pays a human captioner per audio minute, the low and high end of its range, not what a client pays. Log scale, because the groups are two orders of magnitude apart.
The chart above puts one minute of video on all three lists. The machine steps: transcription at $0.003 a minute on gpt-4o-mini-transcribe or $0.006 on Whisper, the model that returns SRT, or $0.22 an hour on ElevenLabs' Scribe v2 API; a translation step at a fraction of a cent per language; $0.034 a minute if you want OpenAI's live translation instead. The human service: ElevenLabs' Productions team from $2.20 per minute, with its language teams translating if you choose. The person: Rev pays captioners $0.54 to $1.10 per minute. The AI job cost calculator runs the machine rows for any length.
Rev pays weekly via PayPal for all approved work, within 1 week of approval, and applications sit on a waitlist due to high application volume. The rate is per minute of video, not per minute worked, so the wage is $0.54 to $1.10 divided by the minutes you spend on each minute of video. Nobody on this page publishes that division.
response_format=srt, and time the fix pass on each.Earns.io (2026). A Human Service Charges $2.20 a Minute for Subtitles. The Machine Does the Minute for $0.006, in Three Languages. Figures checked 2026-09-14. Retrieved from https://earns.io/en/methods/subtitles-and-translation-by-api
https://earns.io/en/methods/subtitles-and-translation-by-api
openai.com
GPT models billed per token.
This link pays this site nothing today. It goes to openai.com.
ai.google.dev
Google's model API with a free tier whose prompts are used to improve Google's products; paid usage is per million tokens and stays private.
This link pays this site nothing today. It goes to ai.google.dev.
mistral.ai
European model maker with a per-token API and open-weight models you can self-host under Apache 2.0 for research and individual use.
This link pays this site nothing today. It goes to mistral.ai.
elevenlabs.io
Voice and audio generation on a credit subscription; the free plan carries no commercial license, and a shared voice clone can earn Voice Actor Payouts.
This link pays this site nothing today. It goes to elevenlabs.io.
rev.com
Transcription, captioning and subtitle work paid per audio minute, $0.40 to $1.10 for transcripts, weekly via PayPal; applications currently sit on a waitlist.
This link pays this site nothing today. It goes to rev.com.