Transcription, recognition and file analysis: the reverse direction of working with AI

Transcription, recognition and file analysis: the reverse direction of working with AI

Call recordings, hour-long talks and scanned contracts sit unused, because an hour of recording takes an hour. The reverse direction of working with AI gets discussed far less than generation, though it pays back faster. Below are five uses and the reliability rule.

Call recordings, hour-long talks and scanned contracts sit unused, because an hour of recording takes an hour. The reverse direction of working with AI gets discussed far less than generation, though it pays back faster. Below are five uses and the reliability rule.

---

Everybody has an archive nobody opens

Recordings of client calls. Hour-long talks. Scanned documents. Photographs of contracts taken on a phone. Video from events.

All of it sits there unused, because to make use of it you have to listen to it or read it. And an hour of recording takes an hour, and nobody spends it.

So valuable information exists but is inaccessible: formally stored, effectively lost.

Most conversation about AI is about generation — how to create something new. The reverse direction gets discussed far less, though it pays back faster.

What can be extracted

Four types of task.

Speech to text. Audio or video becomes a transcript, including with subtitles and timestamps.

Text from an image. Recognition: a scan, a photograph of a document, a shot of a whiteboard after a meeting.

Video analysis. Not only the speech but the content: what happens, what is shown.

Audio analysis. The content of a recording with conclusions rather than only the verbatim text.

The price range: $0.15 per generation. Paid from the wallet on use, the cost visible before you press.

An hour of recording gets processed in minutes.

Five uses that pay back immediately

1. Reviewing a client call

The most lucrative of the five.

You recorded the conversation, uploaded it — and got not only the text but an analysis: what the client said about their losses, what they fear, which lever matters most to them, where you blurred the focus.

Why that is valuable. During a conversation you are busy having it. You miss half of what the client says — not through inattention but because you are working out what to answer.

The review finds what you did not notice. In my own practice a review like that found two mistakes I had not seen with five years of experience.

And it is the raw material for a proposal: the pains the client named become the blocks of a commercial proposal.

2. An archive of talks turned into content

An hour-long recording holds fifteen to twenty completed thoughts, each of which works as a separate piece.

Previously nobody got to them: to find them you had to listen. Now the transcript reads in ten minutes and the fragments are visible immediately.

3. Documents that do not exist as text

Scans, photographs, old files. Recognise them and you can work with them: search, compare, go through them point by point.

Particularly useful before a legal review: a contract photographed on a phone becomes text that can be checked against a checklist.

4. Meetings and internal discussions

A recording becomes structured outcomes: what was decided, who is responsible for what, what the deadlines are.

A side effect valued more than the outcomes themselves: the discrepancies. One person thinks one thing was agreed, another something else. The transcript shows what was actually said.

5. Feedback in bulk

Reviews, enquiries, survey responses. Going through the whole body rather than sampling.

A sample gives an opinion, full processing gives statistics. The difference is fundamental: a problem occurring in a third of reviews may not appear at all in ten read.

The rule that makes this reliable

Here is the main advantage of the whole direction.

When all the material has been supplied by you, there is nothing to invent.

That is the safest mode of working with a model there is. The request "here is a recording, go through it against these points" is more reliable than "tell me what happens in conversations like this".

A phrasing worth adding always: "work only from the attached material; if the answer is not in it, say the data is insufficient".

Without that line the model will fill the gap with something plausible — that is how it is built.

What still gets checked: names, job titles, figures, dates, titles. Recognition goes wrong on those more often than on ordinary text.

Where the technology falls short

A poor recording. Echo, background noise, several people talking at once — the transcription quality drops sharply.

Terms and surnames. Specialised vocabulary gets recognised worse. Those get checked separately.

Who is speaking. Speaker separation does not always work precisely, especially with similar voices.

Handwriting. Recognised noticeably worse than print.

The practical conclusion: five minutes preparing a recording — a quiet room, one microphone — saves half an hour correcting the transcript.

How this sells

Processing archives as a service. Experts, companies and podcasts have recordings sitting idle. The argument is direct: "you have an archive with six months of content in it".

They are aware the archive exists and unaware of its value. That is your sale.

What gets delivered: transcripts, structured outcomes, a list of fragments with timestamps, finished material.

Analysing feedback as a service. A body of reviews and enquiries becomes a list of what to fix, with priorities. For a business that is research which usually does not happen at all.

The price range. By my scrape of 848 listings, the median for one-off work is $30–75, a system or pipeline $490–1 460.

A one-off transcription is per-item work and cheap. A configured process — recordings handled regularly, outcomes arriving by themselves — is the second category.

And the cost: $0.15 per processing against a fee in the tens and hundreds of dollars.

What is available on the platform

Gen AI, the "file → text" mode — four models: speech transcription with subtitles, text recognition from an image, video analysis, audio analysis. $0.15 per generation.

AI Workflow — the "document analyst" block: a file becomes text inside the chain and the result feeds the next step. Plus the "knowledge base" block — retrieval across uploaded documents.

The chain gets assembled as: transcription → analysis → structure → finished material. The iterator runs it across a list of recordings in one go.

Scheduled runs — the Full plan. If recordings appear regularly.

The media assistants — nine of them, including the analytical journalist: it works through a transcript and assembles the structure of a piece. The sales category — reviewing client calls. The legal category — going through recognised documents. All from the Basic plan.

Registration is free and opens three days of full Basic access.

Where to start today

One action. Find one recording you meant to listen to and never did.

Run it through transcription and read the text. Ten minutes instead of an hour.

Usually one of two things emerges: either there is nothing valuable there and it can safely be deleted, or there are three specific thoughts you have not been able to get at for six months.

Both answers are useful — and both take ten minutes.

---

What to transcribe and analyse with

Three tools for different tasks.

MyMeet — turns hour-long calls into text. The first step, after which you work with a transcript rather than a recording.

AssemblyAI in the conversation analytics — recognises voice messages inside the messenger and returns the analysis immediately: what is troubling the client, where the salesperson fell short.

Qwen VLM for diagrams — a separate case: it turns an image of a flow chart into structured data, nodes and links. A scan of a diagram becomes a set of objects.

In Gen AI the "file → text" mode holds four models at $0.15 per generation: speech transcription with subtitles, recognition from an image, video and audio analysis.

And what module 5 is really for

The twelve lessons of the "automator" module are almost all about working with files: analysing reports and advertising analytics · analysing website performance · turning books and documents into summaries · rewriting posts · analysing a brand book · transcription · generating a digital avatar.

What that changes. The reverse direction is not one function but a whole class of tasks, and it pays back faster than generation because you already have the material.

Module 5 — the Plus plan.

---

What to read next

[Reviewing a call with AI](/en/blog/call-analysis-with-ai) — the most lucrative of the five uses.

[AI for journalists](/en/blog/ai-for-journalist) — how to build a structure from a transcript.

[Editing expert video](/en/blog/expert-video-editing-pipeline) — what to do with an archive of recordings.

[An X-ray of your calls: auditing a sales department](/en/blog/sales-calls-audit) — how this sells to a company.

[AI video generation in 2026](/en/blog/ai-video-generation-guide) — the full guide to video models.