Recording a two-minute clip took two minutes. Turning it into something worth posting took thirty to forty. That work wasn't interesting — it was just long.
— Andrew Ward, Founder, Scorchsoft
Before you record
A Script Queue and a Teleprompter
A queue, not a notes app
Scripts carry a state — planned, used, archived — and a reading time calculated at your own pace, so a content backlog reads as a schedule rather than a pile. There are tags, favourites and search, and a batch importer built for the way people actually work now: draft twenty scripts in a chat app, copy the format the app gives you, and paste them straight in. Anything with a matching ID is updated rather than duplicated.
Thirteen takes, one sitting
The teleprompter sits just below the selfie lens so your eyeline stays close to the camera, and you set the pace in words per minute — between 90 and 175 — rather than in scroll speed, so changing the text size doesn't change your cadence. Drag the script to pause it. Fluff a line and simply say it again: the tidy-up stage later picks the better take and drops the false start.
Your words stay yours
Scripts and recordings never leave the phone and are never sent to a provider, even though the rest of the pipeline talks to one. Scripts do ride Android's backup so they survive a new handset — your API keys deliberately don't.
The brief
Direction, Not a Timeline
Pick a shape, pick a look, say what you want
There is no timeline in SpeakCut, and that is the point. You choose one of five formats — straight to camera, explainer cards, full-screen infographics, director's choice, or fully AI-edited — pick one of sixteen card looks, choose a shape, and set how aggressively the AI may cut. Then, optionally, you type a line of direction: “Audience is founders — emphasise the main point.”
Facts the director can use, kept apart from instructions
A separate background-facts field takes the names, numbers and spellings the director should know but shouldn't act on — a small design decision that matters a lot in practice, because mixing reference material into an instruction prompt is how AI features start doing things nobody asked for.
Revise in plain words
When the first edit isn't quite right, you say so: “Snappier cuts. The second card is too busy.” A revision agent reads the edit you just watched alongside your note, and re-runs only the stages that note touches. Everything else is reused from cache, so a revision costs a few pence rather than the price of a fresh run — and the app itemises that before you commit to it.
The interesting bit
An AI Director Making the Creative Calls
It reads, and it looks
The director is an agent, not a filter. It reads the full transcript of what you said, and — this is the part that changes the output — it also looks at still frames sampled from your footage, so it can see what you were holding up, pointing at or demonstrating. That is why a card lands on the moment you start enumerating rather than at a fixed interval. The vision half is a toggle: turn it off in advanced options and the director works from the transcript text alone.
It returns a plan you can read
The agent's output isn't a finished video, it's a structured plan — where the punch-ins go, which moments want a card, and what each card should say. The app shows you that plan in plain English (“11 punch-ins and 3 explainer cards”) and stops for confirmation before it spends anything on images. You can accept it, skip the cards entirely, or reframe and re-punch any shot by hand on the filmstrip.
Every AI stage is a switch
Captions, the AI director, explainer cards and moving cards each have their own toggle, their own provider and their own line on the bill. You can run SpeakCut with everything on, or with the AI off entirely and use it as a fast on-device jump-cutter. Building AI as separable stages rather than one opaque button is the design decision we carry into client work most often.
Generative imagery
Explainer Cards, Drawn to Order by an Image Model
Sixteen looks, and the brief is yours to edit
When the director decides a moment needs a picture, an image model draws one. The house looks run from Red flat and Red isometric through whiteboard, notebook, newsprint and bold dark, and each one is a style brief you can open and rewrite. The text that goes to the image model is right there on screen — “Flat 2D vector infographic card designed for video: it must read on a phone screen in three seconds” — and you can duplicate any look and change it.
Style is specified; content comes from what you said
The split is deliberate. Your look describes how a card should be drawn; what to draw is written by the director from your own words. It means the cards stay on-message without you writing an image prompt per clip, and it means a look you like keeps working across every video you make.
Looks are told apart by shape, not just colour
Each look also carries a tile shape — blocks, isometric, sketch, circles, poster, schematic, halftone — so two looks in similar colours are still distinguishable by anyone who can't tell those colours apart. That's an accessibility decision inside a generative feature, which is not a place people usually think to make one.
The render
What Happens Between Stop and Share
Cuts that land on word boundaries
SpeakCut cuts from the transcript rather than the waveform. Because it knows where every word starts and ends, it drops silences and false starts at the gaps between words instead of guessing from audio levels — so sentences survive intact and the cuts don't clip your consonants. The app shows its working: a filmstrip of which frames were kept, and a line like “1:42 raw — 4 jump cuts remove 13s of silence.”
Framing, punch-ins and captions
Face detection runs on the phone and positions the crop as you move. Zoom punch-ins snap to whole words so emphasis lands where you put it, with a strength you can set from subtle up to bold. Captions come in three styles — clean, spoken word, and spoken word pop — burned into the pixels and accompanied by an .srt file for the platforms that want one.
Your own b-roll wins
Multi-clip mode takes several takes plus up to twelve of your own clips and photos. Where you've supplied real footage for a moment, it's used in place of a generated card — the generative step fills the gaps rather than competing with material you already have.
After the render
The Caption, the Shapes and the Share
A post drafted from your own words
The last AI stage writes the caption to go with the video, per platform — a main post, a shorter one for X, a YouTube Shorts description — leading with the strongest line in the take and carrying a character count against each platform's limit. It's a draft, framed as one: “Read it before you post — it goes out under your name.” The app never posts anywhere itself; sharing hands the file to whichever app you pick through Android's share sheet.
A free reshaping tool that pays its own way
Resize & Crop is a separate, entirely on-device tool: nine combinations of shape and resolution — 9:16, 4:5, 1:1 and 16:9 at 480p, 720p or 1080p. It samples twelve frames, finds the band containing every face in shot, and anchors the crop there with a little headroom, telling you plainly what it found — “Framed to keep all 3 faces in shot” — or falling back to a centre crop when there's nobody there. One sample clip went from 184 MB to 21 MB. It costs nothing to run and needs no keys at all.
Architecture
Nine Stages, and Most of Them Never Leave the Phone
On-device by default
Reading the video, finding the cuts, detecting the face, rendering and saving all happen locally. Only the language and image work reaches a provider, and then only as compressed audio, the transcript, small still frames, card briefs and the generated cards. The full video file stays on the device. Scorchsoft runs no server for this: there is no account system, no analytics, and no crash reporter.
It fails soft
No single provider outage kills a run. A stage that fails is skipped and the short is still made — a failed transcription costs you the captions, never the edit. That property is worth more than it sounds: most AI features break badly because one model call is load-bearing for the whole feature.
It survives the real world
Processing runs in a foreground service, so you can background the app or have Android kill it, and the run resumes without billing you twice for work already done. Completed stages are cached, which is also what makes cheap revisions possible.
Economics
Your Keys, Your Bill, Quoted Before the Run
Bring your own key
There is no subscription and no margin on the AI. You connect your own OpenAI and Google keys and those providers bill you at their published rates — transcription, for instance, runs at $0.006 per minute of audio. The model picker shows each model's published rate next to it, so choosing a cheaper director is a visible trade rather than a hidden one. Keys live in the Android Keystore, are never redisplayed, excluded from device backups, and stripped from the diagnostics log before anything is written to it.
Estimated before, gated during, itemised after
Every run is quoted before it starts: a lighter, typical and heavier figure, itemised by stage and by provider, with cached work shown as free reuse. Mid-run, the pipeline stops before the first paid image and asks — “Nothing has been charged yet” — with skipping the cards always an option. A spend guard asks again above a threshold you set, and the cheaper option is the default. Afterwards a ledger shows what each stage actually cost, marked measured or estimated.
Honest about what it can't do
The guards are an alert and a fuse, not a hard cap — the real ceiling is the budget you set in your provider's own dashboard, and the app says so rather than implying protection it can't deliver. Edits that run entirely on the phone are quoted at $0.00, and mean it.
Inside the App
Real screens and a real finished frame from the Android build — the AI showing its working, asking before it spends, and accounting for itself afterwards.






What This Proves We Can Build
SpeakCut is a portfolio piece rather than a product pitch. These are the four things it demonstrates that clients most often need and most often struggle to buy.
Agentic AI that makes decisions, not just text
The director doesn't return prose. It returns a structured, reviewable plan — punch-in positions, card placements, card content — that the rest of the system executes deterministically. That shape, an agent proposing and a pipeline disposing, is how we build AI features that behave predictably enough to ship.
Multimodal pipelines that hold together
Audio, transcript, sampled video frames, generated images and written copy all move through one workflow, across two providers, with cached intermediate state. Most AI proofs-of-concept handle one modality and fall over on the second.
AI features you can afford to run
Per-stage estimates before the run, a consent gate before the expensive step, a spend guard, cache reuse on revisions and a measured ledger afterwards. Unit economics designed in at the start, not discovered in the first month's invoice.
Privacy as architecture, not policy
The heavy, sensitive asset — the video itself — never moves, because face detection and rendering run on-device. There is no server to breach and no account to compromise. When the data boundary is a requirement rather than a preference, this is how it gets met.
What We Built It With
A native Android app running a real media pipeline on the handset, calling out to language and image models only where it has to, with the cost, the failure modes and the privacy boundary all designed rather than assumed.
- Native Android (Kotlin)
- On-device media decode & render
- On-device face detection
- Agentic AI director (plan-then-execute)
- Vision input — sampled frames to a multimodal model
- Generative image cards with editable style briefs
- Word-level transcript alignment
- Burned-in caption renderer + .srt export
- OpenAI — transcription, tidy-up, planning, copy
- Google Gemini — card and motion generation
- Foreground-service job pipeline with resume
- Per-stage cost estimation, gating & caching
- Android Keystore credential storage
- Local-only storage — no backend
Frequently Asked Questions
Is SpeakCut a client project or a Scorchsoft product?
Neither, strictly. It started as an internal tool — we were making short videos about software development and losing half an afternoon per clip to editing. It worked well enough that we polished it and put it out as a public beta. We include it in the portfolio because it shows what we can build, not because we are trying to sell you video software.
Can Scorchsoft build something like this for our business?
Yes — that is the main reason this page exists. The transferable parts are listed above: an agent that returns a reviewable plan rather than raw output, a multimodal pipeline with per-stage costing and caching, on-device machine learning, graceful degradation when a provider fails, and a privacy boundary that keeps the bulk of the data on the user's hardware. If you have a workflow people currently do by hand, tell us about it and we'll quote it.
What exactly does the AI decide, and what stays deterministic?
The AI decides the creative calls: which retakes to keep, where to place punch-ins, which moments deserve an explainer card, what each card should say, and how to word the social post. Everything downstream of that plan — the actual cutting, cropping, face tracking, caption rendering and encoding — is deterministic code running on the phone. Keeping that line clean is why the same plan renders the same video twice.
Which AI models does it use?
OpenAI handles transcription, the retake tidy-up, the director's plan and the written post. Explainer cards come from OpenAI or Google depending on the look you choose, and moving cards use Google. Models are selectable in-app with each one's published rate shown next to it, and the stages are written to be swappable rather than hard-wired to one vendor — which is how we tend to build AI features for clients too.
Does the AI actually see the video, or just read a transcript?
Both, by default. Small still frames are sampled from your footage and sent to the director alongside the transcript, so it can take account of what's actually on screen. That's a toggle in advanced options — turn the vision half off and it plans from the transcript text alone, which is cheaper and sends less.
Where does our video go? Is it private?
The video file itself never leaves the phone. What does go out — to your own provider account, on your own key — is compressed audio, the transcript, small still frames, card briefs, the generated cards, and any background facts you typed. We run no server, so there is nowhere for us to store any of it. Retention is governed by your provider's terms, not ours.
What does a run actually cost?
It depends on the length of the clip, how many cards you generate and which models you pick, so we don't publish a headline figure we can't stand behind. Instead every run is quoted before it starts, itemised by stage; the pipeline pauses for consent before the first paid image; and a ledger shows the measured cost afterwards. Edits that run entirely on the phone cost nothing at all. Note that the spend guard is an alert, not a hard cap — the real ceiling is the budget you set in your provider's own dashboard.
Why Android only? Is an iPhone version coming?
No iOS version is in development. SpeakCut leans on Android's media stack and on running heavy work in a foreground service, and building it properly for one platform was the faster route to something usable. If you need the same idea on iOS as a commissioned build, that is a different conversation and one we're happy to have.
How long did it take to build?
It was built alongside client work rather than as a funded project, which makes a headline figure misleading. The more useful answer for scoping your own build: the deterministic media pipeline took considerably longer than the AI integration did. That is usually the case, and it is usually the opposite of what people expect when they budget an AI feature.
Can we try it?
Yes. SpeakCut is free during early access and available directly from speakcut.com rather than the Play Store. It is an openly-labelled beta: finished enough to make shorts with, not finished enough that nothing will surprise you. The app says as much on its own beta screen, which felt like the honest thing to do.