Documentation

vcut proposes cuts as data and renders only what a human approved. These pages cover the commands, the transcript requirement, and the edit decision list you are approving.

Getting started

vcut turns a raw recording into a clean master in four steps, and stops at every point where a human should look.

npm install -g @crafter/vcut

Or run it without installing:

npx @crafter/vcut recording.mp4

Requirements

vcut shells out to ffmpeg and ffprobe. Both must be on your PATH.

brew install ffmpeg     # macOS
vcut init                # installs everything a first run needs, ffmpeg included on a fresh machine

vcut init reports anything it could not install and exits non-zero, so it works as a precondition check in a script. vcut doctor reruns the same check afterwards, any time something looks wrong.

The four steps

# 1. Find what is worth cutting. Writes candidates, decides nothing.
vcut detect recording.mp4 --preset clean > detect.json

# 2. Draft an edit decision list. Every segment is proposed.
vcut edl build --detect detect.json --output master.mp4 --campaign my-video

# 3. Iterate on the audio, which is where the decisions are.
vcut render --edl edl.json --audio-only --output cut.wav

# 4. Preview it, watch it, and only then render a master.
vcut render --edl edl.json --mode preview

The EDL from step 2 is not optional plumbing. It is the artifact you read and disagree with before anything gets rendered.

Step 3 exists because a round of edits asks audio questions, and rendering the picture to answer them costs far more wall clock than the round needs. Audio-only is not instant either: it costs roughly 1 second per 14 seconds of audio the cut keeps. See Iterate on audio in the cutting loop for why this is where most of a round’s work actually happens.

Editing across several calls

Silence removal alone is rarely the whole edit — see The cutting loop for why it takes several rounds on anything longer than a quick take. vcut open starts the session a real edit runs inside instead of re-detecting and re-passing paths every round:

vcut open recording.mp4 --preset clean --lang es --transcript words.srt   # once
vcut cut recording.mp4 --refs b042..b044 --kind repetition --reason "..." # per finding
vcut commit recording.mp4 --output master.mp4 --campaign my-video         # build + render, one call

open caches the same detect pass step 1 above runs, and turns its silences into stable block refs (b001, b002, …) a later cut points at by name instead of a hand-typed millisecond pair. The four-step quickstart above is the escape hatch: correct for a one-off cut with no second round, or a script with no long-lived working directory. open is where an edit that is going to take more than one round starts instead — see The stateless pipeline is an escape hatch, not an alternative for why the session is the default and not the quickstart above.

Reading the summary

In a terminal, vcut detect prints a summary rather than the raw candidate list:

recording.mp4  6m 22s
  detected dead air       ###.................  16.5%  (1m 03s)
  net after margins       ##..................  10.3%  (~39s once 100ms is kept on each side)
  silences                119 spans, 1m 03s
  longest silence         1s at 6m 20s
  filler words            not scanned; a word list cannot tell filler from ordinary use. Run vcut semantic.
  review candidates       1 (never cut automatically)
                          clipping: peak level -0.24 dB exceeds -1 dBFS

Two numbers appear because they answer different questions. Detected dead air is how much silence exists. Net after margins is how much actually gets removed once the padding around each cut is given back, and that second number is what vcut edl build should land near. A large gap between them means the margin is eating the cuts.

Choosing a preset

The preset sets the loudness floor below which audio counts as silence.

PresetThresholdUse
noisy (default)-20 dBEvents, rooms with ambient noise
clean-30 dBStudio, talking head
podcast-35 dBDeliberate pauses you want to keep

A preset that is too aggressive cuts into breath and delivery. One that is too conservative leaves the pauses in. Start with the one that matches the room, then tune --min-silence if the result is close but not right.

What comes next

Nothing here approves anything. Segments are written as proposed and the EDL as draft, and vcut render --mode master refuses to run until a human changes that. See Commands for the full surface and The EDL for what you are approving.

Commands

vcut <input>                       Shorthand for: vcut detect <input>
vcut detect <input> [flags]        Find silences and review candidates
vcut suspects --detect <path>      Where to look first, ranked, without reading the file
vcut edl build [flags]             Turn a detect report into a draft EDL
vcut semantic export|check|review|merge  Hand the transcript to a model, take proposals back
vcut render --edl <path> [flags]   Render an EDL to video
vcut locate --edl <path> [flags]   Translate between master time and source time
vcut audit --edl <path> --render <path>  Check a render against the EDL it came from
vcut joins --edl <path> --render <path>  Verify every semantic join in one call
vcut say <media> [flags]           Read back what is spoken at a position
vcut silences <media> [flags]      Speech/silence blocks over a range, at a chosen resolution
vcut converge <media> [flags]      Find where a repeated phrase stops coming back
vcut nonspeech <render> [--verify] Find audible sound that is not language
vcut open <media> [flags]          Open or resume a session, map its blocks with stable refs
vcut peek <media> (--ref|--at)     The four views of a position, aligned, disagreement named
vcut cut <media> --refs|--span|--start-ms/--end-ms  Propose a semantic cut against a session
vcut commit <media> [flags]        Build + render a session's proposals into a draft EDL
vcut rounds <media> [--diff N M]   A session's committed rounds, diffed between two
vcut session list|gc [flags]       See what a session store holds, and clear it explicitly
vcut schema [name]                 Print the JSON contract for a command
vcut skills list|get|path [name]   Read the bundled agent manual, or one section of it
vcut doctor                        Check external dependencies
vcut init [--no-skills]            Install everything a first run needs
vcut setup classifier              Fetch the optional non-speech classifier
vcut version                       Print the version

Global flags: --json forces machine output, --human forces the summary, --fields <a.b,c,d> projects JSON output to those dot paths (comma separated, implies --json), --jq <expr> filters or reshapes JSON output with a jq-subset expression (implies --json, mutually exclusive with --fields), --help works on any command.

Every JSON output carries vcutVersion, the version of the binary that produced it, so an agent working from a cached manual can tell the tool changed underneath it. Selected outputs (suspects, detect, edl build, semantic review, nonspeech, render --audio-only) also carry next, a short list of {question, verb} naming what to run next — a hint, not an instruction.

--fields removalPercent,semanticCuts.removedText reads back only those paths, keyed by the path string itself; an array field projects across every element, so semanticCuts.removedText over several cuts returns the array of that field. A path that does not exist becomes a fieldErrors entry naming it rather than failing the call, and vcutVersion always rides along regardless of selection. Exists because a real 11.7-minute run made roughly 40 python3 -c calls extracting 2-3 fields each from full payloads it had already read, with jq on PATH the whole time — a native flag is part of the output contract, discoverable in --help and schema, rather than a second syntax recomposed per call.

--jq <expr> does the structural work --fields cannot. --fields projects to dot paths; it has no way to filter, select, or sort. --jq adds field access, array iteration, select() with comparisons, and sort_by(), so a payload can be filtered or its arrays reordered without leaving the call. It exists for the same reason --fields does: a caller should never reach for python3 -c to filter or merge a JSON payload vcut already produced.

vcut cut recording.mp4 --list --jq '.proposals[] | select(.kind == "repetition")'
vcut rounds recording.mp4 --diff --jq '.semanticCuts | sort_by(.startMs)'

This is a documented subset of jq, not jq itself. No variables, no user functions, no reduce/foreach, no string interpolation, no regex, no object or array construction ({...} or [...]). An expression outside the subset throws naming exactly what was not understood, rather than silently returning something close. The full subset listing ships in vcut --help.

vcut detect

Runs the deterministic pass. Never edits media, never writes an EDL.

vcut detect recording.mp4 --preset clean --lang es --transcript words.srt
FlagDefaultWhat it does
--input <path>—Source recording. Also accepted positionally.
--preset <name>noisynoisy (-20 dB), clean (-30 dB), podcast (-35 dB)
--min-silence <sec>0.3Shortest silence worth cutting
--margin <sec>0.10Padding kept on each side of speech
--lang <code>esFree-form language tag, passed through to the semantic export
--audio <path>—Separate audio recording; silence is measured on this
--transcript <path>—Word-level SRT, used to keep cuts off word edges
--skip-video-scanoffSkip black and frozen frame detection

It reports three kinds of finding: silences measured from audio energy, review candidates (clipping, black frames, frozen frames), and warnings for conditions worth reading before trusting the run.

Review candidates are never cut automatically. They exist so a human looks.

Filler words are not detected here. A word list matches tokens, not intent: Spanish este is filler in “y este, entonces” and a demonstrative in “en este caso”, and no list survives a language nobody wrote one for. Filler words are proposed by a model through vcut semantic, like every other judgement call.

--audio when the sound was recorded separately. Silence is then measured on that file rather than on the camera track, which matters because the camera track is the one being discarded: cutting against a waveform nobody will hear puts the cuts in the wrong places. The path travels in the report, so edl build writes both sources without being told twice.

vcut detect screen.mp4 --audio mic.wav --preset clean

A source with no video stream is a legal source (#42), and the black/frozen-frame scan skips it automatically. A meeting-recorder mic track, a podcast export, an m4a in an mp4 container — detect probes for a video stream once, and a source without one gets hasVideo: false on the report and a distinct warning (no video stream on this source; black and frozen frame candidates not collected) instead of attempting a scan that has no frame to read. This is separate from --skip-video-scan, which still carries its own wording for the case where a video stream exists but the scan was explicitly skipped. Silence detection, clipping, and word clamping are unaffected.

vcut silences

vcut silences recording.mp4 --from 327.3 --to 330.5 --noise -33 --min 0.08

detect’s silence list is the cutting instrument, at one threshold and one minimum — the preset proven in production, and what edl build cuts against. silences is the placing instrument: the same measurement, a threshold and minimum you choose, over whatever sub-range you name.

It exists because the gap separating a filler from the next word can measure 80-150ms, well under detect’s 0.3s default minimum. Answering “what does the audio do right here, at that resolution” used to mean running raw ffmpeg silencedetect by hand and converting --ss-relative timestamps back to absolute media time yourself, repeated once per boundary.

FlagDefaultWhat it does
--from <sec>0Start of the range to measure
--to <sec>end of mediaEnd of the range to measure
--noise <dB>-30Silence threshold
--min <sec>0.25Minimum silence duration to report

blocks covers the whole requested range in absolute milliseconds, already offset — no arithmetic left for the caller. Never writes an EDL and never changes what gets cut; edl build still cuts against detect.silences.

vcut edl build

Turns a detect report into a draft edit decision list.

vcut edl build --detect detect.json --output master.mp4 --campaign my-video
FlagDefaultWhat it does
--detect <path>requiredReport produced by vcut detect
--output <path>requiredWhere the rendered master will go
--campaign <id>requiredCampaign identifier, stored in the EDL
--edl <path>./edl.jsonWhere to write the EDL
--width, --height, --fpssource valuesOutput geometry
--edge-fade <ms>50Audio ramp at each segment edge; 0 disables
--crop <spec>—top|bottom|left|right:<fraction>, or x,y,width,height
--semantic <path>—Model proposals from vcut semantic
--audio-offset <ms>0Shift the separate audio; positive delays it
--report-json <path>—Write the full JSON build report to this path, regardless of stdout mode

--report-json composes with --human. Without it, a build that prints the readable summary on stdout has no JSON report to hand vcut joins --report, so getting both used to mean running the build twice. --report-json <path> writes the same JSON edl build would otherwise print — the shape vcut commit already writes as rounds/round-N/report.json — to disk in one call, so one build produces both the summary a human reads and the report a later joins call needs.

--crop frames the whole edit at once, which is why it lives here and not per segment. A traditional editor makes you set the frame per clip, so remembering the menu bar after cutting means redoing every segment by hand. Here the crop is one decision applied to all of them, and changing it never touches a cut boundary. Fractions, not pixels, so the same EDL survives a source at another resolution.

The command inverts the cut intervals into the spans worth keeping, so the EDL always describes surviving material rather than deleted material.

A source with no video stream builds a legal, video-less EDL (#42). This used to hard-fail with source has no video stream, which forced every audio-only source through a fake black-video mux just to reach the rest of the pipeline. It no longer does: the built source carries hasVideo: false, output omits width/height/fps/videoCodec/pixelFormat/colorSpace entirely (there is no picture for the V1 contract to describe), and segments are bounded by the audio stream’s own duration. --crop is refused outright on a video-less source rather than silently doing nothing.

The build report includes semanticCuts, one entry per accepted semantic proposal: removedText, the transcript words that fall inside its final span, and boundariesInSilence, whether each edge lands inside a silence detect measured. Read removedText before rendering — it is the corrective for a span drifting onto the wrong words unnoticed, which happened on a real cut: a repetition proposal removed “todos estamos” instead of the stutter “en nuestra propia” because measured blocks were mis-assigned, invisible until a render and a windowed re-transcription caught it. A warning fires when removedText shares fewer than half its carrying words (4+ letters) with the proposal’s reason and has 4 or more of them itself, the same threshold that keeps a short filler cut from firing on a reason that never repeats it word for word.

driftSuspect: true on a semanticCuts entry says removedText sits on drifted cues. removedText inherits transcript drift the same way the whole-file transcript does: detect’s own drift check flags a cue whose claimed start lands inside measured silence, and edl build reuses that exact check, scoped to a span’s own words, rather than reimplementing it. On a recording with 326 drifted cues, removedText cried wolf three times in one run, each costing a say --transcribe to refute. driftSuspect is present and true only on a suspect span, absent (not false) on a clean one, and comes with a matching warning naming the span. It does not re-transcribe anything automatically — vcut peek and say --transcribe already answer that on demand — and on a heavily drifted recording it can flag most or every span, which is detect’s own no-invented-tolerance rule applied at span granularity rather than a bug in the derivation.

It also reports a removal percentage. Compare it against the content type:

ContentExpected removal
Event or interview30-45%
Tutorial or screencast15-25%
Scripted talking head10-20%

A number far below target usually means the source was already edited.

vcut suspects

vcut suspects --detect detect.json

Where to look first, ranked, computed from the silences detect already measured. No transcript, no model, no second pass over the audio.

A speaker correcting themselves breaks delivery into short pauses that land close together; fluent speech spaces them out. The threshold is a fraction of this recording’s own median gap, so it adapts to the speaker instead of needing a number per file. Measured across four recordings: hesitant material fires 5.3 to 6.3 times a minute, a take read from a script fires 1.0, and a speaker whose median gap was 8916ms against another’s 1170ms did not saturate it.

Longer sources fire less per minute rather than more, because a long take carries more thinking pauses and the bar rises with the median: 6.3 a minute at three minutes, 2.8 to 3.5 at four and six.

FlagWhat it does
--detect <path>Report produced by detect (required)
--pause-ratio <n>How close two pauses must be, as a fraction of the file’s median gap (default 0.4)
--limit <n>Return at most this many positions, tightest first

It says where, never what. Telling a discarded retake from a speaker pausing to pick a related thought lives in content, and rhythm is all this measures. Run vcut say --transcribe on a position to find out what is there.

vcut open

vcut open recording.mp4 --preset clean --lang es --transcript words.srt

Opens or resumes a session keyed by the content of the source, not its path: ~/.vcut/sessions/<sha256-16>/. The same bytes at two paths share a session; the same path with new content gets one of its own. Everything inside is disposable cache, not an artifact — the EDL a human approves still lives where they wrote it.

open runs detect once and caches the report. A second open on unchanged media at the same preset reuses that cache instead of re-running ffmpeg (cached: true in the output); a preset this session has never used re-detects and assigns it a new gen.

gen derives from the effective preset, not from whether the immediately previous open differed: a session remembers every preset it has ever used and the gen each was first assigned, so returning to a preset already used returns to that preset’s own generation rather than minting a new one. noisy → clean → noisy reads gen 1, 2, 1 — never 1, 2, 3.

Those silences become refs: the speech blocks between them, numbered b001, b002, … in time order — something a later verb can point at instead of a raw millisecond pair. Refs derive from detect’s own silence list, never from vcut silences.

FlagWhat it does
--preset <name>noisy (-20 dB, default) | clean (-30 dB) | podcast (-35 dB)
--lang <code>Recording language, free-form (default es)
--transcript <path>Caches an SRT into the session and points the cached detect report’s own transcript.path at that copy, so every later reader gets a path guaranteed to still resolve. Without it, open still works — refs come from silences, not words

open’s output is counts, not content: duration, preset, gen, silence and block counts, whether a transcript is cached, and the top 10 suspects (same ranking as suspects, each with the nearest block ref). No spoken text appears anywhere in it. Reading what a ref actually says, and cutting against refs, are later verbs.

vcut peek

vcut peek recording.mp4 --ref b042
vcut peek recording.mp4 --at 550.0 --window 5 --lang es

The four views of one position, aligned in a single call: what the session’s cached transcript claims is there (transcript), what the audio actually says when asked again over the span (heard), the speech/silence shape at fine resolution (blocks, -33dB/0.08s min over the span padded by a second either side), and the level (level). Resolves the session for <media> the way open does, creating it if none exists.

FlagWhat it does
--ref <ref>A block ref from this session’s refs.json (from vcut open)
--at <sec>A position in seconds, instead of a ref
--window <sec>Width of the span when using --at (default 4, centred on --at)
--lang <code>Language passed to the transcriber

viewsDisagree compares transcript against heard on carrying words (4+ letters, the same comparison converge uses) and names transcript-claims-more, heard-more, aligned, or soft-speech-below-threshold — the last one firing when the fine-resolution blocks read silence for the whole span but heard still carries words: speech under the level threshold that neither silences nor detect alone can see. A disagreement is a place to look, not a verdict — a short window transcribes noisily, the same caveat say --transcribe already carries.

vcut cut

vcut cut recording.mp4 --refs b202..b207 --kind tangent --reason "sneeze, speaker says cut it"
vcut cut recording.mp4 --span 0..13.25 --kind tangent --reason "pre-roll before the take begins"
vcut cut recording.mp4 --start-ms 61192 --end-ms 62000 --kind repetition --reason "retake boundary"
vcut cut recording.mp4 --list
vcut cut recording.mp4 --drop 0

Proposes a semantic cut against a session’s own refs, and shows what it removes at propose time rather than after a build. --refs takes a single ref or an inclusive range (b042..b044, from the first ref’s own start to the second’s own end); --span <startS>..<endS> is the escape hatch for a raw span when no ref fits; --start-ms <n> --end-ms <n> is a third, equally first-class way in, taking the raw milliseconds say, silences, and semantic export already emit. All three are mutually exclusive.

FlagWhat it does
--refs <ref[..ref]>A block ref or an inclusive range from this session’s refs.json
--span <s..s>A raw span in seconds when no ref fits
--start-ms <n>Raw span start in milliseconds, the unit say/silences/semantic export emit. Requires --end-ms. Mutually exclusive with --refs and --span
--end-ms <n>Raw span end in milliseconds. Requires --start-ms
--kind <kind>Required: false-start | repetition | tangent | filler
--reason <text>Required, non-empty. Read by a human deciding whether to approve
--listPrint the session’s accumulated proposals with their removedText
--drop <index>Remove the proposal at this 0-based index

--start-ms/--end-ms is not a lesser path than --refs. It takes the same session-tracked route: the proposal accumulates in proposals.json, shows in vcut rounds --diff, and echoes removedText at propose time the same as a ref-based cut. It exists because a finding born in milliseconds — from say, silences, or peek — previously had no way into a session’s refs without converting to seconds by hand for --span, and the gap was wide enough that agents abandoned the session entirely and hand-built EDLs through the stateless pipeline instead. Bounds are validated against the session’s own source duration, and an inverted range is a usage error rather than a silent swap.

The session must already exist — cut never creates one, and a ref from an earlier generation is a usage error naming the ref and the session’s current gen, the same enforcement peek already applies. removedText is quoted from the session’s cached transcript, not re-transcribed. Proposals accumulate in the session’s proposals.json; --list/--drop read and edit that list without hand-editing JSON.

--list (and the accept response, and --drop’s echo) carry the same driftSuspect flag edl build’s report computes, reusing its driftSuspectSpan check rather than reimplementing it: present and true only when a proposal’s own span is built from cues that claim a word starts inside the session’s cached measured silence, absent (not false) when clean. --human prints it as a warning line under the proposal, the same convention edl build’s own warnings use. Scoped to each proposal’s own raw span, not the merged span commit produces once it fuses a proposal with a neighbouring silence cut or another proposal — a clean read here does not guarantee a clean read once commit builds it, only that this proposal’s own claimed words do not already contradict the cached silences. Never persisted to proposals.json: recomputed from the session’s cached transcript and detect report every time a proposal is read back, so it can never go stale relative to the cache it is checked against. Confirm a specific span with peek first if detect’s drift warning fired broadly on this recording.

Proposing and --drop take the session’s advisory lock for the write and release it after; --list never locks. A session already locked by a live process fails naming the holder’s pid, verb, and age — see vcut session below for the full lock story.

vcut commit

vcut commit recording.mp4 --output master.mp4 --campaign my-video

Builds the EDL from a session’s cached detect report and its accumulated proposals, then renders it — byte-identical to running vcut edl build --detect <cached> --semantic <path> by hand, since commit calls the same build seam internally rather than a second implementation.

FlagWhat it does
--output <path>Where the eventual master will go (required)
--campaign <id>Campaign identifier, required
--edl <path>Where to write the EDL (default ./edl.json, the current directory — the user’s artefact, not the session)
--audio-onlyRender audio only, .wav beside the EDL (default)
--videoRender the preview video instead
--fps, --width, --height, --edge-fade, --cropPassed through to the build, same as edl build

Records the round in the session (rounds/round-N/: the EDL copy and the build report); renders and wavs stay out of it. Master mode never happens here. Approval is a human edit to the EDL followed by the existing vcut render --edl <path> --mode master — this command only ever drafts and previews.

On a session opened from a source with no video stream (#42), --video still renders audio. There is no picture to render, so render’s own implied---audio-only behaviour applies underneath commit the same way it would to a standalone render call — a stderr note, not an error.

Takes the session’s advisory lock for the whole build+render, released in a finally. On success, marks the session committed — the signal vcut session gc reads as a candidate to clear, never a trigger that deletes anything itself.

vcut commit recording.mp4 --output master.mp4 --campaign my-video \
  --fields build.removalPercent,build.semanticCuts.removedText

vcut rounds

vcut rounds recording.mp4
vcut rounds recording.mp4 --diff 1 2
vcut rounds recording.mp4 --diff

A session’s own history of what got built. Without --diff, lists every committed round number, ascending. With --diff <N> <M>, compares round N’s build report against round M’s; omitted, diffs the latest two.

FlagWhat it does
--diff [N M]Compare two rounds’ build reports. Omit N and M to diff the latest two

Reports removalPercentDelta, segmentCountDelta, and semanticCuts matched between rounds by span overlap, not array position — a proposal whose edges shifted slightly between rounds (absorbed by a neighbouring cut, re-clamped) still reads as the same cut. Each entry is added, removed, changed, or unchanged.

This diffs what each round’s build asked for (rounds/round-N/report.json, the same data commit writes), not what either round’s render actually says: a text-level diff needs a transcript of each render, and a session never stores renders or their transcripts (cheap to regenerate, expensive to keep). Confirm a semantic diff against the actual renders with peek or say --transcribe before trusting it alone.

The session must already exist with at least two committed rounds for --diff — like cut and commit, this reads a session’s history rather than creating one.

vcut session

vcut session list
vcut session gc
vcut session gc --apply
vcut session gc --older-than 14 --apply

list shows every session under ~/.vcut/sessions/: source path and whether it still exists, size on disk, creation time, committed round count, and whether a live process currently holds its lock.

gc classifies every session against the reasons it could be cleared, without deleting anything unless --apply is given — dry-run is the default, not a flag to remember.

FlagWhat it does
--applyActually delete what was classified deletable. Without it, gc only reports
--older-than <days>Also classify sessions older than this many days. Omitted, age alone never qualifies a session

A session is a gc candidate when:

ReasonCondition
orphanIts source file no longer exists
committedAt least one commit ran successfully against it
older-thanOnly with --older-than <days>, and it qualifies
locked-protectedA live process currently holds its lock — always wins, never deletable regardless of any other reason

The EDL a human approved is never at risk. commit writes it wherever --output/--edl pointed, never inside a session directory gc manages, so clearing a whole session directory can only ever remove the disposable detect cache, transcript copy, refs, proposals, and round history behind it.

Advisory lock. cut’s mutating paths and commit take a lock (lock.json: { pid, startedAt, verb }) before writing and release it after, even on error. Readers (open, peek, cut --list, rounds) never lock. A second writer finding a live pid’s lock is refused with an error naming that pid, its verb, and how long ago it started; a lock whose pid is no longer alive clears itself automatically on the next attempt, with no need for a human to delete lock.json by hand. This is a courtesy between cooperating writers, not a kernel-level guarantee — two writers racing the exact same instant could both pass the check before either writes the file.

vcut semantic

Repeated lines, false starts, digressions and filler words need something reading the transcript. vcut never calls a model. It exports the lines and takes proposals back, so the judgement stays with whoever is reading.

vcut semantic export --detect detect.json > lines.json
# read lines.json, write proposals.json
vcut semantic check --proposals proposals.json --detect detect.json
vcut edl build --detect detect.json --semantic proposals.json ...
SubcommandWhat it does
export --detect <path>Numbered lines with timings, rebuilt into words and split on measured pauses. Every line carries nearestRef, the session’s block ref closest to that line’s own startMs
check --proposals <path> --detect <path>Validates proposals without building
check --review <path>Also fails the round while a repeated phrase goes unnamed
review --edl <path> --detect <path>Reads an EDL back: what survives, and where nobody looked
merge <a.json> <b.json> [more.json...] [--out <path>]Folds two or more proposal files into one, dropping spans identical on all four fields, re-sorted by startMs
--terse (export, review)Omits the instructions block, identical every call and 72% of one measured payload

A proposal is {startMs, endMs, kind, reason} where kind is false-start, repetition, tangent, filler, or non-speech. Every semantic cut lands as semanticRisk: material on the segments around it, so a reviewer can find them without reading all of them. Measured against each proposal’s merged span (after it absorbs a neighbouring silence cut or another proposal), not the raw span proposed — a segment touching the wider, real cut boundary reads material even when its edge sits past where the raw proposal ended.

nearestRef closes the loop back into a session. A proposal written from export goes straight into vcut cut --refs <nearestRef> instead of retyping the line’s raw startMs/endMs — the same lookup vcut open already attaches to suspects. When no ref lands close enough, --start-ms/--end-ms on vcut cut takes the line’s own milliseconds directly.

Nothing malformed passes: an inverted span, a span past the end of the source, an unknown kind, or an empty reason is refused by index and aborts the build. A proposal that vanished between check and build would read as the model choosing not to cut there, which is worse than a refusal.

merge combines two or more rounds of proposals into one. Every proposal file has the same shape, so folding a second round’s proposals.json into the first is a structural operation — exact duplicates on {startMs, endMs, kind, reason} collapse to one, and the result is re-sorted by startMs before it goes anywhere near edl build. --out <path> writes the merged array there, the same shape as an input file, instead of printing the summary object to stdout.

check --review is the gate on a round. Hand it the JSON review wrote and it exits 2 while any phrase in repeated goes unmentioned by every proposal reason. Naming is the bar, not agreeing: keeping a repeat is often right, and saying why in a reason puts the decision where a human approving the EDL can find it. A phrase still present in the render is reported as survivingRepeats and does not fail the check, because a callback repeats on purpose and nothing counting words can tell one from a retake. When repeats are named and kept, the status reads valid-with-kept-repeats and the exit is 0: a finished round, not a pending one.

repeated only lists runs dense enough in content words to read as a candidate retake. A repeated 3-word run built mostly of stopwords (articles, prepositions, pronouns, the copula) is connective tissue, not a repetition anyone needs to answer for — “que la gente”, “va a ser”, and “qué sé yo” all repeat across unrelated sentences in ordinary conversational Spanish and gated a round for nothing before this split. Those land in discountedRepeats instead, alongside the reason (content-word count against the floor), visible but not gating. Language-aware: review reads detect’s lang field and picks the Spanish or English stopword list; anything starting with en (en, en-US, english, …) gets English, everything else gets Spanish, the same default --lang uses everywhere in this CLI.

review closes the loop. With --master it measures silence on the render itself, and with --master-transcript it returns the lines of the render rather than the source projected forward. It also reports unreviewed: the stretches between two cuts that no proposal ever touched, which is where a defect survives round after round because its neighbours look worked on.

vcut setup

vcut setup classifier

Fetches the AudioSet model that skills/core/scripts/non-speech.py uses to find breaths, mic bumps and other audible sound that is not language. Around 320MB into ~/.vcut/panns, and idempotent.

Nothing else needs it: detect, edl build and render all run without it. vcut doctor reports whether it is installed, as optional rather than missing.

vcut render

vcut render --edl edl.json --mode preview --dry-run
vcut render --edl edl.json --mode preview
FlagDefaultWhat it does
--edl <path>requiredThe EDL to render
--output <path>from EDLOverride the output path
--mode <name>previewpreview or master
--audio-onlyoffRender the audio alone, for iterating
--dry-runoffPrint the ffmpeg command without running it
--quietoffSkip the progress lines on stderr

preview accepts proposed segments. master refuses unless the EDL is approved, every segment is approved, every source hash still matches, and the output path is free. It will not overwrite.

render blocks in the foreground until ffmpeg exits. There is no background mode: the call you make is the render, start to finish. While it runs, one progress line lands on stderr per report from ffmpeg’s own -progress output — time rendered, percent of the EDL’s own duration, encode speed — so a render that takes tens of seconds to minutes never sits silent. stdout stays reserved for the result: nothing to poll a file for, nothing to grep a process table for. --quiet drops the progress lines and renders exactly the same file.

After rendering, vcut probes the file it produced and validates it against the EDL. A mismatch fails the run instead of shipping a bad file.

Iterate with --audio-only. Nearly every question a round of edits asks is about sound, and answering it through the video path re-encodes every frame for nothing. Audio-only costs roughly 1 second per 14 seconds of audio the cut keeps (set by kept audio, not segment count), against a video render that runs near real time per minute of source. The audio graph is unchanged, edge fades and loudness included, so what you hear is what the finished render will sound like: -16.4 LUFS on both paths from the same EDL. It writes lossless audio, because a codec artifact heard while iterating reads as a defect in the cut. Refused in master mode.

The result runs a few tens of milliseconds short of the segment sum (31ms on a 54.6s cut). That is loudnorm latency draining trailing decay, not missing material; a video render hides it because the picture sets the container duration.

The audio-only verification loop: render --audio-only, then run vcut audit and vcut nonspeech against that same .wav. Both read the waveform only and accept it wherever they accept a video render, so a full video mux per round is dead wall clock. Render video once, at the end, for the master.

A source with no video stream is a first-class source (#42). A meeting-recorder mic track, a podcast export, an m4a in an mp4 container — anything edl build built without a video stream — renders --audio-only by implication, not by requirement: the flag’s absence produces a stderr note (this EDL has no video source; --audio-only is implied) rather than an error, since there is no picture to render. --mode master on a video-less EDL produces an audio master (AAC, the same codec the video path’s own audio track already uses) instead of a video; the “refused in master mode” rule above is specific to a video-bearing EDL, where a scratch render and a finished video really are two different things. Approval semantics are unchanged.

vcut locate

vcut locate --edl edl.json --master 50.2 --explain
vcut locate --edl edl.json --source 80.07
vcut locate --edl edl.json --sources 20,53.86,61.2      # several at once
vcut locate --edl edl.json --all

Translates between a position in the master and the source it came from.

Positions are seconds. The JSON that comes back speaks milliseconds, so passing those back in is the natural mistake, and it used to be answered as if it made sense: a run asked about nine positions in milliseconds, got removed: true for all nine, and read that as nine spans it had cut. Both flags now refuse a position past the end of the file and name the unit they expected. --sources takes a comma-separated list, which is a round asking about every boundary it proposed without a shell loop around it.

Do not derive this by hand. Accumulating outMs - inMs across segments gives a total that can match the rendered file to the millisecond while individual positions land seconds away, and nothing in that agreement warns you. --explain reports the neighbourhood a position sits in, and --render <path> measures the file rather than trusting the EDL, which records intent.

master 50.200           -> source 84.239  (segment-020)
segment                 source 83.942-85.308, 0.297 in
previous                segment-019 ends master 49.903
cut before it           0.367 of source removed

Asking --source about material that was cut reports it as removed with the next surviving segment, rather than failing.

vcut audit

vcut audit --edl edl.json --render cut.wav

Every check the renderer runs on itself is an aggregate: dimensions, frame count, duration. A render whose segments carried the wrong material passes all of them, because the durations are right whatever ended up inside them. This compares the audio itself, segment by segment, against the source span the EDL points at.

audit  22 of 22 segments compared
  agreeing         21 at or above 0.8 correlation
  segment-022      correlation 0.330 at master 52.186 (source 86.842)

--render accepts an audio-only render. Every comparison decodes a waveform, never a frame, so the .wav vcut render --audio-only writes is enough for every round — no reason to hold this check for a video render.

A low score is a place to look, not a verdict. Envelope correlation is weak over short or quiet windows, and loudness normalisation lifts quiet passages by several dB. On the run above, the segment that scored low was carrying exactly the right words. It reports rather than fails, and stays out of render, for that reason.

vcut joins

vcut joins --edl edl.json --render cut.mp4 --report report.json --lang es

The post-render twin of edl build’s removedText: one call that verifies every semantic join instead of N x (locate + say --transcribe). On a real 11.7-minute run, verifying 9 joins by hand cost about 14 calls.

Each join is the EDL segment that opens right after a semantic cut, derived the same way edl build’s own boundariesAfterSpeech finds it — by the kind of cut, not a distance. Runs on --render, never the source, checks the EDL’s own master-time total against the render’s measured duration first, then re-transcribes a window around each join and reports a reading: lands, removed-text-leaked (the window’s carrying words majority-overlap the cut’s removed text), or check-by-ear (the window carries too little to judge).

joins  cut.mp4
  64.83s (segment-020)   lands  "Que haga lo que tú exactamente querías, pues entonces el workaround del"
  160.18s (segment-047)  removed-text-leaked  "En este principio de agregar, agregar verificabilidad a los"

removed-text-leaked is a place to look, not a verdict. On the run above, that reading was a false positive: the cut’s removed text was the speaker stumbling on the same phrase three times before landing it, and the surviving sentence legitimately reused that phrase as its real content. A wider vcut say --transcribe --window 8 confirmed the join read clean — joins names that exact command in next.

--report <path> (a build report from edl build, commit, or the report.json commit writes into rounds/round-N/, the default lookup beside --edl) adds removedText, reason, and driftSuspect to each join. Its absence is a supported state: joins still runs, with those three fields null.

vcut joins --edl edl.json --render cut.mp4 --report report.json \
  --fields joins.reading,joins.joinMasterMs,joins.removedText

vcut say

vcut say cut.mp4 --transcript cut.srt --at 50.2 --edl edl.json    # read the transcript
vcut say cut.mp4 --transcribe --lang es --at 57.5 --window 4      # ask the audio
vcut say cut.mp4 --transcribe --positions 19.5,30.0,41.9          # sweep several positions

Reads back what is spoken at a position, with the level there and, with --edl, which segment it falls in.

FlagWhat it does
--at <sec>Position to read around, or the start of a range with --through
--through <sec>Read everything from --at to here rather than a window around it
--positions <list>Several positions at once, comma-separated seconds. One object per position, same shape --at returns, in order. Mutually exclusive with --at/--through
--transcript <path>Word-level SRT to read from (required unless --transcribe)
--transcribeCut the window and run the transcriber over it instead of reading
--lang <code>Language passed to the transcriber (--transcribe only)
--window <sec>How much context to include (default 2)
--media <path>Media to measure level on, if not the positional argument
--edl <path>Report which segment the position falls in

Reading is the default and the cheap path. A window under about two seconds transcribes as noise regardless of what the audio holds, so a nonsense result from a slice cannot tell a real word from a guess. The existing transcript already knows.

--transcribe is for the case reading cannot answer. A whole-file pass averages: where a speaker said a line three times it can write it once, and no amount of re-reading recovers the difference. Measured on one recording, reading at 57.5s gave “la que conocemos, ya llegamos a” where transcribing the same window gave “Y a la que conocemos, ya llegue. Y a la que conocemos” — the repetition four runs failed to find. Use a window of four seconds or more, and note it costs one transcriber call. vcut still calls no model of its own: it runs the transcriber already on your PATH, the same way it runs ffmpeg.

A window with no words but real level is the case worth stopping on: something audible the transcript never saw, which is what the non-speech classifier is for.

--positions answers several windows in one call, because sweeping several spans was a shell loop of individual --at calls: one session swept 18 classifier spans exactly that way. With --transcribe, positions transcribe strictly sequentially, never concurrently — each call loads a Whisper model into memory, and racing several is the load that chokes a machine already carrying a video editor.

vcut converge

vcut converge source.mp4 --phrase "a la que conocemos" --from 59 --lang es

Finds where a repeated phrase stops coming back, which is the boundary of a retake. Steps a window forward from --from, transcribing each one, and reports the first that no longer carries the phrase along with every window it read getting there.

FlagWhat it does
--phrase <words>The wording that keeps recurring (required)
--from <sec>Where to start stepping (required)
--to <sec>Where to give up (default: 12s past --from)
--step <sec>How far to move each try (default 0.5)
--window <sec>How much audio each try transcribes (default 3.5)
--lang <code>Language passed to the transcriber

It exists because that judgement went wrong more often than any other: three runs cut the same retake at 61000, 61020 and 61192ms, all about 1772ms short, and each had verified its number. Every attempt at a retake says the same words, so a window opened anywhere inside one comes back complete and convincing.

boundaryMs is not where to cut. A retake and the telling that survives it overlap, so the point where the wording disappears sits past the start of the line worth keeping. Cutting to it beheads that line: on one recording, ending at the reported 62000ms gave “Conocemos, ya llegamos a mil miembros” where ending at 61192ms kept “Y a la que conocemos, ya llegamos a mil miembros”. Both were rendered and listened to; neither transcript reads as broken. lastWithPhraseMs carries that telling in full and sat 308ms from the correct boundary against 808ms for the far edge.

Exit 1 with a null boundaryMs means the phrase was still recurring at --to, which is a reason to widen the span rather than evidence there is nothing to cut.

vcut nonspeech

vcut nonspeech cut.wav                           # spans only, the classifier's own output
vcut nonspeech cut.wav --verify --lang es         # each span read back through a window

Runs the bundled classifier (skills/core/scripts/non-speech.py) against a rendered preview and reports audible sound that is not language: a breath, a mic bump, a stretched hesitation the transcript cleans away even with a verbatim preset. Run it on the render, not the source: on raw footage every pause scores as non-speech, correctly and uselessly.

The render can be the --audio-only .wav. The classifier and --verify both work from audio alone, so there is no reason to hold this check for a video render — use it every round.

FlagWhat it does
--verifyRe-transcribe a window around each span with trx and attach a reading
--lang <code>Language passed to the transcriber (--verify only)

--verify is not optional in practice. Without it you get positions and nothing else, and closing each one against the whole-file transcript is circular: that transcript is exactly the instrument that could not see this class of sound. --verify cuts a window of the span plus 1.2s of context on each side and re-transcribes it, attaching text, peakDb, meanDb, and a reading:

  • vocalization-suspect — the window names a hesitation sound (eh, ehm, mmm, aah, tolerant of a stretched vowel), or the span carries real level with no words inside it.
  • words-around — the window transcribes to ordinary words either side of the span: a breath in a pause.
  • empty — no words and no real level.

Measured on a real 7.5-minute run: 18 spans closed by reading the whole-file transcript were all read as breaths, and seven turned out to be audible “eeeh” fillers a listener caught on the first playback. --verify against the same render named them by their text instead.

The classifier is optional: python3, panns-inference/scipy/numpy, and a ~320MB model under ~/.vcut/panns fetched by vcut setup classifier. Its absence is a supported state — nonspeech reports it and exits 0, the same policy vcut doctor already applies — and invariant 7 falls back to a human ear. --verify additionally needs trx on PATH. vcut still calls no model of its own: python3 and trx are binaries on the caller’s PATH, exactly like ffmpeg.

vcut schema

vcut schema            # lists the commands with a contract
vcut schema detect     # the field-by-field contract for detect

Versioned, so an agent can introspect the output shape at runtime instead of parsing help text or reading source.

vcut skills

Install the skill into Claude Code, Cursor, or any agent that reads them:

npx skills add Railly/vcut

What gets installed is a thin stub. It carries the description an agent matches against and then points at the CLI:

vcut skills list
vcut skills get core                  # the small, always-loaded usage guide, raw markdown on stdout
vcut skills get core --section cut    # one deep-dive section instead of the whole manual
vcut skills path                      # filesystem path to the skills directory (or one skill's directory)

The guide ships inside the npm package and is served by the CLI itself, so it always matches the installed version. A copy pasted into an agent’s config would go stale the moment you upgrade; a stub that points at skills get cannot.

The manual is sectioned. vcut skills get core used to serve one document, every command’s full detail concatenated, whether the clip being edited needed any of it or not — a fixed ~35-40k token tax paid up front regardless of clip length. core is now small: the flow, the invariants, and enough to run a first edit. Everything past that — per-command detail, the methodology for working a round, why the classifier needs a model and not a statistic — lives in a section, loaded only when a caller actually needs it:

vcut skills get core --section cut     # ref ranges, --start-ms, --span, --list/--drop
vcut skills get core --section render  # audio-only, progress, loudness, reproducibility

vcut skills list prints the available sections with a one-line blurb for each — what question it answers, not just its name, since cut or joins alone does not tell a caller whether their question is inside it. A section name that does not exist is refused with the list of ones that do, rather than a bare error.

vcut doctor

Checks that ffmpeg and ffprobe are reachable and reports their versions. Exits non-zero when either is missing.

It also reports the optional non-speech classifier, which is a supported absence rather than a failure: without it the check it performs falls back to a human ear.

And it reports sessions: how many exist, their total size, how many are orphans (source file gone), and the oldest one’s creation date — one line in human mode, a sessions object in JSON. An absent sessions directory reports zero across the board rather than an error: a machine that has never run vcut open has nothing wrong with it. This is the detector class that was missing when a different cache directory grew to 609MB unnoticed — vcut session gc is the fix once doctor reports something to clear.

Exit codes

CodeMeaning
0Success
1The run failed
2The invocation was wrong

One command overloads 2 deliberately: semantic check --review exits 2 when a repeated phrase in the round went unnamed by every proposal reason. The invocation was fine; the round was not finished. An agent driving the loop should treat that case as “answer the repeats and run again” rather than as a usage error.

Data always goes to stdout, diagnostics always to stderr.

The cutting loop

Silence removal is the first round of several, not the job. Most of what needs cutting is invisible until the noise around it is gone, so the work is a loop: cut, render, transcribe the render, read it again, cut again.

Each round sees what the last one uncovered

RoundOnly visible now because
1Nothing hides long silence or an obvious stammer
2A pause two adjoining segments create together did not exist in either of them before
3A join reads as broken only once both sides are adjacent, and a surviving redundancy only once the passage is short enough to hold in your head
4A discourse marker is inaudible inside loose speech and obvious inside tight speech

Stopping after one round leaves work that looks like polish and is not.

The round

A round starts with a session, opened once:

vcut open recording.mp4 --preset clean --lang es --transcript words.srt   # once

open caches the detect pass and turns its silences into stable block refs (b001, b002, …). From there, cut and commit are the round — no proposals.json to open and hand-edit, no --detect/--semantic paths to re-type each pass:

vcut cut recording.mp4 --refs b042..b044 --kind repetition --reason "..."  # per finding
vcut commit recording.mp4 --output master.mp4 --campaign my-video          # builds + renders

trx transcribe master.wav --words --language es -m large-v3-turbo --output-dir "$(dirname master.wav)"

vcut semantic review --edl edl.json --detect detect.json --terse \
  --master master.wav --master-transcript <the .srt trx wrote> > review.json

vcut rounds recording.mp4 --diff   # what changed since the last round

commit renders --audio-only by default and records the round in the session’s own rounds/round-N/ — no N=1 counter to bump by hand, and vcut rounds --diff compares it against the previous round without you tracking which file was which.

A finding enters the session no matter which coordinate system it was born in. That is the whole point of running the round through cut rather than hand-building an EDL: --refs takes a block ref, --span takes raw seconds, and --start-ms/--end-ms takes raw milliseconds — the unit say, silences, peek, and semantic export already emit. A line from semantic export carries its own nearestRef and goes straight to vcut cut --refs <nearestRef>; a boundary from vcut converge or vcut silences comes back in milliseconds and goes straight to vcut cut --start-ms <n> --end-ms <n>. Nothing you found has to be converted by hand or reasoned about in terms of which workflow “fits” it — every coordinate system has a direct path into cut.

vcut audit and vcut nonspeech --verify belong inside this loop, against the same --audio-only .wav commit already rendered: neither reads a frame, audit correlates waveforms and nonspeech classifies audio, so holding them for a video render answers no question either one is asking. Render video once, at the end, for vcut joins and for the human who watches the file.

vcut joins --edl edl-$N.json --render cut-$N.mp4 --report report-$N.json --lang es replaces the locate + say --transcribe round for every semantic cut in one call — a real 11.7-minute run verified 9 joins that way for about 14 calls, before joins existed. Each reading of removed-text-leaked or check-by-ear is a place to look, not a verdict: confirm with the wider say --transcribe window next names before folding anything back into a proposal.

Close a nonspeech hit with --verify, not by reading the whole-file transcript. That transcript is exactly the instrument that could not see this class of sound in the first place, so checking a hit against it is circular: measured on a real 7.5-minute run, 18 spans closed that way were all read as breaths and seven were audible “eeeh” fillers a listener caught immediately. --verify re-transcribes a short window around each span instead and reports which are vocalization-suspect, words-around, or empty.

Read semanticCuts[].removedText in the build’s own JSON before rendering. Every accepted proposal reports the transcript text its final span actually removes, which can drift from what the proposal asked for once it merges with a neighbouring cut. A repetition cut once removed “todos estamos” instead of the intended “en nuestra propia” this way, invisible until a render and a re-transcription caught it — removedText makes that visible at build time. edl build also warns when a span’s removed text and its own reason share too little in common.

--terse drops the instructions block, which is identical every round and was 72% of one measured payload. Read it once on the first call, then leave it out.

The stateless pipeline is an escape hatch, not an alternative

Before a session, the same round was edl build + render + hand-edited proposals.json, called directly:

N=1   # bump every round: the renderer refuses to overwrite

vcut edl build --detect detect.json --semantic proposals.json \
  --output cut-$N.mp4 --campaign my-video --edl edl-$N.json
vcut render --edl edl-$N.json --audio-only --output cut-$N.wav

trx transcribe cut-$N.wav --words --language es -m large-v3-turbo --output-dir "$(dirname cut-$N.wav)"

vcut semantic review --edl edl-$N.json --detect detect.json --terse \
  --master cut-$N.wav --master-transcript <the .srt trx wrote> > review-$N.json

vcut semantic check --proposals proposals.json --detect detect.json \
  --review review-$N.json          # exit 2 while a repeated phrase goes unnamed

This still works, commit calls the identical build seam internally, and it is not deprecated. But it is the exception, not the default: use it for a one-off cut with no second round, or a script with no long-lived working directory — a CI job that renders once and exits has nothing to gain from a session it will never revisit. Anything that expects more than one round should open a session and stay in it.

A fresh agent given the whole manual with no steer on this ran vcut open once, at the start, and then never called peek, cut, or commit again — it read the session’s own JSON files under ~/.vcut/sessions/<id>/ directly and hand-built EDLs through the stateless path for every round after. Its own retro named the cause precisely: which path it took “is decided by which coordinate system your finding happens to be in, not by which workflow is actually better for the edit.” A finding in milliseconds had an obvious raw---span path and no obvious ref path, so it took the path that needed no lookup, one round at a time, until the session was doing nothing but sitting unused beside a hand-rolled pipeline that duplicated its job.

Reading ~/.vcut/sessions/<id>/ files directly is unsupported. cut, commit, peek, and rounds are not a thin wrapper around that directory — they check the block generation (gen) against what a ref was resolved from, resolve stale refs as a named usage error instead of silently against boundaries that no longer exist, take the advisory lock before a write, and patch cached transcript paths so a moved file still resolves. None of that runs when a file under a session directory is opened and read by hand. The layout itself is free to change between releases precisely because the verbs are the contract, not the files backing them.

Where to look first

vcut suspects --detect detect.json

Ranked positions computed from the pauses detect already measured, no transcript involved. On a short take, read every line the export gives you. On anything long, this is the order to read in: measured across four recordings it fires 5.3 to 6.3 times a minute on hesitant material and 1.0 on a take read from a script, and the rate falls as sources get longer rather than rising.

When a round finds a retake, the boundary is its own question and the one that goes wrong most:

vcut converge source.mp4 --phrase "the recurring words" --from 59 --lang es

Three runs cut the same retake at 61000, 61020 and 61192ms, all about 1772ms short, each having verified its number against a window that read like a clean start. What the command reports is the far edge of what is safe to remove; the cut usually ends nearer lastWithPhraseMs, where the telling being kept begins. Ending at the far edge on one recording gave “Conocemos, ya llegamos a mil miembros” instead of “Y a la que conocemos, ya llegamos a mil miembros” — both transcripts read fine, and only one sounds right.

It replaces neither the reading nor the loop. A repetition delivered fluently leaves no rhythmic trace and only the prose shows it. What it replaces is scanning a file you have not read to decide where to spend attention.

Once you know roughly where a boundary goes, placing it exactly can need a finer measurement than detect gives you:

vcut silences source.mp4 --from 327.3 --to 330.5 --noise -33 --min 0.08

detect’s silence list is the cutting instrument, at the preset threshold and a 0.3s minimum — the one edl build cuts against. silences is the placing instrument: same measurement, a threshold and minimum you choose, over the sub-range you name. The gap separating a filler from the next word can measure 80-150ms, under detect’s default minimum and invisible to it; answering “what does the audio do right here” used to mean running raw silencedetect by hand and converting its range-relative timestamps back to absolute media time yourself, once per boundary.

Iterate on audio

Every question in the round above is about sound: whether a filler survived, whether a boundary clipped a word, whether a pause is still there. The picture cannot answer any of them, and rendering it costs far more wall clock than the round needs. Audio-only is much cheaper but not instant: roughly 1 second per 14 seconds of audio the cut keeps, set by kept audio rather than segment count, with loudness normalisation accounting for most of it. Because every commit re-renders the whole cut, batch the cuts you have evidence for into one commit instead of committing per find.

--audio-only uses the same audio graph the video path uses, edge fades and loudness included, so what you hear while iterating is what the finished file will sound like.

Check the transcript path trx reports rather than assuming it: it names the file after its own normalisation step, so the .srt beside cut-1.wav can arrive as cut-1_clean.wav.srt.

Audit the render before you call it done

vcut audit --edl edl-$N.json --render cut-$N.wav

Everything the renderer validates about itself is an aggregate, so a render whose segments carried the wrong material passes every one of those checks. audit compares the audio segment by segment against the source the EDL points at, and names what to inspect. Read the words at any position it flags before believing the number.

--render takes the --audio-only .wav directly — the comparison decodes a waveform on both sides, never a frame, so there is nothing a video render would add here.

Verify every semantic join in one call

vcut joins --edl edl-$N.json --render cut-$N.mp4 --report report-$N.json --lang es

Every accepted semantic cut has a join: the EDL segment that opens right after it, the same boundary edl build’s boundariesAfterSpeech warns about at build time. Checking each one by hand is locate to find the master position, say --transcribe to hear the window — once per cut. joins derives every join from the EDL itself, checks the EDL’s own master-time map against the render’s measured duration, and re-transcribes each window in one call: on the real recording that motivated this, 8 joins came back in about 15 seconds.

A removed-text-leaked reading means the window’s carrying words majority-overlap the cut’s removedText — a real signal, and also one that can misfire: a false-start whose removed text is the speaker repeating a phrase before landing it can leave a surviving sentence that legitimately reuses the same words. Confirm before folding anything back in.

Transcribe the render every round

Not the previous transcript, not the source transcript projected forward. Every cut shifts everything after it, so the two timelines diverge by the whole removed duration and a span written against stale timings lands somewhere nobody chose.

The fresh transcript is also the only place a mangled join is visible as text: the source describes what was said, only the render describes what is left.

Read unreviewed first

A pass reads what it went looking for, so cuts land where the attention was and the stretches between two cuts are where nothing was ever read. They look reviewed because their neighbours are, and that is where a marker survives round after round while both spans around it get examined closely.

vcut semantic review reports those stretches with their text. Work that list before scanning anywhere else.

Widen before adding

When a proposal fails to remove what it named, the usual cause is a boundary set too tight, not a wrong call. A restart is only obvious once you see the attempt that follows it, so the earliest attempts read as content while you are looking at them and as preamble once the last one is in view.

Extending an existing span usually removes more than any new cut placed beside it.

Invariants

Hard rules: each is a defect if it survives a pass, not a matter of taste. What makes them rules is that they are stated about the render rather than about the plan, so they can be checked after the fact instead of argued before it.

  1. No idea is stated twice. Distance between two passages is not evidence they differ; the edit removes that distance.
  2. No sentence begins and does not land.
  3. No pronoun outlives its antecedent.
  4. No fragment survives alone. A clause that only made sense inside a removed passage is a leftover.
  5. Nothing survives that can be deleted without changing what the sentence says. Delete the candidate, read what remains, ask whether a listener learns anything less.
  6. The last line lands. Ending on an abandoned start is worse than ending four seconds sooner.
  7. Nothing audible is left that is not language. A breath, a mic bump, a lip smack.
  8. Every stretch has been read at least once.

Being a rule is not the same as being mechanical. Only rule 8 is machine-decidable: review prints the list and either it is empty or it is not. Rule 7 is decidable when the classifier is installed and a listening task when it is not. Rules 1 through 6 are read by judgement, and their value is in naming a defect precisely enough that you can tell whether you looked for it.

Stop when a round proposes nothing, not when the removal percentage looks respectable, and not when the rounds start finding less. A round that finds three things instead of ten is still a round that found something, and what it found was invisible until the previous one ran.

Filler words are a deletion test, not a list

vcut used to carry a list of six tokens per language. Measured on one Spanish recording, that list caught 3 spans while the finished cut still carried 19 fillers in 332 words. What it missed were ordinary words that happened to carry no meaning in that one sentence, which is most of them.

A list also cannot tell filler from real use. Spanish este is filler in “y este, entonces” and a demonstrative in “en este caso”. Extending the list makes it worse, not better: the same token is filler or content depending on the clause around it, and a list has no clauses in it.

The test that works is a deletion. Read the line without the candidate and ask whether a listener learns anything less. That needs no vocabulary, so it works on a construction nobody named and in a language nobody wrote a list for.

Two things fail the test and must stay anyway: a word carrying emphasis the speaker meant, and a beat that gives a listener room before a heavy point. Removing those is what makes an edit sound like a script.

What has already been tried and failed

Non-verbal sound needs a classifier, not a statistic

A breath is audible, meaningless, and invisible to both instruments: the silence pass hears energy and calls it speech, and the transcript has no word for it.

Four energy statistics were tried and all four failed:

AttemptWhy it cannot work
Sound with no word covering itThe transcript stretches every cue to the next word, so the noise lands inside a word’s span
Gaps between consecutive wordsThe largest gap in a tight edit is a fraction of a second
Energy swing inside one wordA word holding a breath swung less than an ordinary word did
Median level inside one wordRanks unstressed function words first, which is a different question

Each measures a proxy for non-speech, and every proxy is dominated by ordinary variation in speech. Periodicity gets closer, since voiced speech has vibrating folds and a breath is turbulence, but unvoiced consonants are turbulence too and every sibilant becomes a false positive.

A general voice-activity detector is not enough either: one scored a breath at 0.87 voice, indistinguishable from words. What worked was an AudioSet classifier keyed on the absence of speech rather than the presence of breathing. vcut nonspeech runs that classifier (skills/core/scripts/non-speech.py) as a subprocess rather than a built-in dependency: it needs Python and a 300MB checkpoint, and vcut otherwise runs anywhere ffmpeg does. --verify, which turns a raw span into a reading, lives in the CLI itself and needs no Python beyond what the classifier already needed.

Noise reduction is not offered

Measured on one recording, the background floor already sat at -54 dB and a denoiser at a default setting pushed a weak syllable from -45 dB to -57. That is the same defect as a threshold set too high, and it lands on exactly the words a silence detector already struggles with.

There is no safe default because the right amount depends on the room, and unlike a cut it cannot be undone by editing the EDL. Loudness normalisation is the part that is safe to automate, and it runs by default.

A better transcript can mean fewer cuts

Worth knowing, because it looks like a bug the first time. whisper --max-len 1 stretches each cue to the start of the next word, so a word’s range routinely swallows the pause that follows it.

On one recording, 118 of 119 detected silences overlapped some word’s range. Clamping every cut strictly inside word boundaries erased 57 of them outright, and removal collapsed from 10.5% to 2.7% purely because a better transcript was supplied.

vcut treats silence measured from audio energy as stronger evidence than a boundary inferred by a model. When clamping would shred a cut into a remainder shorter than your own --min-silence, that remainder is overlap residue rather than a real pause, and the measured span wins.

Ask for a large transcription model

One cue per word means one cue per token, and what counts as a token depends on the model. Measured on the same three minutes of Spanish: small returned 26% of its cues as word fragments, splitting “Crafter” into Cra + fter, while large-v3-turbo returned 0% and cost 13 seconds.

Fragments weaken the word clamping that keeps cuts off speech, and they make the semantic export unreadable.

When a word loses its opening sound

A silence detector decides by level, so a soft consonant under the threshold is cut like a pause. This reads as a transcription error rather than as a cut, which is what makes it hard to spot: the word comes back missing its first syllable.

Measured on one recording, the same word appeared six times. The three that survived intact sat at -10 dB; the three that lost their opening consonant sat at -43, -56 and -60 dB, with no overlap between the groups. Cr is an unvoiced stop, so it begins with 30 to 60 milliseconds of genuine silence before the burst, and at -60 dB that is indistinguishable from a pause.

The fix is a lower threshold or a better recording, not a larger margin: a margin pads around a cut that should not have been there.

Compression does not rescue it. On that recording the weak syllable sat 25 dB below the room’s own noise floor, and a compressor lifts the noise with the syllable, so the distance between them never changes. Measured: removal fell from 31% to 15% because the raised noise stopped registering as silence at all.

The EDL

The edit decision list is the artifact between detection and rendering. It exists so there is something a human can read and disagree with before any file gets written.

It describes the material that survives, not the material that gets deleted.

Shape

{
  "version": 1,
  "campaignId": "my-video",
  "createdAt": "2026-01-15T18:00:00.000Z",
  "timebase": "milliseconds",
  "sources": [
    {
      "id": "recording-mp4",
      "path": "/absolute/path/recording.mp4",
      "sha256": "…",
      "durationMs": 381760,
      "hasVideo": true,
      "hasAudio": true
    }
  ],
  "segments": [
    {
      "id": "segment-001",
      "sourceId": "recording-mp4",
      "inMs": 1100,
      "outMs": 4633,
      "reason": "approved-line",
      "handlesMs": { "before": 100, "after": 100 },
      "approval": "proposed",
      "semanticRisk": "none",
      "crop": null
    }
  ],
  "audio": {
    "speechTargetLufs": -16,
    "truePeakMaxDbtp": -1,
    "externalAudioSourceId": null,
    "syncOffsetMs": 0,
    "edgeFadeMs": 50
  },
  "output": {
    "path": "/absolute/path/master.mp4",
    "width": 1920,
    "height": 1080,
    "fps": 60,
    "videoCodec": "h264",
    "pixelFormat": "yuv420p",
    "colorSpace": "bt709",
    "audioTrackPolicy": "required",
    "overwrite": false
  },
  "approval": {
    "status": "draft",
    "approvedAt": null,
    "approvedBy": null
  }
}

The full JSON Schema ships in the package under schemas/edl.schema.json.

Audio-only sources

A source with no video stream — a meeting-recorder mic track, a podcast export, an m4a in an mp4 container — produces a legal EDL, not an error. Its source entry carries "hasVideo": false, and output carries only path, audioTrackPolicy, and overwrite: width, height, fps, videoCodec, pixelFormat, and colorSpace are all absent, since the V1 video contract they describe has no picture to apply to on this EDL. durationMs on the source is the audio stream’s own duration, since that is what segments are actually trimmed against.

{
  "sources": [
    {
      "id": "meeting-mp4",
      "path": "/absolute/path/meeting.mp4",
      "sha256": "…",
      "durationMs": 1_145_000,
      "hasVideo": false,
      "hasAudio": true
    }
  ],
  "output": {
    "path": "/absolute/path/master.m4a",
    "audioTrackPolicy": "required",
    "overwrite": false
  }
}

render reads the absence of a video source and implies --audio-only rather than requiring it; --mode master on an EDL shaped this way produces an AAC audio master instead of a video. --crop is refused at build time on a video-less source rather than silently accepted and ignored.

Approval

Two levels, and both start closed.

  • segment.approval is proposed, approved, or rejected.
  • approval.status on the EDL is draft, approved, or rejected.

vcut render --mode preview accepts proposed segments, so you can watch the cut before committing to it. --mode master requires the EDL approved, every segment approved, and an approval identity recorded.

vcut never writes approved. There is no flag for it and no --yes. Changing it is a human act, whether by hand or through a tool a human drives.

What the renderer refuses

A master render aborts on any of these:

  • the EDL or any segment is not approved
  • a source file is missing
  • a source hash no longer matches what the EDL recorded
  • the output path already exists
  • a segment references an unknown source or an interval outside the source duration
  • a crop falls outside the frame
  • audioTrackPolicy is required but a source has no audio

Hashing sources and refusing on a mismatch is what makes an approval mean something: you approved that footage, not whatever now sits at that path.

Audio recorded separately

externalAudioSourceId names a second entry in sources that carries sound and no picture. The renderer reads every segment’s audio from there instead of from the video, which is the case a separate microphone exists for.

syncOffsetMs corrects two recorders that did not start together. It slides the window the audio is read from rather than shifting the audio afterwards, because shifting changes the length and the length is what the duration contract checks. Two recordings started by the same app share a clock, so zero is the common case.

edgeFadeMs ramps the last and first milliseconds of each segment to zero. It is not a crossfade: overlapping the two sides would shorten the render against concatenated video and drift the audio out of sync, so each side fades within its own segment.

speechTargetLufs and truePeakMaxDbtp are applied to the concatenated result rather than per segment, so a quiet passage stays quieter than a loud one instead of every piece being dragged to the same number.

Fields that are rejected, not ignored

An externalAudioSourceId naming nothing, or naming a source with no audio, is refused rather than silently falling back to the camera track. So is a crop outside the source bounds, a segment whose source is unknown, and a master render whose EDL is not approved.

A tool that silently drops a field you set is worse than one that refuses to run.

Self-validation

After rendering, vcut probes the file it just produced and compares it against the EDL: dimensions, pixel format, colour metadata, decoded frame count within one frame, sample rate, channel count, and the audio track contract.

A render that quietly produced two extra frames is a bug, and without this check it ships as a working file.

Reproducibility

Renders pin the thread count, fix the creation timestamp, and avoid anything nondeterministic, so the same EDL produces a byte-identical file. The sha256 in the render result exists so you can verify that yourself rather than take it on faith.

Frame boundaries

Cut points are milliseconds, but frames are not whole milliseconds: at 60fps a frame is 16.666…ms. ffmpeg rounds each trim to the nearest frame on its own, and near a frame edge that rounding can go either way.

vcut places every boundary in the middle of its frame, as far as possible from the point where the rounding flips, so the same frame is chosen regardless of how many segments the EDL has. Aiming at the frame edge instead held for eight segments and drifted past tolerance at ten.

The end of the source is clamped to the real duration rather than snapped: a segment must never claim material past the end of the file.