Skip to content
SUSHIL
← All work

2026 · Product

№09 — Tool

Shorts Studio

Paste a YouTube link, get a publish-ready Short: transcription, clip picking, framing, captions and scheduled uploads on one GPU.

Livetypescriptpythonwhisperxffmpeg
Role
Solo
Runs on
One consumer GPU
Stack
TypeScript · Python · ffmpeg
Source
Open on GitHub
A long podcast video and its YouTube link turning into three vertical Shorts: a single speaker with captions, a split screen of two speakers, and gameplay with the player below.

// 2 min read

What it does

Paste a YouTube link and get a publish-ready Short. Shorts Studio takes a long video, finds the moments worth clipping, edits them into vertical 9:16 clips with layouts, animated captions and thumbnails, and can upload them to several YouTube channels on a schedule. It all runs from a local web UI on one consumer GPU.

The pipeline

  1. Download and transcribe: yt-dlp, then WhisperX for word-level timings.

  2. Find the cuts: clip edges snap to real scene changes and pauses, never mid-sentence.

  3. Pick the moments: a language model (Claude, GPT, Gemini or a local Ollama model) reads the transcript and proposes clips, titles and hooks.

  4. Frame the shot: active-speaker detection and a smooth camera path keep whoever is talking in frame.

  5. Render and publish: a Python worker with ffmpeg and OpenCV renders the clip and burns the captions; uploads go out on a staggered schedule.

The Shorts Studio pipeline in eleven stages and three lanes. Find: download with yt-dlp, transcribe with WhisperX, snap boundaries to cuts and pauses, and analyze with a language model that proposes clips, titles and hooks. Frame: classify with rules from measured signals, bind voices to faces, compute a smooth camera path, and route to a concrete layout. Ship: render with ffmpeg and OpenCV with captions burned in, pick and grade a thumbnail, and upload to several YouTube channels on a staggered schedule. Progress streams to the browser over SSE.
Eleven stages: a Node control plane runs the job, a Python worker does the GPU work.
Six 9:16 layouts the router can choose: fullscreen-follow for one speaker, split-screen for two people talking over each other, camera-switch to cut to whoever holds the floor, group-crop for a panel of three or more, gameplay-facecam-stack with gameplay on top and the player below, and blurred-fill for b-roll.
The same source, framed differently depending on who is on screen and who is talking.

Why it's built this way

6 GB of VRAM is a hard limit. Every stage runs as its own process and only one model is on the GPU at a time.

A timeline of the GPU, the CPU and the language model. WhisperX, speaker detection and NVENC encoding take turns on the GPU while face detection and the ffmpeg filter graph run on the CPU, and the clip plan and taste calls go to the model. Notes: each stage is its own process, faces stay on the CPU, and every stage logs wall time and peak VRAM.
Only one model is ever on the GPU at a time.

Rules own the facts, the model owns the taste. Which layouts are possible comes from measured signals, never from a model's guess. The model only chooses among options already proven possible, and a malformed reply always falls back to a valid render.

Rules own the facts, the model owns the taste. 1, measure signals: faces on screen, who is speaking, overlap between speakers, subject motion, facecam or action. 2, classify: talking-head, multi-speaker, screen-rec or b-roll. 3, route to allowed layouts, for example real crosstalk to split-screen, turn-taking to camera-switch, three or more faces to group-crop. 4, the model picks among the allowed layouts, effects and pacing. A broken model reply falls back to a valid render.
The model chooses only among layouts the rules have already proven possible.

Stages only talk through files. Each stage writes a typed, versioned JSON artifact, so a crash or a restart resumes exactly where it stopped instead of re-running (and re-paying for) finished work.

The job folder storage/jobId with ingest.json, transcript.json, scenes.json, trends.json, clips.json, analysis and asd done, composition crashed, and render and out waiting. Notes: every stage writes typed, versioned JSON, files are written atomically with a temp file and a rename, and a restart skips finished stages.
A crash costs one stage, not the whole job.

// see it for yourself

Shorts Studio

View on GitHub
Next projectNamesteadFree identity subdomains on community-funded domains. Sign in, claim yourname.devportfolio.com and go live in seconds.Read the case study →