Skip to content
SUSHIL
← All writing

· 9 min read

Article#Content creation#AI#Motion graphics

I can't edit video. I make 3D reels anyway

Summary

I've never opened Premiere or After Effects. My first two videos came out of a folder on my laptop: Claude Code writing the scripts and scenes as code, a voice model on my own GPU, and Buffer to post them. Here's every piece.

A terminal running Claude Code with a prompt to make a 3D reel, next to three phone screens showing frames from my intro video and the Shopify explainer.

I've never opened Premiere or After Effects. I don't know what a keyframe panel looks like. Until a few days ago, "make a 3D motion graphics video" sounded like a skill that takes years.

My first two videos are out anyway: a 27-second personal intro, and a two-and-a-half-minute 3D explainer on how Shopify stops selling the last hoodie twice. Neither was edited on a timeline. They came out of a folder on my laptop, written as code, and that's the part I want to explain, because I think it changes who gets to make this kind of content.

Pipeline diagram in ten steps: 1 idea and research with Claude Code, 2 script in script.md, 3 voice with Kokoro TTS, 4 word timings with faster-whisper, 5 storyboard, 6 3D scenes with HyperFrames, 7 music and sound effects with MusicGen, 8 render and master with FFmpeg, 9 checks with hyperframes check, 10 publish with the Buffer MCP.
Every video goes through the same ten steps. The rest of this post walks through them.

The honest starting point

What I don't have: any editing or animation experience, a microphone setup, or a designer.

What I do have: I write code every day, a laptop with an RTX 4050 (6 GB of video memory), and Claude Code, an AI coding agent that runs in my terminal. My bet was simple: if a video can be described in code, I can make it, the same way I build any other project.

It turns out almost all of it can be. The scenes are web pages. The voice is a model. The timing is data. The sound mix is numbers.

The studio is a folder

Everything lives in one folder, video-editing-helper. There's no app to learn, just files that Claude Code and I both read.

The studio's folder tree: CLAUDE.md, a scripts folder with new-video.sh, tts_lines.py and transcribe.py, and a videos folder with _shared, intro-sushil and shopify-oversell. Each video has notes.md, script.md, audio, captions, hyperframes and exports. Side cards explain CLAUDE.md (the studio's rules), notes.md (each video's memory) and voices.md (every voice measured).
One folder per video, and one CLAUDE.md that every session reads first.
  • CLAUDE.md holds the studio's rules: which GPU, which voice by default, reels at 1080×1920 and 30 fps, final files as H.264 MP4 in exports/.

  • scripts/new-video.sh scaffolds a new project with the same folders every time.

  • notes.md in each video is its memory: what's done, what I approved, what broke and how it was fixed. The next session picks up exactly where the last one stopped.

Apart from Claude itself, everything runs locally on my laptop: the voice, the transcription, the music and the rendering.

Step 1: a script with stage directions

Each video starts as script.md. Plain lines are spoken. Lines that start with > are visual direction and never read out loud:

## 6. THE TRICK: TOKENS + SKIP LOCKED
> VISUAL: the single "6 left" row shatters into six 3D tokens in a glass bowl.
> ON SCREEN: "Before: 1 row that says 6  ->  After: 6 rows, one per hoodie"

And here's that trick. Instead of one row that says six,
make six rows, one for each hoodie. Like six tokens in a bowl.

The first version of the Shopify script lost anyone who wasn't already a developer. The rewrite follows one rule: show an everyday object first, then put its real name on screen. A notebook is a database, a bank transfer is a transaction, an "occupied" sign is a lock, and tokens in a bowl are rows. Claude drafts and rewrites; I approve the hook and the final wording.

Step 2: a voice without a microphone

The narration comes from Kokoro, a small open text-to-speech model that runs on my GPU in a few seconds. No studio, no retakes.

Picking the voice was the first real decision. Every voice was measured for pace (words per minute) and pitch, and the numbers went into voices.md. The first one I heard sounded too light for me, so the final voice is a blend of two: Kokoro can average voices, and am_michael plus am_onyx came out deeper and steadier than either alone.

The voice is generated one line at a time, with pauses I design instead of pauses the model guesses. Each line's start and end time is saved, so scene cuts can land exactly on a sentence:

# scripts/tts_lines.py (core loop)
for i, line in enumerate(lines):
    audio = np.concatenate([c for _, _, c in pipe(line, voice="am_michael,am_onyx", speed=a.speed)])
    nz = np.where(np.abs(audio) > 0.01)[0]            # trim the model's own silence
    audio = audio[max(0, nz[0] - 240): nz[-1] + 1200]
    meta.append({"text": line, "start": round(t, 3), "end": round(t + len(audio) / SR, 3)})
    t += len(audio) / SR + a.gap                       # then a pause I chose

Kokoro said "Riddies". the script says "Red iss".

Small fixes like that respelling live in the script, and the captions still show the right word.

Step 3: every word gets a timestamp

Next, faster-whisper (a speech-recognition model, also on the GPU) listens to the finished voice and returns every word with its start and end time. That does two jobs:

  • A pronunciation check. If the transcript doesn't match the script, the voice said something wrong. One voice turned "picked" into "pick"; this is how it got caught.

  • The clock for everything else. Captions highlight the word being spoken, and animations start on the word they illustrate.

The caption text comes from the script, not the transcript, lined up against Whisper's timings with Python's difflib. So the screen says "Redis" and "$5.1 million" even though the voice said "Red iss" and "five point one million dollars".

Step 4: a storyboard before anything moves

Before a single animation, Claude draws a storyboard: one static sketch per scene, with the line it belongs to and how it hands over to the next. I approve the sheet, and only then does anything get animated.

The first act of the Shopify storyboard: four 9:16 sketches. 01 Hook, Stop selling it twice with a Redis cube and a MySQL cylinder. 02 Two buyers pressing Pay at the same second. 03 The hold behind the counter with Paid and Failed outcomes. 04 Two notebooks becoming Redis and MySQL. Each sketch has its caption and timing.
Act one of the Shopify storyboard: 16 frames in total, each approved before it was animated.

A design.md sits next to it, like a mini design system for the video. It sets the colours (each idea keeps one colour for the whole film: Redis is always coral, MySQL always blue), the fonts from this site, and the safe zones. Instagram puts buttons on the bottom 360 pixels and the right 130, so no text goes there.

Step 5: the 3D scenes are web pages

This is the part that would have taken me years to learn by hand. The scenes are built with HyperFrames, HeyGen's tool for making video from HTML: each scene is a web page with a GSAP animation timeline, and HyperFrames renders it frame by frame in headless Chrome, straight to MP4 on the GPU.

The "3D" is real CSS 3D: a cube is six faces in a preserve-3d container, a database is a cylinder of stacked discs, and the camera is a slow perspective drift so no frame is ever frozen. I didn't have to learn a 3D package, because it's the same CSS I write for websites.

Scenes don't contain hard-coded times. They contain the words they react to:

// tools/scenes/f12-skip-locked.html: the term pill appears when the voice says it
const ts = {{t:called skip locked}};
tl.fromTo($("pill"), { scaleX: 0, opacity: 0 }, { scaleX: 1, opacity: 1, duration: 0.35, ease: "power3.out" }, ts);
tl.fromTo($("sql"), { opacity: 0, y: 40 }, { opacity: 1, y: 0, duration: 0.4, ease: "power3.out" }, ts + 0.3);
Diagram: the voice is the clock. 1, the script line 'Then use a MySQL feature called skip locked'. 2, the voice and its word timestamps: called at 110.32 seconds, skip at 110.68, locked at 111.02. 3, the scene code with the placeholder {{t:called skip locked}}, which build_scenes.py replaces with 1.2 seconds from the start of the scene.
A small build script swaps every placeholder for the real second. Change a line, re-voice it, and every scene re-times itself.

The intro used the same approach with a different job: it's about me, so it uses my photos and this site's look. The same compositions render in both 9:16 for reels and 16:9 for YouTube, because each scene adapts its layout to the frame width. I made three alternative scripts too, and all four versions rendered in both sizes.

Four frames from the 16:9 version of the intro video: the name scene with the portrait and the dev badge, a sketch of an app labelled idea.sketch, 'Picked tech early' with a photo at the desk, and the manifesto lines I build front, I build back.
The 16:9 cut of the intro, rendered from the same scenes as the reel.
Two scenes from the Shopify explainer: six green hoodie tokens, and a glass bowl where a locked token is skipped and the next buyer takes token 2.
From the Shopify explainer: one row per hoodie, and SKIP LOCKED shown as a hand skipping a locked token.

Step 6: sound, by the numbers

Sound is where I expected to fail, because I can't hear a mix the way an editor can. So it's done with measurements instead of ears:

  • Music is generated locally with MusicGen, a music model, and looped under the voice at a steady low level.

  • Sound effects: the Shopify video has 68 cues (pops, whooshes, impacts), each set from its measured loudness rather than by guessing.

  • Loudness: the voice is normalised to -16 LUFS, and the final master to -14 LUFS with peaks under -1 dB, which is about where Instagram and YouTube expect it.

The delivered Shopify reel is H.264 at 1080×1920 and 30 fps, with AAC audio at 256 kbps and the index at the front so it starts playing straight away.

Step 7: checks before anyone sees it

A video can't be unit-tested, but it can be checked:

  • hyperframes check lints every scene and measures text contrast. The Shopify reel passed with 0 errors and all 46 text checks.

  • Contact sheets: frames from every scene on one image, so a broken scene shows up at a glance.

  • Re-transcribing the final audio to make sure nothing got mangled in the mix.

  • Loudness over time, second by second, to find spikes like the one in the lessons below.

Step 8: one reel, three platforms

Each platform wants different words, so each gets its own copy, written into post-copy.md. Instagram gets the hook in the first line and five hashtags at most. X gets a short post and a thread. LinkedIn gets the longer version: the problem, the fix, and what I take from it. The cover image is also a web page, rendered to a 1080×1920 PNG.

One reel goes to Buffer, connected to Claude as an MCP server, which queues it to three channels: Instagram @sushilk.dev with the hook first and five hashtags at most, X @sachu0dev with a short post and a 4-part thread, and LinkedIn in/sachu0dev with a longer post.
Written once per platform, scheduled from the same conversation.

Posting goes through Buffer, connected to Claude as an MCP server (a plug-in that lets the agent use Buffer directly). The same conversation that made the video can queue it to Instagram, X and LinkedIn, with the right caption on each.

What I decide, what the tools do

This doesn't mean the AI makes the videos and I watch. It means I do the part that needs taste, and the tools do the part that needed years of practice.

Two columns. Me: the topic and what the viewer should walk away with, approving the script, the hook and the storyboard, picking the voice by ear, watching every render and saying what feels off, and what gets posted where and when. Claude and the tools: research, script drafts and rewrites, voice, word timings and captions, every scene as code, rendering, loudness and checks, and captions for each platform queued in Buffer.
I direct; the tools build.

Mistakes I only had to make once

Every one of these is written into a notes.md, so the next video starts already knowing them:

Four lessons. The v1 script lost beginners, so: analogy first, term second. The voice sounded weak, so: measure every voice, then blend two. The music drowned the voice, so: set levels by numbers, voice at -16 LUFS and master at -14. A 'sound bug' at 14 seconds was a 3.5-second glitch effect, so: trim every effect to the moment it marks.
Lessons from the first two videos.
  • An automatic "duck the music under the voice" effect silenced the music completely in my renders, so it's a plain volume level now.

  • A preview tool rewrote my project template and broke the music settings, so the template now lives in its own file that the tool never touches.

  • A 3.5-second glitch sound, used as a quick accent, played for its whole length and sounded like a bug. Effects get trimmed to the moment they mark.

What this means to me

Motion design, voice work and sound mixing are real crafts, and the people who are great at them are still far ahead of what I do. But the gap between "I have an idea for a video" and "it's posted" has become something a developer can close in days, not years. The skills that got me here are the ones I already use at work: break the problem into steps, write each one down, automate it, and check the output.

It's early: two videos, one studio that's days old, and a lot to learn. If you want to follow how it goes, or tell me which system I should explain next:

next video: you pick the system.

other articles