用 AI 工具组合,1 小时做一条短视频
A decent short needn't mean an all-nighter. Chain AI tools into a pipeline and go from blank to finished in an hour.
The stack
- ChatGPT / Claude 写文案
- ElevenLabs 生成真人语音
- Runway / Midjourney 出素材
- 剪映/CapCut 自动字幕合成
Time budget
- 定选题、写 30 秒脚本
- ElevenLabs 出配音
- Runway 生成画面片段
- 剪映合成+字幕+导出
Prompt templates
Script: a 30s hook-first tip video on '3 AI time-savers'. Voice: calm male, medium pace.
Gotchas
- 用平台曲库或 AI 生成
- 数字人/Sync 才需要对嘴
- 前 3 秒必须抓人
Wrap-up
Once the pipeline is set, several videos a day is easy — standardize each step.
Step-by-step detail
0–10 min: script
Draft a 30-second voiceover script with ChatGPT/Claude: audience, topic, length, tone and CTA. The first 3 seconds must hook.
10–25 min: voiceover
Paste the script into ElevenLabs, pick a matching tone, tune pace, export. This is usually the slowest step.
25–45 min: footage
Generate 3–5 five-second clips in Runway, or images in Midjourney with added camera motion. Specify shot types or the visuals go astray.
45–60 min: edit
Cut to the script in CapCut, lay in audio and footage, auto-generate captions, and duck the background music under the voice.
Tool choices
- Talking-head:数字人(Synthesia)可直接替代真人出镜,省掉拍摄。
- B-roll heavy:Runway 更合适,Midjourney 出图再运镜也可。
- On a budget:全程用免费档也能出片,只是要多等生成。
Boost retention
- 3-second hook:开头别自我介绍,先抛结论或痛点。
- Change frame every ~5s:视觉上避免单调。
- Large captions:多数人静音刷,硬字幕是刚需。
- End with a CTA:关注/收藏/评论任选其一。
Real case: a knowledge short in 57 minutes
For a 60-second "3 AI time-savers" short: script 8 min, ElevenLabs voice 12 min (incl. two pace tweaks), Runway clips 22 min, CapCut edit + captions 15 min — 57 minutes total.
- Slowest step:不是生成画面,而是配音的语速与停顿调整——建议先定好语速再批量生成。
- Efficiency key:把脚本拆成 3 段分别配音再拼接,比整段生成更容易对齐画面。
- Day two:模板固定后,同系列第二条只需换脚本与画面,可压到 40 分钟内。