字节 Seed 官方:一镜成片,灵活参考 —— Seedance 2.5 发布全文(中文翻译)
字节跳动 Seed 团队 Seedance 2.5 官方发布博客的完整中文翻译,含 10 段官方演示视频与 8 段完整官方提示词原文。
关于这篇翻译
原文:One-take Creation, Flexible Referencing: Introducing Seedance 2.5 · 字节跳动 Seed 团队 · 2026-07-31
这是非官方中文翻译,版权归字节跳动 Seed 团队所有。翻译目的是方便中文读者阅读,文中所有视频均直接引用官方原始地址、未做转存。如原作者认为不妥,请告知,我会立即撤下。
为什么值得完整读一遍:这篇发布博客里嵌了 8 段完整的官方提示词原文。市面上关于 Seedance 的提示词教程绝大多数是二手转述,而这些是模型团队自己写的、并且配着对应成片。想搞清楚多参考素材到底怎么绑定、时间戳到底怎么写,这是最短的路径。
⚠️ 视频说明:官方原片是 4K 高码率(开篇那支 4 分 22 秒的片子原文件 1.37 GB),页面已设为点击才加载,不会在打开时消耗流量。点开后需要一点缓冲时间,这是原片规格决定的。
今天,我们正式发布新一代视频创作模型 Seedance 2.5。
自 Seedance 2.0 发布以来,我们注意到用户对视频创作模型的期待发生了变化:从「生成一个片段」,转向「完成一件创作作品」。Seedance 2.5 建立在 Seedance 2.0 统一的多模态音视频联合生成架构之上,围绕基础生成与参考生成两条主线,在长叙事、多模态参考和编辑三个方向上取得重要突破。它立足真实使用场景,打开了更大的创作想象空间与控制力,并进一步释放生产力。
核心亮点包括:
单次最长 30 秒,支持多轮延长:Seedance 2.5 能一次生成高质量的 30 秒音视频片段,并支持多轮延长。它同时改善了镜头转场与场景切换,让长视频的连续性更强,在画面、音频和运动质量上都有明显提升,最终呈现出比常见 AI 生成视频更自然、更精致的观感。用户因此可以产出音画语言一致的数分钟内容,一镜之内讲完一个完整故事。
多模态参考全面升级:用户现在可以在单次生成中输入最多 30 张图片、10 段视频和 10 段音频作为参考素材。模型同时强化了一系列参考能力,包括白模渲染(clay render)参考、运动参考和创意参考,使其能更好地理解创作者意图,实现跨多主体、多场景、多镜头切换的复杂想法。
更精确稳定的编辑能力:Seedance 2.5 提供时间戳级别的控制,可对音视频内容做定向编辑,显著提升效率与可控性。模型还增强了绿幕、镜头视角、参考编辑等高级编辑功能,以满足影视和广告这类专业复杂领域的严苛要求。
在长叙事、多模态参考和编辑三方面的进展之下,Seedance 2.5 的意义不止于「单次能生成更长的视频」。模型更懂创作意图,把从想法到成片这段路,交付得更可控。下面这支短片,是完全由 Seedance 2.5 端到端制作的创意影片。
Seedance 2.5 今日起在即梦 AI、豆包专业版等平台陆续开放,API 也将很快通过 BytePlus ModelArk 提供。欢迎试用并反馈。
项目主页:seed.bytedance.com/seedance2_5
使用入口:即梦网页版 → 视频生成 → 选择 Seedance 2.5;豆包专业版 → 视频生成 → 选择 Seedance 2.5
30 秒长叙事与多轮延长:一次讲完一个完整故事
Seedance 2.5 把单次生成从 15 秒延长到 30 秒,并进一步强化了长视频中的叙事能力。在 30 秒内,模型能够组织多个逻辑上相互关联的镜头,让故事经由铺垫、发展、转折和收束展开,而不只是把某一个瞬间拉长。
举个例子:在一支歌手登台的一镜到底片段里,模型呈现的是完整故事——歌手在化妆间与工作人员互动,穿过后台走廊,与舞者会合,最后一同走上舞台开始表演——而不只是走上台的那一个瞬间。
文生视频(T2V)提示词原文:
One-take handheld gimbal tracking shot. The camera slowly pushes in through a gap in a heavy red curtain and enters a warm-toned backstage dressing room. A young female singer, with her back to the camera, is adjusting her earpiece as a staff member reminds her it's time to go on. She turns toward the camera and starts singing citypop. The camera pulls back and tracks her as she passes through the curtain into a dim backstage corridor, interacting naturally with her dancers along the way; one staff member hands her a microphone. She and the dancers then step onto the stage, and the camera arcs around to the back, gradually revealing the red-and-black stage design, LED screens, spotlights, haze, and reflective floor. The camera finally pulls out to a wide shot of the arena, showing the packed audience, light boards, glow sticks, and cheering crowd, capturing the youthful, free-spirited climax of the concert.
中文: 一镜到底手持稳定器跟踪镜头。镜头缓缓从厚重红色幕布的缝隙推入,进入暖色调的后台化妆间。一位年轻女歌手背对镜头,正在调整耳返,工作人员提醒她该上场了。她转向镜头,开始唱一首 citypop。镜头后拉并跟随她穿过幕布,进入昏暗的后台走廊,沿途与舞者自然互动,一名工作人员递给她一支麦克风。她与舞者一同走上舞台,镜头绕到背后,逐渐揭开红黑配色的舞台设计、LED 屏、追光、烟雾和反光地面。镜头最终拉出成为场馆的大全景,呈现满场观众、灯牌、荧光棒和欢呼的人群,捕捉这场演唱会青春而自由的高潮。
得益于模型的多轮延长能力,用户可以在已有视频输出的基础上平滑地续接后续镜头。在延长过程中,主要角色、环境和叙事节奏都保持一致。这让用户可以一次性输出数分钟长度的视频,减少了拆分片段、反复拼接、修补转场所需的工作量。
参考生视频(R2V)提示词原文:
Extend the video. Continue from the visuals and subjects in @Video 1 and generate another 30-second clip, keeping the character subjects, scene, visual style, and sound effects consistent. The little boy runs along the train carriage holding a soccer ball. When the subway stops, the side door opens and he immediately dashes out, with the male lead chasing after him. The two run across the platform and out onto the street, startling passersby and vehicles along the way. The male lead finally catches up and grabs him. The boy looks up, aggrieved. The male lead's anger slowly fades; he pats the boy's head and shows a helpless smile.
中文: 延长该视频。承接 @视频1 中的画面与主体,再生成一段 30 秒片段,保持角色主体、场景、视觉风格和音效一致。小男孩抱着足球沿车厢奔跑。地铁停靠时侧门打开,他立刻冲出去,男主角在后面追赶。两人跑过站台冲上街道,一路惊动路人和车辆。男主角终于追上并抓住他。男孩抬头,一脸委屈。男主角的怒气慢慢消退,他拍了拍男孩的头,露出无奈的笑。
在视觉呈现上,模型实现了更平滑的镜头运动过渡。主体在多次切换之间保持稳定,音画保持同步,最终得到高度连贯的长视频。例如在一段京剧场景中,镜头跟随主角舞动的水袖完成一次优雅的环绕运镜,主体与背景始终保持一致,水袖的甩动在空中形成自然的弧线,紧贴真实世界的物理规律。
R2V 提示词原文(注意这段的时间戳写法):
16:9 widescreen, cinematic texture, single continuous take, smooth camera movement, no cuts. Scene reference: @Image 4. 0–5s: Open with a close-up of the Overlord from @Image 2. The camera slowly circles his upper body and transitions into a medium shot. The Overlord spins and turns, his body and back flags sweeping quickly past the lens to form a natural occlusion, and the camera follows through to Consort Yu's side in @Image 1. 6–10s: The camera steadily circles Consort Yu in a medium shot from @Image 1, following her water sleeves through the arc. She raises her arm, flicks her wrist, unfurls the sleeves, and half-turns. She then draws the sleeves back, holds the pose, and looks sideways toward the Overlord. 11–20s: The male warrior from @Image 3 enters with an aerial flip. The Overlord takes center stage while the warrior advances and retreats on the opposite side in a combat exchange. Consort Yu stands slightly behind and to the side of the Overlord, weaving in water-sleeve movements to set softness against strength. The camera slowly pulls back from a medium-close shot of the warrior to a full stage view. At the end, all three face the audience and strike a synchronized Peking opera finale pose.
中文: 16:9 宽银幕,电影质感,单一连续镜头,运镜平滑,无剪切。场景参考:@图片4。0–5秒:以 @图片2 中的霸王特写开场。镜头缓慢环绕他的上半身并过渡为中景。霸王旋身转向,身形与背后靠旗快速扫过镜头形成自然遮挡,镜头顺势跟到 @图片1 中虞姬一侧。6–10秒:镜头以中景平稳环绕虞姬(@图片1),跟随她的水袖划过弧线。她抬臂、抖腕、甩开水袖、半转身,随后收袖定势,侧目望向霸王。11–20秒:@图片3 中的男性武生以空翻入场。霸王居于舞台中央,武生在对侧进退交手。虞姬略靠后侧站位,以水袖动作织入柔劲对比。镜头从武生的中近景缓缓后拉至全舞台视角。结尾三人面向观众,同步亮相定势。
此外,针对 AI 生成视频常见的「过于人工」的观感,Seedance 2.5 系统性优化了物体质感、皮肤与眼部特征、光照和色彩饱和度等细节,同时减少了字幕和背景音乐上的失控现象,让成片更接近实拍影像的电影质感。
多模态参考全面升级:复杂创作有了更强的控制力
Seedance 2.5 进一步强化了多模态参考生成能力,允许用户在单次生成中输入最多 30 张图片、10 段视频和 10 段音频作为参考素材。更大的素材量和更丰富的素材种类,能更好地捕捉用户意图,产出主体更多、场景更丰富、运镜更灵活的复杂视频。
模型会全面理解所有素材中的视觉构图、场景、风格、角色和道具等元素,并按指令应用到视频生成过程中。即便在多角色镜头或群像叙事这类复杂场景下,它也能保留多个角色的外貌与声音,同时保持各主体特征稳定。
R2V 提示词原文(这段一次绑定了 18 张参考图,是全文最值得研究的一段):
A 30-second concert sequence in 16:9 landscape, with cinematic realism, authentic concert hall lighting and shadows, warm golden stage lighting, and the atmosphere of a formal classical concert. Use @Image 1 for the venue. Reference @Image 2 for the pianist. Reference @Image 3 for the cello. Reference @Image 4 for the violin. The lead vocalist must strictly follow @Image 5. Reference @Images 6 to 10 for the rest of the orchestra. Reference @Images 11 to 14 for the choir. Reference @Images 15 to 18 for the audience seating. The lead vocalist walks from center stage toward the front edge. The pianist is positioned by the piano. The orchestra is arranged on both sides and toward the rear. The choir stands at the back of the stage. Open with a high-angle wide shot of the full concert hall. The pianist strikes the keys, and the lead vocalist steps into the spotlight and begins singing. The camera naturally moves across the violin, cello, and orchestra as they perform together, with the violin feeling bright and the cello warm. In the latter part, the choir joins in. The lead vocalist briefly makes eye contact with front-row audience members, who respond with a smile and a slight nod. In the closing shot, the camera pulls back. The singing ends, and the audience joins in the applause.
中文: 一段 30 秒的音乐会序列,16:9 横画幅,电影级写实,真实的音乐厅光影,温暖的金色舞台照明,正式古典音乐会的氛围。场地使用 @图片1。钢琴家参考 @图片2。大提琴参考 @图片3。小提琴参考 @图片4。主唱必须严格遵循 @图片5。乐团其余部分参考 @图片6 至 10。合唱团参考 @图片11 至 14。观众席参考 @图片15 至 18。主唱从舞台中央走向前沿。钢琴家位于钢琴旁。乐团分布在两侧和后方。合唱团站在舞台后部。以音乐厅全景的高角度大全景开场。钢琴家落键,主唱步入追光开始演唱。镜头自然掠过小提琴、大提琴和共同演奏的乐团,小提琴明亮、大提琴温暖。后半段合唱团加入。主唱与前排观众短暂对视,观众报以微笑和轻轻点头。收尾镜头后拉。歌声结束,观众一同鼓掌。
Seedance 2.5 还增强了白模渲染(clay render)参考、运动参考和创意参考等特定参考能力,让用户对画面中的主体、动作和运镜有更精细的控制。以白模渲染参考为例:用户可以用无材质的 3D 模型搭建场景的空间结构、角色姿态、运动路径和机位角度,模型再依据这套结构生成视频,确保复杂镜头的构图与调度贴近创作者预期。
此外,Seedance 2.5 改进了光照控制。借助白模中的空间信息,它能生成符合物理规律的真实光照效果——包括光源方向、色温、强度和阴影投射——让最终成片的光影更自然。
R2V 提示词原文(白模参考的标准写法,注意「继承什么」被拆得很细):
Refer to @Clay Render 1 for camera movement, pacing, shot-size transitions, subject trajectory, and blocking. Refer to @Image 2 for character design, scene, materials, lighting, color, and fairy-tale atmosphere, and render the white model as a dreamy, warm, 3D animated short with a childlike fantasy feel. The story unfolds as follows: flight through a fantasy sky → mythical beasts flying alongside through a sea of clouds → a dive into the ocean → weaving through the deep with manta rays → passing through a mirrored rift in spacetime → picking stars from the cosmos → transforming back into the bedroom → father tucking in the blanket → the picture book closes and holds on the final frame.
中文: 参考 @白模1 的运镜、节奏、景别过渡、主体运动轨迹和场面调度。参考 @图片2 的角色设计、场景、材质、光照、色彩和童话氛围,把白模渲染成一支梦幻、温暖、带童真幻想感的 3D 动画短片。故事如下展开:飞越幻想天空 → 神兽在云海中并肩飞行 → 俯冲入海 → 与蝠鲼一同穿行深海 → 穿过镜面时空裂隙 → 从宇宙中摘下星星 → 变回卧室 → 父亲掖好被角 → 绘本合上并定格在最后一帧。
更精确可靠的编辑:创作效率的提升
在视频创作中,用户通常既需要在生成过程中控制节奏,也需要在生成之后打磨细节。某个动作发生在第几秒、镜头切换的精确时机、某一段里角色的动作是否需要调整——这些都会深刻影响最终结果。更精确可靠的编辑能力,让创作者能准确地把想法落地,提升效率并降低反复生成的成本。
Seedance 2.5 支持通过时间戳进行精确的内容编辑。在生成阶段,用户可以用提示词控制特定时间段内的叙事、镜头视角、运动和整体节奏,让输出更贴近创作意图。生成之后,用户还可以对特定片段内的角色、动作或情节做定向修改,同时保持编辑前后的连续性和真实感。
Seedance 2.5 也提升了绿幕编辑、镜头视角编辑和参考编辑等多项编辑功能,以满足影视、广告等专业领域的严苛要求。以绿幕编辑为例:模型可以在保持主体不变的前提下替换背景、讲述完全不同的故事,并且擅长呈现主体对新环境物理规则的响应——包括衣物飘动方向、头发状态、步态节奏和光照互动,确保主体与场景和谐融合。
R2V 提示词原文:
Using @Video 1, render the green-screen background, obstacles, wardrobe, and supporting characters. 0–4s: outdoor training, replace the obstacles with rocks, bricks, tires, and wooden crates. 4–10s: locker room, friends offering encouragement. 10–15s: international match, replace the training poles with original defenders and a goalkeeper, and the protagonist scores. Overall photorealistic, cinematic quality.
中文: 使用 @视频1,渲染绿幕背景、障碍物、服装和配角。0–4秒:户外训练,把障碍物替换成岩石、砖块、轮胎和木箱。4–10秒:更衣室,朋友们给予鼓励。10–15秒:国际比赛,把训练杆替换成原创的防守球员和守门员,主角进球。整体照片级写实,电影质感。
R2V 提示词原文(只改运镜、其余全部锁死,这段值得反复看):
Edit @Video 1. Keep the characters, actions, and visual style unchanged. Adjust only the camera movement. A 15-second segmented camera plan: 0–4s, a micro-FPV move skims tightly past the pan, then follows the popping toast and whip-pans to the coffee; 4–7s, push in and track laterally along the rim of the pan, following the fried egg as it flips up and lands back in place; 7–11s, rapidly rise to a top-down view, then descend at a steady pace, sweeping across the plate and keys; 11–15s, use a handheld close-up to follow the hands with a fast lateral whip, then push in on the breakfast and pull back to a medium two-shot. Keep the entire sequence smooth, continuous, and stable.
中文: 编辑 @视频1。保持人物、动作以及画风不变,仅调整运镜。15 秒分段运镜设计:0–4秒,微型 FPV 贴锅穿行,跟随弹起的吐司后横甩至咖啡;4–7秒,推近并沿锅边横移,跟随煎蛋翻起又落回;7–11秒,急速升至顶视角,再匀速下降,扫过盘子和钥匙;11–15秒,用手持特写跟随双手快速横甩,然后推近早餐,再拉回为中景双人镜头。整段保持流畅、连续、稳定。
走向更广泛的行业场景
随着模型能力的演进,Seedance 2.5 正在深入教育、制造等更广泛的行业场景。
在教育领域,模型已经开始进入真实的教学场景。例如,Seedance 2.5 可以把一节课背后的历史背景、人物和故事线转化为更生动、更具沉浸感的画面。它也帮助教师更高效地制作教学视频,把科学原理、历史事件、实验流程这些抽象内容变成动态演示。这不仅降低了教学材料的制作门槛,也让内容定制变得高度灵活。
上图为豆包爱学 App「豆包课堂」场景示例
R2V 提示词原文:
Expressive Eastern painterly style. A street scene in Lin'an during the Southern Song dynasty. Several children run and shout through the bustling street, chanting, "I turn around, and there he is, where the lantern lights grow dim." The camera follows the children as they run, sweeping past the lively street. The camera then tilts up to reveal Xin Qiji from @Image 1. Xin Qiji turns his head, and in the distance stands a man among the fading lantern lights. The shot stays continuous throughout.
中文: 写意东方绘画风格。南宋临安的街景。几个孩童在热闹的街上奔跑呼喊,念着「蓦然回首,那人却在,灯火阑珊处」。镜头跟随奔跑的孩童,掠过热闹的街道,随后上摇露出 @图片1 中的辛弃疾。辛弃疾回头,远处灯火阑珊处立着一个人影。全程保持镜头连续。
在工业制造、具身智能和自动驾驶等领域,Seedance 2.5 正在融入高度具体的生产流程。模型可以生成高质量的合成视频数据,帮助训练机器人的感知与操作能力;也被用于工业仿真、流程培训和设备演示。在自动驾驶方面,模型可以模拟长尾场景,例如极端天气和复杂交通状况,为系统测试和训练提供更多样的样本。
R2V 提示词原文:
Reference the camera work, composition, shot scale, spatial relationships, part positions, model structure, assembly order, and motion paths from @Clay Render 1. Reference the materials, lighting, color, reflections, and atmosphere from @Image 1, and turn the clay render into a high-end, photorealistic car assembly sequence.
中文: 参考 @白模1 的运镜、构图、景别、空间关系、零件位置、模型结构、装配顺序和运动路径。参考 @图片1 的材质、光照、色彩、反射和氛围,把白模渲染成一段高端的、照片级写实的汽车装配序列。
总结与展望
Seedance 2.5 标志着模型在理解和呈现真实世界方面迈出了重要一步,把视频生成从「片段级输出」提升为「完整的创作工作流」。同时我们也清楚,仍有改进空间——尤其是复杂运动的物理合理性,以及多主体互动场景的稳定性。
展望未来,Seed 团队将继续探索更连贯的叙事、提供更直观的生成与编辑体验,并进一步深化模型对真实世界物理的理解。我们希望 Seedance 系列模型能变得更生动、更可控、更懂用户意图,帮助更多人表达创意,同时持续探索并服务更广泛的行业需求。
译者补充的两个观察
一是这 8 段提示词里,凡是用到参考素材的,几乎都在做同一件事:给每个素材单独声明「从它这里继承什么」。白模那段拆得最细——运镜、节奏、景别过渡、主体轨迹、场面调度从白模来,角色设计、场景、材质、光照、色彩、氛围从图片来。这个「分素材定向声明」的写法,比笼统地丢一堆参考图有效得多。
二是时间戳的官方写法是 0–5s: 6–10s: 这种直白格式,且每一段里都同时交代镜头怎么动、主体做什么。京剧那段和早餐运镜那段是最好的范本。
原文出处:ByteDance Seed 官方博客