Vidu S2 includes separate Avatar and Editing models for continuous generation and incoming video streams. ShengShu Technology released Vidu S2 as two models: S2-Avatar for continuous interactive characters and S2-Editing for changing an incoming video stream. The launch shifts emphasis from waiting for a completed clip toward generating or editing while an interaction continues.

Vidu S2: verified facts

Verified event facts
Released September 15, 2026 ShengShu Technology
Models S2-Avatar and S2-Editing Vidu; arXiv
Avatar output Real-time 720p ShengShu; BJNews
Editing scope Style, clothing, character and background references Vidu platform

Vidu S2 implementation flowFour stages show scope, controls, pilot and measured rollout.Vidu S2: evidence flowScopeControlsPilotMeasure

What the update changes

ShengShu Technology released Vidu S2 as two models: S2-Avatar for continuous interactive characters and S2-Editing for changing an incoming video stream. The launch shifts emphasis from waiting for a completed clip toward generating or editing while an interaction continues.

S2-Avatar raises stated real-time output from 540p in Vidu S1 to 720p. Users can introduce a new reference image while a character is speaking, which allows a scene, object or appearance to change without restarting the whole session. The model also accepts motion instructions such as dancing.

S2-Editing applies references for style, clothing, character or background to a live input stream. That could support virtual production, live commerce and interactive entertainment, but buyers need to test identity stability and failure recovery when lighting, motion or references change abruptly.

The technical report also explores spatial video for virtual-reality use. This is an experimental direction, not evidence that the system can generate an unrestricted navigable world. Field of view, continuity, headset comfort and control response all matter beyond the visual quality of selected demonstrations.

Published frames-per-second figures should not be confused with interaction latency. A stream can render many frames while taking too long to respond to a new voice command or reference. Evaluations should record end-to-end delay from user input to a visible, stable change under realistic network and hardware conditions.

Creative teams also need provenance controls. Dynamic references may contain people, brands or copyrighted material, while a live editor can make convincing changes faster than a conventional approval workflow. Systems should preserve source references, operator actions, model versions and exported-watermark choices.

A practical trial should test long sessions, interruptions, changing speakers and adversarial references rather than only polished prompts. Measure dropped frames, visual drift, response latency, moderation behavior and the time a human spends correcting output. Public demos establish capability, not production reliability.

Vidu S2 makes generated video more stateful and interactive. Its real value will depend on whether creators can control that state predictably, disclose synthetic output and recover when a live transformation fails. Real-time generation is useful only when the control loop is trustworthy.

Related Lapaas Voice coverage

Read our coverage of Cohere encrypted inference and NVIDIA CUDA-Q logical quantum codesign for adjacent context.

Frequently asked questions

What is Vidu S2?

Vidu S2 includes separate Avatar and Editing models for continuous generation and incoming video streams.

What changed?

The Avatar model raises real-time output from 540p to 720p and accepts new references during interaction.

What should users verify?

The company is also experimenting with spatial video, but frame rate does not by itself prove low interaction latency.

Sources

Get the day’s top stories in your inbox

One concise email. No spam, unsubscribe anytime.