No Corners Cut: State-Grounded Transitions for Mid-Stream Prompt Switches in Video Generation
Organizations: University of Chinese Academy of Sciences · Kling AI · University of Science and Technology Beijing
Abstract
Streaming video generators allow users to dynamically modulate video synthesis via mid-stream prompt switching. Existing streaming methods can respond to the updated instruction while still cutting corners, prematurely realizing goals or taking heuristic shortcuts that bypass necessary intermediate state changes needed for a plausible transition. In this study, we present SEGUE, a novel framework that makes this process explicit and trains the generator to execute these transitions faithfully. At each switch, a training-free planner parses the latest frame and prompts, writes a few segue prompts with roles and durations, and then hands control back to the user's prompt. Furthermore, to address the inherent difficulty of training causal models on short-lived temporal schedules without corrupting preparatory supervision, we introduce SPANDMD, which evaluates each active prompt using the full rollout as temporal context while retaining its DMD residual only within the prompt's assigned span. On OpenTrans-360, a benchmark of 1,800 switches that scores how the old state exits and the new one begins, SEGUE ranks first on all eight transition metrics and raises the overall score over the strongest baseline from 0.866 to 0.887. It also ranks first on four of six instruction-response metrics of StreamAV-Bench, while the planner transfers to frozen autoregressive generators without retraining.
Figures & tables
| Segment Quality | Transition Quality | |||||||||||
| Method | SPF | SMC | SPP | VBS | MBS | OSE | NSE | SP | ScP | SUF | TP | Overall |
| Closed-source Models | ||||||||||||
| PixVerse R1 | 0.899 | 0.852 | 0.984 | 0.984 | 0.872 | 0.783 | 0.853 | 0.751 | 0.809 | 0.773 | 0.936 | 0.863 |
| HappyOyster | 0.901 | 0.864 | 0.996 | 0.961 | 0.861 | 0.762 | 0.854 | 0.780 | 0.840 | 0.769 | 0.941 | 0.866 |
| Open-source Models | ||||||||||||
| Self-Forcing | 0.793 | 0.799 | 0.923 | 0.993 | 0.865 | 0.628 | 0.779 | 0.726 | 0.773 | 0.647 | 0.847 | 0.798 |
| Method | VBS | MBS | OSE | NSE | SUF | TP |
|---|---|---|---|---|---|---|
| Segue (full) | 0.995 | 0.897 | 0.802 | 0.890 | 0.796 | 0.961 |
| w/o TD-Schema | 0.989 | 0.881 | 0.758 | 0.885 | 0.788 | 0.954 |
| w/o SpanDMD | 0.988 | 0.869 | 0.759 | 0.865 | 0.790 | 0.950 |
| w/o SpanDMD & planner | 0.976 | 0.824 | 0.717 | 0.861 | 0.747 | 0.931 |
| Self-Forcing | 0.993 | 0.865 | 0.628 | 0.779 | 0.647 | 0.847 |
| + planner | 0.995 | 0.877 | 0.684 | 0.829 | 0.721 | 0.880 |
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
| Baseline | Visual quality | Transition naturalness | Overall preference |
|---|---|---|---|
| Self-Forcing | 76.7% | 73.3% | 80.0% |
| LongLive | 61.7% | 58.3% | 65.0% |
| Rolling Forcing | 78.3% | 75.0% | 81.7% |
| Deep Forcing | 66.7% | 63.3% | 70.0% |
| MemFlow | 63.3% | 60.0% | 66.7% |
| Causal Forcing | 83.3% | 80.0% | 85.0% |
| Visual Quality | Instruction Adherence and Interactive Response | |||||||||
| Method | VA | VQ | VA-D | VQ-D | VIF | VID | VUF | PVC | PVUAR | PVRL |
| PixVerse R1 | 0.529 | 2.428 | 0.032 | 0.253 | 2.976 | 0.484 | 2.573 | 4.278 | 0.350 | 12.992 |
| HappyOyster | 0.529 | 2.785 | 0.032 | 0.417 | 3.565 | 0.510 | 2.975 | 3.744 | 0.518 | 14.484 |
| Self-Forcing | 0.585 | 2.753 | 0.025 | 0.323 | 3.019 | 0.373 | 2.237 | 3.912 | 0.185 | 14.164 |
| LongLive | 0.605 | 2.768 | 0.023 | 0.367 | 3.164 | 0.317 | 2.341 | 3.631 | 0.241 | 11.792 |
| Rolling-Forcing | 0.572 | 2.757 | 0.043 | 0.427 | 3.009 | 0.425 | 2.315 | 4.086 | 0.217 | 13.234 |
| Planner input | VBS | MBS | OSE | NSE | SUF | TP |
|---|---|---|---|---|---|---|
| Text only | 0.984 | 0.867 | 0.756 | 0.860 | 0.774 | 0.920 |
| Text + current frame | 0.991 | 0.890 | 0.788 | 0.892 | 0.783 | 0.965 |
| Component | Description |
|---|---|
| Visual observation | The final frame of the generated segment, providing evidence of ongoing actions, subject pose, physical contacts, held objects, visible entities, and spatial relations. |
| prompt update | The preceding and updated prompts, specifying the change in the desired event. The observed state takes precedence when it differs from the preceding prompt. |
| Transition budget | The available duration for intermediate actions before control returns to the target prompt. |
| Grounded source state | A concise description of the observed conditions relevant to planning the transition. |
| Intermediate plan | An ordered sequence of segue prompts, each specifying its role, action description, duration, and intended start and end states. |
| Schema element | Definition and applicability |
|---|---|
| Terminate | Concludes or safely interrupts an ongoing event when its continuation conflicts with the updated prompt. It resolves the incompatible activity through an observable completion or interruption, avoiding an unexplained discontinuity in the subject’s behavior. |
| Release | Removes interaction dependencies that prevent the next event, such as held objects, physical contacts, occupied hands, or incompatible subject–object relations. It is invoked only for dependencies that obstruct the target event; compatible interactions can persist. |
| Align | Establishes prerequisites of the target event, including subject pose, spatial arrangement, interaction configuration, or viewing condition. It may comprise several distinct preparatory actions when multiple prerequisites remain unsatisfied. |
| Entry | Establishes required entities absent from the current visual state through a plausible observable process. It accounts for how an entity becomes available for the target event, maintaining continuity with the visible scene. |
| Execute | Denotes the beginning of the target event specified by . It is not instantiated as an additional segue prompt; instead, it marks the handoff from transition planning back to the original target prompt. |
| System message |
|---|
| You are a transition planner for streaming video generation. Given the last generated frame, the preceding prompt, and an updated prompt, plan the intermediate actions needed to reach the conditions for the target event. |
| Ground the source state in the image. Identify ongoing activity, held objects, physical contacts, occupied hands, subject pose, spatial relations, and visible entities. Use the observed state when it differs from the preceding prompt. Do not assume any future visual outcome. |
| Select only preparatory roles justified by the difference between the observed state and the target event: |
| TERMINATE: Conclude or safely interrupt an ongoing event when its continuation conflicts with the updated prompt. |
| RELEASE: Remove interaction dependencies that prevent the target event, including obstructing held objects, contacts, occupied hands, or incompatible subject–object relations. Preserve compatible interactions. |
| ALIGN: Establish missing prerequisites of the target event, including subject pose, spatial arrangement, interaction configuration, or viewing condition. Use separate preparatory actions when needed. |
| Planning instructions |
|---|
| 1. Ground the state. Describe what the image shows that matters for the target event, including unfinished actions, interaction dependencies, subject pose, spatial configuration, and available entities. |
| 2. Identify unmet conditions. Determine whether an ongoing event conflicts with the update, whether an interaction obstructs the target, which prerequisites remain unsatisfied, and whether a required entity is absent. Do not introduce a transition action without a corresponding need in the observed state. |
| 3. Construct the intermediate plan. Select the necessary TERMINATE, RELEASE, ALIGN, and ENTRY actions and order them according to their dependencies. Each action should start from the preceding action’s resulting state and move toward the conditions needed for the target event. |
| 4. Allocate durations. Distribute the transition budget among the intermediate actions. For each segue prompt, provide its role, action description, duration, and brief start and end states. |
| After the intermediate plan, return control to the original target prompt. EXECUTE denotes this handoff and must not appear as an additional segue prompt. |
| Generation Setting | Existing Evaluation Focus | Semantic Transition Evaluation | ||||||||
| Benchmark | Multi- | Runtime | Update | State | Boundary | Prior-Event | Transition | State | Event | Overall |
| Prompt | Switch | Fulfillment | Retention | Smoothness | Resolution | Plausibility | Readiness | Initiation | Coherence | |
| General video generation benchmarks | ||||||||||
| VBench ( Huang et al., 2024 ) | ||||||||||
| EvalCrafter ( Liu et al., 2024 ) | ||||||||||
| T2VBench ( Ji et al., 2024 ) | ||||||||||
| Prompt template |
|---|
| System message |
| Write short, concrete video captions describing visible events. Use present-tense sentences without metaphors, camera jargon, or narration of intentions and emotions. Return only the requested JSON object. |
| User message |
| Subject and appearance: {subject} . Visual style: {style} . Setting: {scene} . Category: {category} . Relevant prop or secondary entity: {entity} . |
| Construct six consecutive steps, each lasting approximately 10 seconds, following this event outline: {category_outline} . Preserve the assigned appearance and scene. Describe a distinct visible event in each step. Where a secondary entity is introduced, show it arriving and keep it present for subsequent interaction. |
| Return a JSON object with scene , identity , and six actions . The identity sentence describes the persistent subject and setting; each action refers back to the subject with a pronoun. For subject-free scenarios, return only six actions describing the supplied scene. |
| Metric | Frames supplied | prompt text |
|---|---|---|
| SPF, SPP | 3 uniformly spaced frames within the 10-second segment. | Current prompt |
| SMC | 6 consecutive frames near the segment midpoint. | None |
| VBS, MBS | 2 frames in and 2 in . | None |
| OSE | 1 reference at and 4 in . | Both prompts |
| NSE | 1 reference at and 4 in . | Both prompts |
| SP, ScP | 2 frames in and 2 in . | Both prompts |
| Prompt template |
|---|
| System message |
| You are a transition planner for streaming video generation. Given a current prompt and an updated prompt, construct a short, physically plausible transition between them using the Transition Dependency Schema. |
| Describe each segue prompt as an observable change in progress. Make each caption self-contained, preserve subject identity and scene context, and ensure that one step establishes the conditions needed by the next. Introduce missing entities through a plausible observable process. Leave the target event to the original updated prompt after the transition. |
| User message |
| Current prompt: {source_prompt} |
| Updated prompt: {target_prompt} |
| Metric | SPF | SMC | SPP | VBS | MBS | OSE | NSE | SP | ScP | SUF | TP |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Spearman | 0.82 | 0.88 | 0.78 | 0.83 | 0.85 | 0.93 | 0.90 | 0.88 | 0.91 | 0.90 | 0.89 |
| Hyperparameter | Setting |
|---|---|
| Sampling and planning | |
| Student sampler | Four-step causal sampling |
| Sampling timesteps | |
| Timestep shift | |
| Generation block size | latent frames |
| Video frame rate | fps |
| No. | Criterion (YES = 1, NO = 0) |
|---|---|
| 1 | The main content described in the prompt is visible in the frames. |
| 2 | The appearance attributes stated in the prompt (clothing, colours, materials, textures, forms) match what is shown. |
| 3 | The core action or change described in the prompt is taking place in the frames. |
| 4 | Every entity or object named in the prompt is present in the frames, with none missing. |
| 5 | The entities and objects in the frames behave, or are used, in the way the prompt describes. |
| 6 | The scene or environment shown matches the prompt description. |
| No. | Criterion (YES = 1, NO = 0) |
|---|---|
| 1 | The position of the moving content changes gradually across frames with no sudden teleportation. |
| 2 | The shape or posture of the moving content transitions smoothly across frames with no abrupt jumps. |
| 3 | No part of the moving content shows clipping, fracturing or abnormal distortion. |
| 4 | There is no flickering between frames (no sudden jumps in brightness or colour tone). |
| 5 | There are no frame skips or scene resets between the 6 frames. |
| 6 | Every moving element in the scene has a continuous trajectory with no sudden appearance or disappearance. |
| No. | Criterion (YES = 1, NO = 0) |
|---|---|
| 1 | The structure of every depicted body or form is coherent, with no parts bent or deformed beyond a natural range. |
| 2 | Weight and support relationships in the frames are physically reasonable, with nothing floating or unsupported without explanation. |
| 3 | Contact with the ground or supporting surfaces is physically plausible, with no clipping through surfaces and no unexplained hovering. |
| 4 | Every object in the frames is positioned and oriented in a physically plausible way. |
| 5 | The relative size and scale of all elements in the frames are realistic and mutually consistent. |
| 6 | Contact points between elements show no clipping or interpenetration. |
| No. | Criterion (YES = 1, NO = 0) |
|---|---|
| 1 | There is no flash, white frame or black frame between frames 2 and 3. |
| 2 | There is no visible tearing, blocking or encoding artifact between frames 2 and 3. |
| 3 | There is no abrupt drop or rise in resolution or sharpness between frames 2 and 3. |
| 4 | The camera focal length or viewing angle between frames 2 and 3 stays on the trend established by frames 1 and 4, with no sudden change. |
| 5 | Overall, the change in brightness, colour tone and contrast from frame 2 to frame 3 does not feel more severe or more sudden than the change from frame 1 to frame 2 or from frame 3 to frame 4. |
| 6 | Overall, the four frames look like continuous footage from one take rather than two clips edited together. |
| No. | Criterion (YES = 1, NO = 0) |
|---|---|
| 1 | The relative motion of scene elements across frames 2 and 3 is consistent with the apparent camera movement, allowing for independent object motion and depth-dependent parallax. |
| 2 | Shape and posture change gradually from frame 2 to frame 3, rather than jumping straight to a completely different configuration. |
| 3 | No part of the content shows clipping, fracturing or implausible distortion in frames 2 and 3. |
| 4 | The motion keeps a consistent speed and direction from frame 2 to frame 3, with no sudden reversal or stall. |
| 5 | The facing direction of the content does not turn implausibly between frames 2 and 3. |
| 6 | The spatial structure of the background stays consistent between frames 2 and 3, with no jump to a differently laid-out space (judge layout only, not colour or style). |
| No. | Criterion (YES = 1, NO = 0) |
|---|---|
| 1 | The process under way in frame 1 has a recognisable ending in frames 2-5, rather than vanishing between two adjacent frames. |
| 2 | The old state is not frozen in an obviously unfinished intermediate form. |
| 3 | The form in which the old state ends is one in which that kind of process could plausibly stop in reality. |
| 4 | The ending of the old state leaves nothing suspended and unresolved (no change stopping halfway with no follow-through). |
| 5 | Every element taking part in the old state in frame 1 has an explicable fate in frames 2-5, with nothing vanishing or resetting out of nowhere. |
| 6 | The way the rate of change winds down is explicable, with no unattributable instant drop to zero. |
| No. | Criterion (YES = 1, NO = 0) |
|---|---|
| 1 | The new state starts from the state shown in frame 1, rather than starting from a completely different starting point out of nowhere. |
| 2 | The spatial configuration of the elements when the new state begins is reached from the configuration in frame 1 through a visible process of change. |
| 3 | The build-up or preparation that the new state requires is visible in frames 2-5 rather than skipped. |
| 4 | The first visible stage of the new state is genuinely its beginning, rather than opening from the middle or from an already completed form. |
| 5 | No necessary intermediate step is missing between frames 2-5 (the viewer does not have to imagine an unshown transition). |
| 6 | Where an element present in frame 1 changes its position or form, that change is visible as a process rather than happening instantaneously. |
| No. | Criterion (YES = 1, NO = 0) |
|---|---|
| 1 | Facial features (face shape, proportions of the features, skin tone) stay consistent across the four frames. |
| 2 | Hairstyle and hair colour stay consistent across the four frames. |
| 3 | Clothing style stays consistent across the four frames. |
| 4 | Clothing colour stays consistent across the four frames. |
| 5 | Body shape and proportions stay consistent across the four frames. |
| 6 | Identity attributes such as gender, apparent age and ethnicity stay consistent across the four frames. |
| No. | Criterion (YES = 1, NO = 0) |
|---|---|
| 1 | The amount by which the spatial structure of the background changes across the four frames matches the environmental difference implied by the two descriptions. |
| 2 | The amount by which the lighting environment (direction, intensity, colour temperature) changes across the four frames matches the environmental difference implied by the two descriptions. |
| 3 | No environmental element appears in frames 3-4 that neither description mentions. |
| 4 | No environmental element present in frames 1-2 disappears in frames 3-4 in a way neither description can explain. |
| 5 | Any environmental change across the four frames happens progressively, with no wholesale instantaneous replacement. |
| 6 | The spatial relationship of the content to background landmarks stays coherent across the four frames, with no teleporting into another space. |
| No. | Criterion (YES = 1, NO = 0) |
|---|---|
| 1 | Frames 3-6 show an observable change compared with frames 1-2. |
| 2 | What changed is exactly what the semantic difference between the two descriptions refers to. |
| 3 | The parts the two descriptions share (content, scene, style) stay unchanged across the six frames and were not altered along with the rest. |
| 4 | The magnitude of the visual change is commensurate with the semantic difference, without being clearly excessive. |
| 5 | The magnitude of the visual change is commensurate with the semantic difference, without being clearly insufficient. |
| 6 | What the new description refers to becomes established in frames 3-6, rather than appearing only as a brief trace that disappears again. |
| No. | Criterion (YES = 1, NO = 0) |
|---|---|
| 1 | The new content begins to appear soon after the switch, with no long lag spent still showing the old content. |
| 2 | The change is not completed instantaneously between two adjacent frames; a transition process is visible. |
| 3 | The duration of the transition suits the magnitude of the change, without feeling rushed. |
| 4 | The duration of the transition suits the magnitude of the change, without feeling drawn out. |
| 5 | The pace of motion around the switch stays natural, with no unreasonable sudden acceleration. |
| 6 | The pace of motion around the switch stays natural, with no unreasonable sudden deceleration or stall. |