From One Photo to a Live Sports Broadcast with COLA

How a simple node-based workflow turns a character reference into a realistic baseball fan reaction video.

Creating a believable AI video usually takes more than a single prompt.

You need to preserve the character, build the right environment, choose a convincing camera angle, and then animate the scene without losing the identity established in the original image.

In COLA, all of these steps can be connected into one visual workflow.

This example starts with a casual character photo and turns it into a realistic sports broadcast shot of the same person watching a baseball game from the stadium stands.

The complete COLA workflow: reference image, scene generation, video animation, and final result connected on one board.
The complete COLA workflow: reference image, scene generation, video animation, and final result connected on one board.

Step 1: Start with a character reference

The first node contains the original photo.

This image is not used as a composition reference. Its main purpose is to define the character’s identity: facial features, hairstyle, age, body type, and overall appearance.

The workflow prompt makes identity preservation the highest priority. At the same time, it gives the image model permission to change the pose, clothing, background, and camera setup.

That distinction is important.

Without clear instructions, an image model may either copy too much from the reference or create a different person who only looks vaguely similar.

The original character image acts as an identity reference rather than a fixed scene or pose.
The original character image acts as an identity reference rather than a fixed scene or pose.

Step 2: Build the stadium scene

The reference image and the first prompt are connected to a GPT Image model.

The model creates a completely new scene:

The generated image already looks like a frame captured from a live television broadcast.

This intermediate step gives the video model a strong and controlled starting point. Instead of asking the video model to invent the character, location, framing, and movement at the same time, the workflow first locks down the visual direction.

GPT Image 2 creates the new stadium environment while preserving the character from the original reference
GPT Image 2 creates the new stadium environment while preserving the character from the original reference

Step 3: Animate the broadcast shot

The generated stadium frame is then connected to Kling 3.0.

A second prompt describes how the final video should behave.

The character casually watches the game, follows the action with their eyes, blinks naturally, changes facial expression slightly, and occasionally shifts their gaze. The movement is intentionally restrained.

This is not meant to look like a staged performance. It should feel like a candid reaction captured by a television camera during a real game.

The prompt also keeps the camera stable and preserves the visual language of sports coverage: shallow depth of field, compressed telephoto perspective, broadcast graphics, and documentary-style movement.

The final Kling 3.0 animation adds subtle eye movement, blinking, and natural reactions while keeping the sports-broadcast framing stable.
The final Kling 3.0 animation adds subtle eye movement, blinking, and natural reactions while keeping the sports-broadcast framing stable.

Use the clearest frame where the character is looking toward the field and the broadcast overlay is visible. Splitting the workflow between an image model and a video model gives each stage a focused task: GPT Image handles identity, clothing, environment, composition, and lighting, while Kling adds blinking, eye movement, subtle reactions, and natural posture changes. This approach provides more control, makes results more consistent, and lets you reuse the same structure for concerts, interviews, football matches, fashion events, documentary portraits, and other scenes. In COLA, every step remains visible and editable on one board, turning a single generation into a reusable creative workflow.

Build visually, iterate quickly

COLA lets you create boards and connect text, image, and video models through a node-based interface.

Every stage remains visible:

This makes complex AI generation easier to understand, test, and reuse.

You are not just generating a single image or video. You are building a workflow that can be adjusted, duplicated, and applied to new ideas.

One reference image. Two prompts. Two specialized models. One finished broadcast-style video.

That is the idea behind this COLA workflow: break a complex creative task into clear stages, give each model a specific role, and keep the whole process visible on one board.