From ComfyUI to LibTV: What Workflow Orchestration Needs in the Video Generation Era

Hyacehila

The most common action in the production of pictures in the last two years is probably a draw card.

The questions in this article can also be addressedTouchDesigner Lights with 3D Gaussian SplattingHow the game industry is introduced AI AgentHow the concept of a relatively close read together is developed in different contexts.

The reason is simple: the model is originally stylish. A new id, a little hint, a LoRA, a reference map, or even nothing, could turn the result from ordinary to usable. The models, Stable Diffusion, SDXL, Flux, have made many people accustomed to the process of "scattering, picking, redrawing, re-expanding" and "scrambling." The ComfyUI can become important at this stage, and it's about this. It breaks the generation process into nodes and links, and transforms the original sub-point button into a working stream that can be saved, modified and reused.

Now, video generation is also entering this phase.

Just video is much more troublesome. The image card failed, at most, this one failed; the video failed, probably the role first second and the third second face floated, or the last and next part could not be connected, and the camera was moving well but the product was deformed. Many of the current video models are still based on short clips, single-scenes, Vincent's or graphic videos. For example, Stable Video Diffusion, which is published with 14 frames and 25 frames, supports a more robust closed-source model such as Seedance 2.0 and HappyHorse for up to 15 seconds. But people obviously need more than 15 seconds.

This is not a ComfyUI curriculum. I'm not a professional user of ComfyUI either. More precisely, I want to re-understand it from a workflow perspective: What is ComfyUI now? Why is it so weird to call it Cost Map Software? When video generation is becoming a common feature, does it need to be supplemented by some long video-oriented organization like LibTV?

My judgment is here: ComfyUI has moved from a living map UI to an open-source workwork/runtime; video generation is pushing the problem to a higher level. It is worth filling some of the camera, segment, bulk-script cards, verification and long-term mission management capabilities, but it does not have to become LibTV. The most interesting place in ComfyUI is open, manageable, localized, model ecology and node combination.

What's the matter now?

If you look at it from the interface, ComfyUI is easily understood as a tool to draw a flowchart of Stable Diffusion. This was never wrong.

Official documentsDescribe ComfyUI now as node-based interface and inference engine for active AI. It deals with workflow: a set of calculations linked to nodes and links. Node can load models, coded hints, sampling, decoding, saving results, or tasks related to controlNet, IP-Adapter, LoRA, magnification, video synthesis, audio processing, 3D or angent. So, ComfyUI's concern is what the steps are, not just the last one.

Traditional WebUI is more like the parameter panel. User fills prompt, selects models, moves, points generation. ComfyUI spreads these steps so that models, parameters, inputs, and outputs are visible nodes. A KSampler node connects prompt embedding, latent, model, sampler and seed; a VAE Decode node restores the latent as an image; and a Save Image node sets the disc. It looks complicated, but it allows users to see the process.

If it's the first time I understand ComfyUI, I'll grab a few ideas first and I'll not rush back to nodes.

Look first. Node graph. The node chart is not for show, it simply puts the process of generation on the surface. The image generation is followed by a model loading, text encoding, latent initialization, sampling, decoding, reprocessing, etc. The chain is longer in video generation: reference maps, front end frames, motion modules, frame patches, magnification, and collages may also be addressed. The nodal charts break down these steps and the user can replace one of them without a whole process starting from scratch.

Look at the workworkwork. Workflow is a generation configuration that can be saved and reproduced. In addition to prompt, it records nodal structure, model selection, parameter connection and processing order. In many cases, the real influence on results is the state of modelling, reference diagrams, ControlNet strength, LoRA weight, denoise, sampler and reprocessing.

The following is a basic set of nodes and their general sequence:

Load Checkpoint:请模型上场
CLIP Text Encode:翻译提示词
Empty Latent Image:准备空画布
KSampler:真正生成
VAE Decode:把 latent 变成图片
Save Image:保存结果
Load Image:导入参考图
VAE Encode:把图片压回 latent
Load LoRA:加载风格/角色补丁
ControlNet:控制结构
Upscale:放大增强
Video Combine:把帧合成视频

Local models are also important. The ComfyUI capabilities are derived from local open source models and community open source nodes. This information can be accessed through the Load Checkpoint, Load Lora, VAE, and ControlNet nodes. As the community expands, images and video links such as Flux, Wan, AnimateDiff, Stable Video Diffusion, HunyuanVideo can be accessed in some form. This is where it's not the same as the pure cloud that webui produces. Users can organize their generation tasks into a workflow on their own machines, using information from the community as a whole.

Finally, programmable.ComfyUI Server API Supports the submission of workwork as task, via queues, history and WebSocket listening performance. External scripts or angents can modify the id, prompt, reference diagrams and parameters, run the mass results, and retrieve the output. Here, ComfyUI is not just an interface, but can be used as a backend to move.

The current ComfyUI can be understood as a stack of three things: visualization to generate a map, local/mele-debate running time, workflow system that can be called by the program. It was first popular with the living map, but it is no longer enough to talk about it.

Here's a living map, Pipeline. The view ahead is based on local models, and many people now generate pictures through external API and partner Nodes. A base Pipeline often uses fast-track patterns to distinguish styles, lot-draw cards, selects them in 4 to 10 diverse images, then cuts to higher quality models to make multiple fine-tuning. The system also needs to be evaluated automatically, and it is common to introduce external LLM Rubric, reduce manual repetition of the hints and leave people in a situation where they really need to be judged.

Where are you going?

If you look at the complex nodes in the community, ComfyUI seems to be a technology user forever. But official moves in recent years are clearly pushing it out. The node chart is still controlled, and complex workwork is beginning to be packaged for call by ordinary users, angents and cloud systems.

A line is apoIization. Local API files and workflow API format: users can export work streams into API format, by /prompt Submit to queue, pass /queue Look in the queue, pass /history Take the results, wire the execution through WebSocket. That means the ComfyUI is fit for a bigger system. An external script can read 100 prompts, assign 20 seeds to each prompt, submit tasks in bulk, and send output maps to CLIP, OCR or VLM.

The other line is partner Nodes.Partner Nodes External API services, closed-source models or third-party hosting models can be fed into ComfyUI workworkworkworkflow. This direction is realistic, as production processes rarely rely on a single model. Local models may be responsible for low-cost cards, closed-source models for high-quality graphic videos, VLM for auditing, OCR for text checking, and human face models for consistency. To become a workforce/runtime, the ComfyUI needs to keep these capabilities in the same picture.

And App Mode, Agent and Cloud. Official Blog From Workflow to App It was mentioned that App Mode could package complex workflows into more apparitional interfaces;Agent Tools Let angent call ComfyUI generate image, video, audio, 3D;Comfy Cloud The project is aimed at the cloud, working towards the cloud, deploying as API, running in parallel, working with multiple workworkwork, working with a team.

These things together, Comfyui is no longer a single-point tool. It is more like a platform for generating media workflows: the bottom is an open nodes, with API, queues, clouds and model connections in the middle, and the upper layer wraps the workflow into apps or angent tool.

That is why long video programming becomes a problem. If ComfyUI is just a tool to run a picture at a time, of course the long video is far from it. But if it is already workflow/runtime, it will sooner or later come before it, succession, screening, re-testing and cluttering.

Videos are not a single creation that can solve this.

The problem with long video is not to lengthen the prompt, nor to fill the data in the video model.

Under current model conditions, many video production is more like a multistage water line. Mr. is a role or product reference map, completing the initial design of the entire script and lens. Using the biograph model, key frames are drawn, key frames are converted into short videos and collage checks are performed. If a section is broken, the corresponding camera is to be reticked, without having to reverse the whole video.

The image generation often requires only this one answer enough. Video production takes a lot of thought: whether the role is drifting, the scene is changing, whether the last frame and the next frame can be connected. The camera structure, video motion, text and product details can be problematic. There's the cost. The video was retweeted, and it hurt a lot more than the picture.

At the same time, it is constrained by capacity and cost constraints to generate models, and the basic actions for long video (greater than 15s) are segments. After the subparagraph, the real problem arose: how could information be shared between the paragraphs?

If the first paragraph confirms the role image, the second paragraph should follow that role reference. If the first paragraph had a good frame, the second paragraph could be used as a frame of reference or reference. If the third shot is product-specific, it should inherit product maps, brands, materials and graphic requirements. If there are 20 candidates for a shot, the system should be able to record which one is selected, why is it selected, which version is being used in the next phase.

And then, workflow itself is not enough. A ComfyUI workworkworkwork can describe how well a camera is generated, but it is difficult to describe the video naturally with 24 shots, with 10 candidates for each shot, the first six shots sharing the same character reference, the seventh shot using the end frame of the sixth shot, and the 10 shots sharing some kind of environment between the 10 shots, but the person may need to change his clothes. The problem has passed the individual nodes to the project level, the lens level, and the task level.

The problem of long video has changed from a problem of generation to a systemic problem. Modelling capacity is certainly important, but the structure of the material, the sequence of implementation, the screening mechanism and the failure to retest are equally important. Without these external structures, models are stronger and can easily be stopped in a good short clip, making it difficult to become a deliverables (e.g. short play).

Why do you take LibTV as a reference?

I didn't take LibTV out to say it was a substitute for ComfyUI, or write an evaluation for it. It's more like a reference here: when a product is directly oriented towards video creation, what is the structure of the upper layer that it naturally grows?

LibTV is more like an AI Video Production Workstation. It's about mining samplers, VAE, Denoise, these bottom parameters, and it puts the interface on scripts, role, reference diagrams, video clips, sessions and results. Official libtv-skills The script capabilities provided in the warehouse include creating sessions, sending creative instructions, referencing the progress of sessions, uploading files, downloading results, etc.;LibTV CLI The page also highlights the possibility of calling LibTV in antgent tools such as Claude Code and Codex to complete pictures, videos and role generation.

This is not the abstract level of ComfyUI. ComfyUI has a more logical node: which model is loaded, which conditions are used, how sampling is done, how decoding is done, how the results are transmitted to the next node. LibTV targets are more like creative assets: what role is in the project, what is the lens that is generated, what is the reference chart, where is the current session going, how results are downloaded, how materials continue to be used. The difference between the two also arises from the stage at which the model is generated: The former comes from the peak of open source mapping models, while the latter is a tool for long video using closed source models.

LibTV's inspiration for ComfyUI, not to hide all the nodes, which would lose his original advantage. On the contrary, ComfyUI should not have lost the node chart. What really deserves to be drawn on is that video production requires a higher organizational unit than a single workworkwork. The graphic age, where a workflow generation is useful, the video age, where users need to work around scene, shot, aset, Candidate, version.

What should we fill in?

If I look at this, I won't first fill ComfyUI with a button. This is complemented by a meso-level organization capability. It is caught between the bottom generation model and the upper App/Agent, and controls the segments, inheritance, screening and long assignments in video production.

I'll fill the Shot/Scene layer first. ComfyUI can continue to keep workflow as a bottom-up performance chart, but add lenses and scene concepts to workworkflow. A video project can contain multiple scene, each scene contains multiple shots, each of which is sometimes long, descriptive, reference, front frame, tail frame, workworkwork, candidate results and current status. User-managed is a video structure, not a pile of scattered output files.

Then there's a batch of the card layers. A shot should not be created once, but should generate a N candidate. The system needs to know which candidates come from the same shot, what the parameters are, what the cost is, which is being phased out and which is going to the next stage. This is not necessarily a complex capability, but it is less eccentric to manually copy workwork, seed, search for documents.

The verification layer also needs to be filled. It doesn't have to be automatic at the beginning, and it doesn't have to be a fake VLM that can judge everything. It is more realistic to combine manual and model ratings. For example, VLM checks for lens description, CLIP checks text relevance, OCR checks image text, human face similarities check role consistency, NSFW models do security filters, and can still be ticked or fork-marked. It is not intended to replace aesthetics, but to structure the selection results so that they can be returned to the next generation.

The inheritance layer of assets may be more fundamental. Long video cannot start from scratch, and what has been identified must be brought down. Role references, product references, scene references, style references, and a frame of last lenses should be passed as assets between different workworkworks. The assets here are not just document paths, but also include uses: This is the role face reference, this is the costume reference, this is the scene depth map, and this is the last schematic that can be prepared.

Long-term mission management is also not in place. Video generation fails, queues, breaks, and retrys. ComfyUI already has queues and history, but long video requires a project-level mission perspective at the lens level. The user wants to know whether a node has been implemented, but "The 12th shot and 3 candidates are still running." "The 5th shot has passed but not been zoomed up." "The 8th shot has failed twice, and the model should be changed."

Finally there's a light timeline. ComfyUI does not need to become Premiere, nor does it need to do the complete editing software. However, if it is to support video production, it can provide at least a light time line arranged by lens, preview, replacement candidate, and export rough clippings. Users should be able to see the whole video from the structure instead of opening mp4 in the folder.

These abilities, taken together, are not the transformation of ComfyUI into LibTV. They are more like a production syntax that is needed to supplement the existing nodes. Address what is generated, in what order, how to inherit, how to select, how to continue.

It doesn't have to be LibTV.

I think it's more rational to go in the direction, maybe three layers.

The bottom is still node graph. Professional users control models, parameters, reference maps, sampling, magnification and reprocessing.

The middle level is orchestration. It is responsible for the lens, scene, mass candidate, verification, asset inheritance, long assignments and versions. It can be a specialized layer of video generation; the living map and video generation evolve separately, and then share roles, products, scenes and style assets through a resource layer. It's just a conversation about a direction, and how to do it depends on what the Ecology of ComfyUI looks like.

Upper tiers are AppMode, Agent or Studio UI. Ordinary users do not need to see all nodes, but simply upload the material, fill in the requirements, select candidates, confirm the results. Complex workworkwork can be encapsulated as an application or called by ant.

A really useful system may draw on both sides of the equation: the bottom is as open as ComfyUI, the upper is as organized as LibTV. The rise of video generation forces these runtimes to face more trouble: How to generate short clips once and for all, organize into long video production processes that can be created, screened, inherited, retested and delivered.

Maybe the future video model will be strong enough to produce a complete short film. Until then, however, the workflow was still in a state of incoherence. Even if the model eventually produces short films for a few minutes or hours, a truly valuable long video production still requires a layer of understanding lenses and project structure, rather than giving a sentence for prompt and leaving everything else to the model.

References

  • Title: From ComfyUI to LibTV: What Workflow Orchestration Needs in the Video Generation Era
  • Author: Hyacehila
  • Created at : 2026-06-18 07:00:00
  • Link: https://hyacehila.github.io//blog/2026/06/18/comfyui-video-workflow-orchestration/
  • License: This work is licensed under CC BY-NC-SA 4.0.
Comments