Every few months a new AI video model gets released and someone announces that 3D animation is finished. 3D designers must pack and go home. At Kasra Design®, we do both type of work so we decided to show our readers how these 3D AI projects compare against the traditional 3D work. This article (or you can even say case study) looks at 2 x 30 seconds commercials we produced, one built entirely in 3D and one built in 3D and then finished with AI (materials, lighting, rendering). We look at the real process behind each type, the time each one took as well as the shortcomings.
Two Commercials, Two Methods
The first project is for Pair Eyewear, a spot featuring five styles of eyewear frames in 30 seconds. It was built the traditional way. We modelled every frame, created materials from scratch and rendered the scenes in 3D with Octane Renderer. No AI anywhere in the pipeline. You can watch it here:
The second is Twex, a chocolate commercial for a fictional brand that we produced in house as part of our R&D. It uses a hybrid method: The base model is created and animated in 3D, then it ran through an AI video model to generate the final realistic look. The finished spot with a before and after comparison built into the end is here:
Before we break things down, it is worth checking what each of these two methods actually means as the terms get used loosely.
What Each Method Actually is?
Traditional 3D Animation
Traditional 3D animation means every element you see was built and rendered inside a 3D software. A 3D designer builds the objects as geometry. Another artist creates the 3D materials for various rendering engines. A lighting artist places the lights around the scene and an animator animates the objects / characters.

The cameras are set and a render engine calculates every frame, pixel by pixel. Nothing is generated or guessed. If the client wants the logo two millimetres to the left, you move it two millimetres to the left, and it stays there in every frame. The upside of this method is that you get total control and exact accuracy. The cost of it is time, because all of that precision is built by hand.
AI Hybrid Animation
The AI assisted hybrid 3D method starts similarly. We create the objects or characters in 3D software. But instead of taking the 3D scene all the way to a polished final render, we settle for a simple clay render. With correct shapes, correct movements and speed as well as correct camera angles. And all in plain unpolished surface like the whole scene was sculpted from clay. That pass locks the things a creative director cares about such as the composition, the timing, speed of movements and the camera angles.

After the above, the clay render is fed into an AI video model (e.g Seedance or Kling Ai) which generates the photorealistic renders on top of it. This way the 3D controls the structure and the AI delivers the final polished look and feel. If I want to describe it in one line, I would say:
The 3D is the skeleton. The AI is the skin.

This is a different thing from typing a sentence into an AI and hoping. That is the workflow most people mean when they say AI video and it is why so much of it looks impressive for three seconds and then falls apart.
The hybrid keeps a human director in charge of everything structural and lets the AI do only the part it is genuinely good at. Our AI assisted video service is built on this approach instead of the prompt-only approach.
Pair Eyewear: Built Entirely in 3D
Starting From the Moodboard
Fortunately, we were provided with a moodboard. That single document was very useful. It told us the colours they were using in their marketing funnels, the type of lighting they wanted (warm, summer vibe), and the materials that mattered to them. They liked to use linen fabrics along with their products.

Pro Tip: If you are a 3D team, always ask if the brand can provide a moodboard. It will save you a lot of guesswork.
Five Products in 30 Seconds
The brief asked for five different frames of glasses inside a thirty second spot. That was quite tough to do given the short duration of the video. Each product needed its own moment to be noticed clearly, and then they all needed to come together. We featured them one at a time and then as full pack-shots.
We linked them with “match cuts”, a type of transition that a shape or a movement at the end of one scene carries into the start of the next. The scenes were never physically connected, but the eye is led smoothly from one to another. It generally makes the film feels like a single continuous idea rather than six separate clips stitched together.
Pro Tip: This technique is particularly useful when you don’t want many elements in your 3D scene which increases the sizes of the scene and can cause loading or rendering issues.
Why Materials Were the Hardest Part
We modelled the frames to match the real product as close as possible. This was important because the spot is full of extreme close-ups and a close-up shows all the flaws you wouldn’t want the viewer to see. At that distance the viewer is literally inspecting the product so any difference between the animation and the real thing becomes obvious.
Getting the geometry right was half of the struggle. The harder half was materials. The product uses some special materials that pass the light through (think SubSurface Scattering type) and some of their models have unique patterns. Materials were the single most time consuming part of the whole project and that’s why traditional 3D takes the time it does. There is no shortcut to achieve a surface that can survive a full-screen close-up.
The Part Nobody Costs In: Social Media Aspect Ratios
The spot had to be delivered in five different ratios. One widescreen 16:9 master and four social media ratios such as 9:16 and 1:1. This is where a lot of people underestimate the work, because it sounds like an export setting. It is not!
You cannot simply crop a widescreen video into a tall one. If you do, products would fall outside the frame. So for every ratio we rebuilt the cameras and adjusted the animation keyframes so that each product sits nicely inside that specific ratio. Five ratios means composing the 3D animation five based on the original scene.
This process is time-consuming, so when there is a need for social media ratios, you need to take to account the time and cost of this properly.
Twex: 3D for Control, AI for Skin
Why We Did Not Just Prompt an AI?
Twex started from this one question: “How can we have the same level of quality as video generated by an artificial intelligence system, while keeping our control over what happens in that video?”
Text-to-video and image-to-video (text-to-image) are both really fast and good enough for many applications where speed is the main goal. However, they require that you continually negotiate with a prompt about filmmaking stuff that a director should fully control. E.g where the camera is placed, the length of a scene, composition of a shot or exactly when the caramel splits (in this case).
These Are Artistic Decisions.
Giving these decisions to a machine that interprets your prompts differently each time is a bad practice. Instead of being the director, you spend your time cajoling the machine.
The Skeleton and the Skin
We modelled and animated the movement of chocolate pieces in 3D. It wasn’t very detailed because it didn’t need to be. In fact, it was proven later on that the simpler the structure is, the less it interferes with AI interpretations.

Giving a too detailed model to Seedance sometimes will make it think it is polished enough, so it stops getting properly prepared based on realistic references. We locked the camera angle and the frame rate for each shot exactly how we wanted them. Then we rendered those frames in clay (no texture/materials).

Next, we sent the clay style rendered video to “video-to-video” Seedance 2.0 through OpenArt platform which generated the final photorealistic details on top of our 3D base. As expected, AI followed our exact camera angles and movements since we have given it all in a video.

All parts of this 30-sec commercial, from 3D setup to post-production and color correction were completed in approximately 1-2 weeks.
Where the AI Still Fights Back
It is dishonest to say this was a straightforward process. The hybrid method is much more controllable than a prompt-only AI model, but it is still less controllable than working entirely in a 3D pipeline. We had to put a lot of actual time and effort into getting the model to behave.
Sometimes the AI rendered the background in the incorrect color. Sometimes it placed the lighting on a shot from the opposite direction (ignoring what the light source should have been). Occasionally, it also invented ingredients for the chocolate that were never even in the reference image, or would place those inventions floating in the air.
These aren’t fatal issues, but can’t go into a broadcast ready piece. Getting a clean shot required trying multiple generations.
So in my opinion, the Hybrid Method provides a controlled process, however, it does not provide an automated process.
Revisions for a Hybrid AI Project
Client revisions are where a workflow is really tested because no project of value is approved on the first pass. But hybrid method can come with a specific structure that keeps changes sane:
Revising a Structural Change
We get approval for the 3D clay animation first, prior to any AI being used on it. If the client has a structural change they wish to make (a camera move, time adjustment, product shape change etc.) at this point, we can go back into Cinema 4D or Blender, adjust accordingly, and then resubmit the updated clay render for further approvals. As this is a “real” 3D layer of our project, we have an exact and predictable result, just as we would when working on a fully traditional project.
Revising the Final Look
If the client wishes to simply tweak the end-look (color, finish, surfaces), we will go back to the AI stage, create a new prompt and reference image based upon their request and run it through the system.
The 3d base remains unchanged.
AI in the Traditional Pipeline Too
AI can be used internally with “traditional” 3d projects.
In our studio, we use some of these AI still image generation tools like Nano Banana Pro for developing concepts before we even start modeling anything for a traditional project. We start by creating an AI generated concept that shows the mood, camera angle, light and color directions. We then show those to the client. Once the direction is agreed by the client, we can then build the entire scene from scratch in full 3D. The benefit here is certainty.
When we begin to model a scene or character, we have already received approval for the mood the client is looking for and what type of lighting will be acceptable to them as well as which colors will be appropriate.
Therefore, by the time we get into modeling, we don’t have to worry about revising anything. This is one of those AI usages that can assists artists rather than trying to replace them.
Where the Time Goes, and Where the Cost Follows
The clearest way to see the difference between the two methods is side by side. Both columns below are for a thirty second commercial and both estimates include post-production, colour correction, sound and music.
| Traditional 3D (Pair Eyewear) | AI Hybrid (Twex) | |
|---|---|---|
| How it is made | Every frame modelled, textured, lit and rendered in 3D | Animated in 3D to a clay render, then AI generates the surface |
| Best suited to | Real products that must match exactly, close-ups, brand-critical accuracy | Organic, textural subjects where richness matters more than exact accuracy |
| Accuracy | Near exact match to the real product | Hyper realism, but the AI can drift and needs constant correcting |
| Timeline (30s spot) | 4 weeks, with 6 to 7 preferred to refine | 1 to 2 weeks |
| Revisions | Exact and predictable, changed directly in 3D | Structure changed in 3D, surface changed with a new AI prompt |
| Relative cost | Higher: more render hours and hand-built materials, plus per-ratio rebuilds | Lower: fewer render hours and a faster turnaround |
Why the Hybrid is Faster
The time difference isn’t just about how fast the AI can work. It’s about skipping over one of the slowest stages in the original process. In the eyewear spot, materials are always going to be the slowest step (all those hours trying to get the 3D material to behave as you want it to when light hits).
So since the AI will supply the surface for the hybrid model, this stage becomes almost non-existent.
Why the Cost Differs
The cost is also based on the same principles. Since traditional 3D animation requires many more skilled man-hours and more rendering time than hybrid animation, on an average campaign, the costs multiply.
Hybrid animation will need less time for preparation and rendering, therefore a shorter production schedule, thus lowering your overall cost.
The Tools Behind Both Commercials
For anyone curious about the pipeline, the 3D work is animated in Cinema 4D and Blender, with final traditional frames rendered using Octane rendering engine. The AI stages use Seedance 2.0 for the video-to-video generation, and Nano Banana Pro for still images and concept development. We have not used Seedance 2.5 yet, but it probably does a better job than its earlier model. The tools change often, and by the time you read this there will be newer ones, which is rather the point. The method matters more than the model, and the method is built to survive whichever model is current.
You can see more of both kinds of work, traditional and hybrid in our portfolio.
Speed or Control – We Offer Both
The conclusion would not be about one method beating the other. It would be about the fact that they answer different briefs and client requirements. If your project needs exact accuracy, brand-critical close-ups and total control, traditional 3D animation is the one you should opt for. If it needs richness and speed but at a lower cost then the hybrid AI-assisted approach is a good choice.
Most of the time the right answer is a conversation about the specific thing you are making. Tell us what your product is, who it is for and when you need it, and we will tell you which method fits and why. Start here.
