Description
About this prompt
FLUX 3 feels very Sora-like in how it generates video.
Explore, Create & Generate Powerful AI Prompts
FLUX 3 feels very Sora-like in how it generates video.
FLUX 3 feels very Sora-like in how it generates video. And the reason is structural not random: Both models first convert your prompt into a conditioning signal (numerical embeddings from a text encoder) which then guides a generation process that happens all at once. Sora did this with diffusion while FLUX 3 does it with “flow matching”. There are small differences but the core idea is the same: create one singular conditioning signal to steer the entiere sequence at once, instead of doing it in stages like with Grok Imagine. This is different from Kling, Seedance, and Grok. Kling and Seedance are also Diffusion Transformers, but they are optimized more heavily for fluid motion, multi shot control, and commercial/social video promoting. It’s so cinematic bro! ^^This is exactly why it’s hard to get nostalgic styles (like VHS content) with Seedance, even when you give it the perfect reference image and use every single trick in the book in your prompt. Grok Imagine is unique in that it refines the output more sequentially, chunking the generation as the video is build in steps. This is one reason why Grok’s extend feature is so powerful. FLUX 3 and Sora share a stronger “world model” training emphasis. This means broader physical and causal relationships, how the world actually works matter a lot more in their generation process. The initial conditioning signal pulls from a richer internal map of how things are actually constructed and behave in the real world. And it’s not just how they behave directly, it’s how disparate concepts can be combined to look real, even if the models had never technically “seen them in their training data”. The idea that you can only create video that has been seen in training data is simply not true with world models. Because the model constantly refers to that single signal while refining the whole clip, the signal needs to be chosen correctly. And most people thing you do this by writing out these long 3000 character prompts. But those only work if you really understand that specific model. In fact: Short, ultra-specific prompts will produce higher-information embeddings if you avoid vague and common language. The problem with long prompts is that they: add common phrases that don’t point to one specific place in the model’s training data and often times create competing instructions that dilute the signal which guides the generation toward the average of the training distribution and makes everything boring or inaccurate based on your intent. For most users, shorter prompts will lock onto specific styles and actions more cleanly on FLUX 3, just like was the case with Sora. Here is an example. This is the prompt for the video attached. ATM security camera, extreme close-up face, fish-eye distortion, keypad visible at bottom, timestamp overlay, gritty street lighting, ban transaction. The character taking money out is an adult seal, we hear the seal muttering that the sharks are after him and he shouldnt have taken money from the sharks. Suddenly this turns into display of the foodchain when two sharks walk in from out of scene and eat seal. All captured on the ATM camera. Drop a comment if you want to see an article about this with extremely accurate 1-2 sentence prompts that replicate almost exactly every time you run them.
Reviews
There are no reviews yet.