People are still figuring out the social norms surrounding "deepfakes" and I'm convinced there are some versions of the future where it becomes normalized for content under essentially fair use/free speech doctrine, and many where it's used to restrict free speech in the small number of jurisdictions where public figures would currently be allowed to be featured in this content.
Obviously there are a lot of ways it shouldn't be used but I want to live in the future where it's something we use to have fun, where reasonable, rather than a dangerous/sketchy taboo
I dont normally like AI videos but this was just amazing, by next year the tech will be perfect, cheaper, we are going to AI videos everywhere whether anyone likes it or not.
people like it. Maybe not older generations, but we're about to get a fresh new generation of people that will not have a lifetime of experience looking at human generated content, and so have no real bias against AI content.
And then every generation after that will be born into increasingly a AI-video rich world.
Older generations absolutely love AI slop. If you haven't used Facebook recently, it's all old people sharing fake videos of animals doing silly things.
Right now we're going through multiple concurrent but slightly non-overlapping transitions where AI is almost good enough for something, and seeing a predictable surge of spam/bad quality crap people don't like, followed by a whiplash effect when it reaches human or superhuman levels of performance.
I think most developers would agree that the Cursor vibe-code era sentiment towards AI for coding was right in the sense that it wasn't that useful then, and a lot of people did and do stupid things with it, but it really did deliver quite substantially on its promise just a couple years after the "consensus" was often "that's never going to work".
It's very strange to me how consistently public opinion shifts on this problem. In the long run we're all dead but whether it takes one or five years, you probably will be there to live through it. I assume whoever is downvoting you is just thinking with their consumer-brain or "I don't want that" (the way it is now) rather than that this will be no different than video games, or wasn't a child who watched youtubepoop or other weird stuff because it was funny.
The quality is quite high, but an observation is the direction they're taking these models correlates heavily to the usage demand of China vs the West. Specifically, they're immensely focused on t2v for action / high effect shots. There's one human reference shot in the entire release page, and that one doesn't focus on dialog at all.
For filmmakers I've talked to in the US, one of the biggest demands they have is v2v where they can carry over an actor's performance and insert it into whatever world they want. The movie market in China is somewhat different though - it's heavily oriented towards high action / high special effect movies. Consider this list of hollywood movies that have flopped in the US, but did great in China:
Warcraft (China: $225m box office, US: $47m)
Resident Evil (China: $159m, US: $27m)
xXx: Return of Xander Cage (China: $164m, US: $44m)
Pacific Rim: Uprising (China: $99m, US: $59m)
One read of this is that action movies / visual spectacles translate more universally than dialog-based movies. Another read is culturally China prefers that type of content in general. I suspect that for Bytedance, their focus is on action / special effects because thats where they see the demand.
I hear this a lot, but I observed this is no longer that simple.
Like Green Book (2018) made over US$70.7 million in China.
And one of the most popular films in 2021 was Hi, Mom (2021), made over US$785 million.
Obsession (2026) is another one that is doing amazing in China, absolutely beating high action films like Supergirl ($12 million and going vs less than $1 million)
I think it makes sense that historically American film focused on narrative based dramas and slow burn horror. Think of classic Hollywood hits such as Citizen Kane and Psycho. This is even apparent in the recent success of Obsession, which really follows strongly in the trend of Hitchcock, magical realism, human drama, psychological horror.
Compare this with historically successful Honk Kong films: Kung Fu Hustle, Police Story, Ip Man. I mean there are examples of Hong Kong films that are slow burn human dramas, like Chungking Express and Eat Drink Man Woman, and successful American action films like Diehard, but I think its clear that Hollywood film didn't start out with action movies, and Hong Kong film didn't start out with slow burn dramas, but adopted these genres afterwards; you could even claim that successful Hong Kong films like the Bruce Lee films of the 70s fueled the American appetite for action in the 80s, which was the most prominent era for that genre in the US.
As a corollary to this, it becomes clear that the Korean and Japanese film making industries are not as culturally distinct from Hollywood as Honk Kong film making is. For Korean film, the reason is obvious: Koreatown in Los Angeles is right next to Hollywood, in some respects is a part of Hollywood, and there are deep ties now between the Korean and American film making industries for that reason, and Korean-Americans have an overly high degree of representation in American media, and Korean productions often come to America to shoot (think of the end of Squid Games: it's almost certainly the case that they shot in LA because a producer has a cousin or something that works in Hollywood).
Im also leaning toward the per capita explanation as being a greater driver than audience vfx preference here. Through 2025 China was esti.ated to have about 800 million people in the middle class, with that growing to 1.2b by year end 2026. This is compared to the USAs approx 175m middle class.
If anything, those numbers for China seem quite low.
Whenever I see the new releases around video generation (and image) generation models, I get goosebumps, because it just feels so fun to work with them. But then I remember that I spend upwards of $10k on inference generating well over 50k images for storyboards, training models; and probably creating almost an hour of video (I assume). Yeah, I get that things can be economic if you don't use the latest models (ran some case studies on this), but the latest models are the most fun to work with. It doesn't scale as well as "vibe coding" stuff together on the weekend. And when things work really well its almost as if you're seeing an zoopraxiscope come to life for the first time; and you just want to keep going.
I got a few offers to work with some startups in this space, but it also seems that many startups work on stuff that just doesn't seem to be very worthwhile (like creating masses of spam for YT or TikTok shorts), or even straight out morally/ethically wrong (cloning/deepfakes, etc). But seeing advances in this space; and coming from a filmmakers background, I might just end up being naturally drawn to this space on an engineering level and figuring something out along the way. As you can see I worked on a lot of stuff just for the fun of it, and documenting the process: https://edwin.genego.io/blog (but I stopped at the beginning of the year .... might.. just pick it up again.
I went to Art Center for film. Loved it. But ended up writing software instead of shooting movies (while still also handling a lot of visual art direction, graphics work, UI, 3D animation, etc). Now I feel like we're starting to be roughly in the same boat as far as using prompts.
What bothers me is that every piece of content generated this way helps flood an already saturated market for content, while slowly degrading the expectations of what people see, to the point that no one will bother with shooting or animating anything anymore. Even if it's 50% worse, it's 90% cheaper, so the economics argur against producing any new physically made content. Simultaneously, it's cannibalizing all existing content. This points toward a feedback loop, like a snake eating its own tail. And even though Hollywood blockbusters have followed that pattern for a couple decades, it's demoralizing to me to see it enshrined as the future of film (or to hear from someone who makes films that it would be a preferred mode of creation).
MiniMax H3 is going to release weights. You can locally run it with definitely less than $10k (and possibly faster than Seedance's queue), and it's fun to train it for whatever you need.
Awesome! Have been a bit in the dark of the latest models, will have a look. I am due an upgrade for my local machine and GPU, so that excites me as well.
I've been dabbling in this space on the application layer and have spent a few hundred myself experimenting with video models.
It can be really entertaining/addicting to build with them because you're essentially pulling the slot machine and having TikTok/Marvel/YouTube come out of it. I think that was the bet with Sora but the problem is mostly that the novelty wears off quick, and most people want to just consume content without typing in what content they want to see, or sifting through mass-generated spam "content" with nothing behind it to make it worthwhile (a lot of people engage with content parasocially)
Once the tools for creators to steer and integrate models in this space get better, it will explode. We've been working on what I think will be one of the first use case for integrating these models, because I think we're approaching a middle ground where they can be integrated in experiences to provide entertainment/engagement/fun experiences without feeling like slop.
Btw, I'm impressed with some of the AI content on your site but I think you might want to pare down the non-demo pages because it has a different impression me (can I trust that this text is true? / I'm reading a lot of words but not really learning about this person) than you might have intended). I'm a bit of a hypocrite here but also speaking from experience.
Thanks for the input! And no I agree with your assessment of the website, I have some plans of a much simpler redesign soon, and I definitely value that critique. I think my "digital garden" has been through at least a few dozen of iterations, and sometimes I just get too carried away with it. The good part being, the next iteration always starts out better than the previous one (or at least I hope so :)).
The video quality is insane, but is there a model that doesn't make video where it looks like the characters are pausing at the end of their lines for a laugh track? There just seems to be an extra couple of beats after someone says something where they just stand dead still. What's the deal with that?
Video generation models generate videos of a predetermined duration, so if a character finishes a line but there's still seconds remaining, then the model still has to fill it in
Seedance 2.5 looks amazing, but MiniMax H3 is going to be open weights within 24 hours: https://fal.ai/minimax-h3. According to the ComfyUI team, it should even work acceptably on mid-range consumer GPUs like the 3080.
I'd honestly take the slight quality hit for more control and lower costs.
This seems extraordinarily good to me, compared to what I've seen before. Their washing machine advert example seems like it's as good as anything else on social media. I'm shocked by the quality and the coherence they're able to maintain, I assume it's really good at using those reference images they mention in their prompts. I think the only one that's noticeably bad is the concert hall, where the first few seconds show almost empty stalls and then toward the end it shows a full audience, and that audience also looks a bit off.
I don't think audio, image or video generation should exist. I haven't seen enough positive applications to justify the amount of harm these tools are being used to cause.
Now they cause harm because there's people that may believe a fake video is real. But people is developing skepticism about what they see in videos and learning to think that they may be generated, before taking them as true.
For everything that exists, there are people who think it shouldn’t exist. Fortunately, society doesn’t ban everything that disagrees with some people, and even more fortunately, we now live in a multipolar world where what is available doesn’t depend on the whims of a single pseudo-moralistic government.
If in the future they can create a reward model to aid with creating super human levels of appeal tailored to me, I would considered it one of the best technologies ever made. I'm not sentimental about the source, pretty videos and pictures make me happy
The sample video looks incredible and the ability to keep details consistent for so long is impressive but it _still_ looks screams of AI in every single shot. I can’t put my finger on why, something about the way the rooms are put together, the expressions of the faces, the movements… It’s very unsettling.
I think we as humans are just really good at being able to tell. CGI has been around for decades and yet we can still look at it and say "yeah that was CGI". Uncanny valley I guess?
It has improved a lot, but these demo reels still have all AI video issues. Flash cut salad (including the scenes that should have longer cuts), unnatural motion that looks animated, unprompted YouTube-face acting, etc. Admittedly it's all a lot less pronounced in this version.
What is much more interesting is how well it behaves off distribution (e.g. how far it can deviate from that movie/trailer aesthetics and still stay coherent)
You know what, in my previous life as a filmmaker I could've only dreamed of such a thing. Filmmaking is an art form which you cannot do alone. Outside of a lot of time and money, you need cooperation of a number of people and each day of production you end up accumulating a set of compromises to your vision. Your taste is what makes you tolerate that or not, and it's exhausting. It matters if you're the author since at the end of the day it's your name on it, not the crew (as much).
Now that we have these tools, I don't know but I am absolutely disinterested in it. Not that it doesn't feel right or anything, but I'm just not excited enough even to want to try to materialize some of my "big game" stuff. Closest thing I'd compare it with is like when you pirate a bunch of games and you play none as a result of abundance.
I think for me it's just missing the "essence". Making a film is hard, but that labor with multiple people who are passionate about the project is just as important as the end result.
I've been trying to generate concept art for a game idea. I figure, I can't draw so why not have something else bring my vision to life. After dozens of failed attempts and burned tokens I decided to just hire someone. The experience was night and day. They provided thoughtful details to help tell a story in the image. They got to tell me what they added and why, etc. This isn't even their project but their passion for it showed in the end result.
These models are impressive, but I think people watch movies because of the essence as much as the movie itself.
The thing is, it has the same essential property that you don't control it any more. The problem with these tools is the lack of fine grained control means there is no room for you to express your individual creative input. Your total input is a few sentences of a prompt, then the AI did all the creative part. If the creative input is what you enjoyed, it is actually not that much more (or even less) here than it was with traditional film.
Another aspect is that if you struggle and struggle and create something great, people can see the work you put in, and you'll visibly stand above the rest who don't have the same level of perseverance. But if anyone can create big studio quality visuals with a prompt, how do you stand out? "When everyone's super, no-one will be."
This is a matter of effort and direction and not an inherent flaw of the tools. Look at Apple's recent image tools to change perspective. Tools to improve artistry are improving.
I often see this kind of passive-aggressive characterization of the process with AI that's meant to bait people who obviously are expressing their creativity with these models.
I hope no one will fall for it: even when you emphatically reply "I spent the last year building the pipelines I use" it'll still be treated as equivalent a couple of sentences relative to the old ways.
-
At the end of the day, the painful lesson some people are (re)-learning is that creative expression doesn't have to be tied to any specific process. A lot of people fall in love with creative expression + process... but some fall in love with just the process, and some fall in love with creative regardless of the process.
AI does not favor those of us who loved the process (that's me with coding), but it's definitely capable of allowing new people to experience what it's like to have creative input in your mind and see it expressed outside of your mind.
I completely agree with you. The question is, are we headed to a situation where it is a new process that still allows the same expression, or is it fundamentally less able to facilitate that expression? we have to get beyond the one-shot style prompt shown here to something much more fine grained. The jury is still out for me on this, but I can believe we will get there. We just aren't anywhere close to it right now.
Was just going to say... Even if all AI gives me is the opportunity to bring my visions to life even somewhat terribly... That's a lot better than I'd be able to do without it. Which is for them to go nowhere. I'll take it.
The quality of AI videos is blowing my mind. I can still see things that seem a bit "off" but it's hard to distinguish AI videos from real videos. Even blockbuster movies are starting to be dwarfed by what AI can create. How long until we see the first full length AI movie hit the theaters? No actors, no development team, just a guy prompting AI...
People said this kind of thing about MiniDV cameras, and then iPhones. The reality is that technology hasn't been a barrier for filmmakers for a very long time. Writing is still the hard part.
Yes, it is awesome. Funny we got this level of capability now and we all take it in our stride. Growing up in the 80's I loved futuristic scifi, feels like I am in it now. And even I am 'just' taking it in my stride.
This will make longer (~30s) narrative add creation a lot better and more interesting at a reasonable price tag--roughly $7 best I can tell. Looking forward to trying it out.
it seems perfect for that - maintaining coherence over 30s-1min type window is achievable and so many ads are built on unrealistic premise to begin with.
To me, just reducing the cost alone seems secondary, far more important is putting the actual creative and marketing people directly in control of the output. Maybe the end result still gets sent to a pro studio for final production, but letting the true stakeholders directly create what they want could be a killer app. It may also be a terrible idea, like Homer Simpson's car - but that won't stop it being successful.
Where can you actually get access to these models that isn't an outright scam? All the sites that promised to have Seedance 2 turned out to be scams. Does anyone know how to actually use it, and is it available to run yourself?
Is it just me or are these video models only actual use case is misinformation and spam? Sure they show us quirky and whimsy samples on the release page, but does anyone really believe that?
I think that will remain the use case of these tools as long as it is still so easy to tell they don't represent reality very well. Maybe that's a good thing, maybe we're not ready for a world where any video longer than a brief clip is indistinguishable from reality.
Also, much of the promotion of these tools come from a population of people unhappy with the current state of the real content creation industry that use real footage and real people because AI movie promoters see traditional media as mainly propaganda machines anyway, so these tools sort help level the playing field.
There are thriving Ai video communities who are trying to replicate big budget productions with indie resources. Look up Gossip Goblin. There's a parallel explosion of memes at the same time, like Balenciaga Harry Potter
I've been having some similar thoughts lately about some application spaces being more harmful than others.
Coding, for example, seems somewhat benign, because code has to fulfil clear metrics: Either it works or it doesn't. Either it performs or it doesn't. As long as you are able to provide those metrics, you can to some extent treat it as a black box without losing much from a system view (lifecycle, long-term maintenance, keeping things working, broader qualified efficiency like re-use are of course other stories).
But on the other hand, delegating decision-making and thinking to models feels plain harmful to me. I have so many stories around me from office settings these days: "We have to pre-pone meeting XYZ, and we don't have enough time to prepare, so I made this AI analysis <power point deck>". This was then mis-prompted to fit a foregone conclusion, no one has time to do it or capacity to refute a 2000 word slop deck (except by slopping back), and crazy stuff becomes plan of record. This is really going to be a problem for organizations ...
That is an Indian English word, most non Indian English speakers are not familiar with it. reschedule or bring forward would be more appropriate for a global audience.
Only if you don’t think about it for more than two or three seconds. Then you remember deepfake revenge porn and elon musk style child sexual material.
I get entertainment value out of watching the videos other haves created. I hear of educators creating educational value from the videos they make as it allows them to make higher quality explanations.
It's just you. Any director, storyteller, or filmmaker can see the potential here to realize scenes that simply wouldn't be possible otherwise. But indeed there's plenty of slop coming out too.
https://x.com/CuiMao/status/2049828401201246395
This is extraordinary lifelike.
Obviously there are a lot of ways it shouldn't be used but I want to live in the future where it's something we use to have fun, where reasonable, rather than a dangerous/sketchy taboo
And then every generation after that will be born into increasingly a AI-video rich world.
I think most developers would agree that the Cursor vibe-code era sentiment towards AI for coding was right in the sense that it wasn't that useful then, and a lot of people did and do stupid things with it, but it really did deliver quite substantially on its promise just a couple years after the "consensus" was often "that's never going to work".
It's very strange to me how consistently public opinion shifts on this problem. In the long run we're all dead but whether it takes one or five years, you probably will be there to live through it. I assume whoever is downvoting you is just thinking with their consumer-brain or "I don't want that" (the way it is now) rather than that this will be no different than video games, or wasn't a child who watched youtubepoop or other weird stuff because it was funny.
For filmmakers I've talked to in the US, one of the biggest demands they have is v2v where they can carry over an actor's performance and insert it into whatever world they want. The movie market in China is somewhat different though - it's heavily oriented towards high action / high special effect movies. Consider this list of hollywood movies that have flopped in the US, but did great in China:
Warcraft (China: $225m box office, US: $47m)
Resident Evil (China: $159m, US: $27m)
xXx: Return of Xander Cage (China: $164m, US: $44m)
Pacific Rim: Uprising (China: $99m, US: $59m)
One read of this is that action movies / visual spectacles translate more universally than dialog-based movies. Another read is culturally China prefers that type of content in general. I suspect that for Bytedance, their focus is on action / special effects because thats where they see the demand.
Like Green Book (2018) made over US$70.7 million in China.
And one of the most popular films in 2021 was Hi, Mom (2021), made over US$785 million.
Obsession (2026) is another one that is doing amazing in China, absolutely beating high action films like Supergirl ($12 million and going vs less than $1 million)
But you're probably just blundering with value comparison since just per capita, your numbers arn't that' extreme, still the point stands.
Compare this with historically successful Honk Kong films: Kung Fu Hustle, Police Story, Ip Man. I mean there are examples of Hong Kong films that are slow burn human dramas, like Chungking Express and Eat Drink Man Woman, and successful American action films like Diehard, but I think its clear that Hollywood film didn't start out with action movies, and Hong Kong film didn't start out with slow burn dramas, but adopted these genres afterwards; you could even claim that successful Hong Kong films like the Bruce Lee films of the 70s fueled the American appetite for action in the 80s, which was the most prominent era for that genre in the US.
As a corollary to this, it becomes clear that the Korean and Japanese film making industries are not as culturally distinct from Hollywood as Honk Kong film making is. For Korean film, the reason is obvious: Koreatown in Los Angeles is right next to Hollywood, in some respects is a part of Hollywood, and there are deep ties now between the Korean and American film making industries for that reason, and Korean-Americans have an overly high degree of representation in American media, and Korean productions often come to America to shoot (think of the end of Squid Games: it's almost certainly the case that they shot in LA because a producer has a cousin or something that works in Hollywood).
If anything, those numbers for China seem quite low.
I got a few offers to work with some startups in this space, but it also seems that many startups work on stuff that just doesn't seem to be very worthwhile (like creating masses of spam for YT or TikTok shorts), or even straight out morally/ethically wrong (cloning/deepfakes, etc). But seeing advances in this space; and coming from a filmmakers background, I might just end up being naturally drawn to this space on an engineering level and figuring something out along the way. As you can see I worked on a lot of stuff just for the fun of it, and documenting the process: https://edwin.genego.io/blog (but I stopped at the beginning of the year .... might.. just pick it up again.
What bothers me is that every piece of content generated this way helps flood an already saturated market for content, while slowly degrading the expectations of what people see, to the point that no one will bother with shooting or animating anything anymore. Even if it's 50% worse, it's 90% cheaper, so the economics argur against producing any new physically made content. Simultaneously, it's cannibalizing all existing content. This points toward a feedback loop, like a snake eating its own tail. And even though Hollywood blockbusters have followed that pattern for a couple decades, it's demoralizing to me to see it enshrined as the future of film (or to hear from someone who makes films that it would be a preferred mode of creation).
It can be really entertaining/addicting to build with them because you're essentially pulling the slot machine and having TikTok/Marvel/YouTube come out of it. I think that was the bet with Sora but the problem is mostly that the novelty wears off quick, and most people want to just consume content without typing in what content they want to see, or sifting through mass-generated spam "content" with nothing behind it to make it worthwhile (a lot of people engage with content parasocially)
Once the tools for creators to steer and integrate models in this space get better, it will explode. We've been working on what I think will be one of the first use case for integrating these models, because I think we're approaching a middle ground where they can be integrated in experiences to provide entertainment/engagement/fun experiences without feeling like slop.
Btw, I'm impressed with some of the AI content on your site but I think you might want to pare down the non-demo pages because it has a different impression me (can I trust that this text is true? / I'm reading a lot of words but not really learning about this person) than you might have intended). I'm a bit of a hypocrite here but also speaking from experience.
I'd honestly take the slight quality hit for more control and lower costs.
What kind of harms do you have in mind?
Think about how many completely fake videos are going to be flying around for the next election. Nobody is going to believe anything the see online.
What is the the next stage, where the only things you believe are things you see with your own eyes.
If in the future they can create a reward model to aid with creating super human levels of appeal tailored to me, I would considered it one of the best technologies ever made. I'm not sentimental about the source, pretty videos and pictures make me happy
What is much more interesting is how well it behaves off distribution (e.g. how far it can deviate from that movie/trailer aesthetics and still stay coherent)
Now that we have these tools, I don't know but I am absolutely disinterested in it. Not that it doesn't feel right or anything, but I'm just not excited enough even to want to try to materialize some of my "big game" stuff. Closest thing I'd compare it with is like when you pirate a bunch of games and you play none as a result of abundance.
Weird-ass times.
I've been trying to generate concept art for a game idea. I figure, I can't draw so why not have something else bring my vision to life. After dozens of failed attempts and burned tokens I decided to just hire someone. The experience was night and day. They provided thoughtful details to help tell a story in the image. They got to tell me what they added and why, etc. This isn't even their project but their passion for it showed in the end result.
These models are impressive, but I think people watch movies because of the essence as much as the movie itself.
I hope no one will fall for it: even when you emphatically reply "I spent the last year building the pipelines I use" it'll still be treated as equivalent a couple of sentences relative to the old ways.
-
At the end of the day, the painful lesson some people are (re)-learning is that creative expression doesn't have to be tied to any specific process. A lot of people fall in love with creative expression + process... but some fall in love with just the process, and some fall in love with creative regardless of the process.
AI does not favor those of us who loved the process (that's me with coding), but it's definitely capable of allowing new people to experience what it's like to have creative input in your mind and see it expressed outside of your mind.
This feels like the serious inflection point for high quality full length feature film productions using this tech.
To me, just reducing the cost alone seems secondary, far more important is putting the actual creative and marketing people directly in control of the output. Maybe the end result still gets sent to a pro studio for final production, but letting the true stakeholders directly create what they want could be a killer app. It may also be a terrible idea, like Homer Simpson's car - but that won't stop it being successful.
IIRC they give Seedance 2.5 yesterday. Byteplus API is the official way to access it as an API but I am not sure if they serve 2.5 yet.
Also, much of the promotion of these tools come from a population of people unhappy with the current state of the real content creation industry that use real footage and real people because AI movie promoters see traditional media as mainly propaganda machines anyway, so these tools sort help level the playing field.
Coding, for example, seems somewhat benign, because code has to fulfil clear metrics: Either it works or it doesn't. Either it performs or it doesn't. As long as you are able to provide those metrics, you can to some extent treat it as a black box without losing much from a system view (lifecycle, long-term maintenance, keeping things working, broader qualified efficiency like re-use are of course other stories).
But on the other hand, delegating decision-making and thinking to models feels plain harmful to me. I have so many stories around me from office settings these days: "We have to pre-pone meeting XYZ, and we don't have enough time to prepare, so I made this AI analysis <power point deck>". This was then mis-prompted to fit a foregone conclusion, no one has time to do it or capacity to refute a 2000 word slop deck (except by slopping back), and crazy stuff becomes plan of record. This is really going to be a problem for organizations ...
That is an Indian English word, most non Indian English speakers are not familiar with it. reschedule or bring forward would be more appropriate for a global audience.