What the world's most mature AI-entertainment market teaches us about engagement and unit economics.
Short dramas may finally be taking off in the U.S. The reason has less to do with a sudden interest among American consumers than with the supply side: AI models have gotten good enough that it's now much cheaper to produce content where a character stays the same person from shot to shot, and where the dubbing lip-syncs into any language. As Coatue sector head @lordbartonjr put it in a June video, China could be the early signal for what reaches the US next.
But short drama is only one slice of China's AI entertainment market, which runs about a year ahead on go-to-market: how these products get packaged, priced, and sold. That lead holds even where China is supposedly behind on the underlying models. Though after GLM-5.2, Seedance 2.5, and Kimi K3, whether it's behind at all is now an open argument on X. There's a lot here for anyone in AI consumer, where investment interest is clearly picking back up, because the way China's market developed exposes a real tension between engagement, cost infrastructure, and monetization.
The models will keep getting better, so what you compete on is no longer production quality or volume. Everyone will have that soon. But consumers still only have 24 hours a day. That's where the problem starts:
The scarcest thing now isn't content. It's attention.
The moat isn't acquiring users; it's getting them to stay, come back, and keep engaging.
But every extra minute of engagement can mean another bill from your cloud and model providers. So how do you price it?
Hold those three in mind as we go.
The four main categories and each of their economics
China has plenty of AI-entertainment niches, but these four are where sustained investor interest and real user attention sit, even though each is at a very different stage. The clearest way to see them is on two axes:
- Is the user leaning back or leaning forward? Watching content (lean-back) vs. interacting with it (lean-forward).
- Is the AI a production tool or the live experience? Does it make the content up front, or does it run on every interaction?
These axes aren't just descriptive. They set your cost structure. Lean-back plus tool means you generate it once and serve it forever at almost no marginal cost. Lean-forward plus native means every interaction is a fresh, paid inference call. That split is what this piece is about.
For each of the four, I'll cover the same four things: what it is, how it makes money and where the compute lands, how far along it is, and where it's heading.
1. AI short drama — lean-back, AI-as-tool
What it is. Vertical, phone-native micro soap operas. Episodes run 60 to 90 seconds, nearly every one ends on a cliffhanger, and a coin paywall lands right when you're hooked. People watch them the way they doomscroll: one more unlock, one more episode. The two big overseas apps, ReelShort and DramaBox, are both Chinese-built, with Kuaishou (Kling) and ByteDance (Seedance) supplying the video generation.
This only became AI-viable in 2026, and the blocker was never resolution. The picture had looked fine for a while. The blocker was consistency: eighty connected episodes starring a character who still looks like herself in episode 40, same face, same clothes, same voice from shot to shot, no strange shimmer between cuts, and who still looks like herself after her dialogue is re-dubbed into Spanish, Portuguese, and Thai.
How it makes money and where the compute lands. This is the cheap, safe corner of the map. AI does its work up front; once an episode exists, every additional view costs you basically nothing. So the math is old-fashioned media math: user-acquisition cost + production cost < what that viewer pays in unlocks and ads. AI's only job here is to make production faster and cheaper — and to localize, by dubbing, lip-syncing, and recutting one hit for ten markets. Marginal cost per extra viewer ≈ zero, which is why this is the one category with a proven path to profit.
How far along it is. The most mature and competitive of the four. Saturated at home; the real growth is overseas, where no vertical-drama industry existed. Outside China, in-app spend on short drama hit roughly $2.98B in 2025 (+115% YoY), and it's mainstream enough that SAG-AFTRA signed a vertical-drama agreement and Disney ran DramaBox through its accelerator. Paramount+ also announced last week that it will experiment with vertical short dramas. The format is proven.
Where it's heading. More overseas expansion. More genres that used to be too expensive, like mythology and sci-fi, now that AI renders them cheaply. And a slow shift from live action with AI assist toward more fully AI-generated episodes as the video models improve. But the ecosystem is increasingly controlled by the two ends, the video models and the distribution apps, because what the middle sold was capacity and execution: shooting, editing hours, dubbing, recutting for ten markets. That's exactly what the models absorb, which leaves less and less margin for the studios and vendors in between.
@deedydas had some great posts on short dramas.
2. AI mini-games — lean-forward, AI-as-tool
What it is. Lightweight, casual, shareable games where AI both lowers the bar to build a game and make it feed-distributable. Feed-native interactive content you can spin up fast. Chinese names here include Aippy, Rezona, and Loopit. Meta's vibe-coding game Pocket is the closest US analog.
How it makes money and where the compute lands. Same cheap corner as short drama, one column over. The AI is spent at build time. At runtime the game largely runs itself, so serving stays cheap. It monetizes like a store or a platform: selling games, taking a cut, running ads. The economics are fine. The open question is the product.
How far along it is. The earliest and least proven of the tool-based plays. AI clearly lets more developers ship more games, but a genuinely new game loop that couldn't exist before AI hasn't shown up. Capital is watching more than betting.
Where it's heading. Either it stays a tooling layer that makes ordinary games cheaper to produce, or someone finds the native loop: AI-driven characters and generated content that make a game you couldn't build any other way. Until that shows up, it's the most speculative of the four.
3. AI companionship — lean-forward, AI-native
What it is. Apps built around an AI character you talk to, for comfort, flirtation, role-play, or company. You're not watching. You're in an ongoing relationship with software that remembers you. The Chinese leaders are MiniMax's Xingye (星野) and its overseas twin Talkie, plus Tipsy and Emochi. This was the first category to prove people would pay for AI entertainment. Text was the first model capability to get genuinely good, so companionship shipped before anyone needed high-quality video.
How it makes money and why it hit a wall first. This is where the tension from the top of this piece shows up, and the most interesting category to watch. The AI isn't a tool that made the content. The AI is the content, running live on every message. Every interaction is a fresh inference call you pay for, so compute stops being a one-time production cost and becomes a cost of goods that rises with engagement. Then stack the freemium reality on top:
- These are mostly free apps with a thin slice of ~$10/month subscribers — and the heaviest users, the ones generating the most compute, are disproportionately the free ones. Your biggest cost center pays you nothing.
- A flat subscription caps your revenue but not the user's usage, so a true power user can burn more compute than they pay.
- Metering usage (the one obvious fix) is the one lever you can't pull, because nobody wants a meter running while they confide in a companion. It kills the intimacy that was the product.
- And making the product better (longer memory, richer context) makes each call cost more.
So the thing that proves the product works, deep daily engagement, is the same thing that runs up the bill. That's why companionship, even while clearly winning on demand, was the first category to hit a hard economic wall. The fix isn't no metering. It's invisible metering: packaging the heavier interactions as a value offer instead of a running counter. Sell the quest, the outfit, the premium character. Charge for the thing the user wants, not for the tokens it took to make it.
That only fixes willingness to pay, not the free power user. That one is a routing problem: cheap models for small talk and the expensive ones for the moments that matter, which is why the companion apps that survive compete on cost engineering as much as on character writing.
How far along it is. Demand is fully proven: these apps see the highest engagement of any category, often over two hours a day. MiniMax's AI-native products (led by Talkie and its domestic twin Xingye) averaged around 27.6M monthly actives in the first nine months of 2025, up from 19.1M in 2024, and consumer apps drove roughly 71% of the company's revenue over that period, most of it from overseas. Domestic content rules constrain exactly the part of companionship that monetizes best, so the same team runs two products: a compliant version at home and a fuller one abroad, with the revenue concentrated in the latter. But the field is crowded and consolidating, and plenty of look-alikes are quietly dying on exactly the economics above. Proven, but in a shakeout.
Where it's heading. The wall lifts as inference gets cheaper, through smaller fine-tuned models and on-device processing, and as teams find monetization that adds value instead of metering it. The winners will be the few who can balance model quality against compute cost. The rest get absorbed. Push that curve further and you arrive at the frontier.
4. Interactive video world, generated live — the deep end of AI-native
What it is. An ambitious middle ground between a movie and a game. The plot branches around you, and characters answer in free-form dialogue generated on the fly. You're not picking from three options someone wrote in advance. You say what you want and the story absorbs it. The clearest current example is 万千, a Chinese otome game (romance, aimed at female players) on PixVerse's real-time world model. It runs on a setup popular in Chinese web fiction: the heroine keeps world-hopping — a new setting and a new love interest each time — but never resets. She's one continuous character the whole way through, carrying her memories and choices from each world into the next, and the story is meant to remember all of it. That's what lands it at the far end of the map: holding one evolving character coherent across hours and across totally different generated worlds is the coherence problem in its rawest form.
Three people built the demo in three weeks. A traditional 3D otome character takes four months just to model.
How it makes money and where the compute lands. Same cost structure as companionship, further along the same curve. A short drama gets generated fifty times offline and you keep the best; an interactive story has to land live, in front of the player, on the first try, because a genuinely open story branches into too many paths to pre-build. Every turn is a paid inference call that has to work immediately and stay coherent across a long session. Monetization is subscription or pay-per-play, but the compute exposure is the steepest anywhere on this map.
How far along it is. The least shipped thing on this map. Everyone points to it as the long-term prize, and almost nobody is running one at scale. Current attempts hedge by pre-generating art, locations, and set-piece scenes offline, and spending live inference only on dialogue and orchestration. That's the one architecture that works today, and it's also an admission that the fully live version isn't affordable yet. Capital is positioning early rather than deploying big.
Where it's heading. This is what most people mean by AI-native entertainment, and why so many treat it as the endgame. It does completely what the other categories do partway: generates the whole experience live and bends it around you, a story genuinely different for every player, with the pull of a game, the immersion of a film, and the intimacy of a companion in one. No earlier medium could deliver that, which is what gives it the highest ceiling on this map: the deepest engagement, the strongest reason to pay, the longest lifespan.
It's also why it hasn't arrived. The payoff depends on two curves crossing: models that stay coherent across hours of branching, unpredictable interaction, and a token cost low enough to generate all of it live. Neither is there yet. So it's the clearest picture of where the field is going and the furthest from being able to get there.
The map is a trajectory, not a snapshot
Putting the four categories on a grid is useful, but the real signal is which way they're moving. The whole field is drifting from AI as a cheap production tool toward AI as the live experience, pulled by the same two curves: models getting better at long, coherent interaction, and compute getting cheaper.
That drift is what makes the pricing question from the top of this piece unavoidable. As more of the field slides from make-once to generate-live, more of it inherits the cost structure where engagement is a bill, not just a metric.
So the real question for anyone building here isn't which category to pick. It's at what level of engagement your unit economics break, and whether people will pay enough, fast enough, to get you past that point before they do.
Other takeaways for U.S. peers in AI entertainment
A few things operators in China keep running into that tend to surprise people coming from Tech.
"Infinite personalization" is a trap. "AI will let every user generate their own perfectly personalized story" only sounds good. For mainstream entertainment it mostly fails. Most people aren't writers and don't want to be. Hand them a blank canvas and full freedom and what comes out is a shapeless story with no tension and nothing worth remembering. The products that work keep a human-authored backbone: a set worldview, real stakes, deliberate emotional beats. Strong structure with a little freedom beats total freedom, which just gets boring. That's not an argument against the interactive frontier. It's a spec for it. The win isn't a blank canvas. It's a strong authored world that bends.
Content is about to be infinite, so the scarcity flips. AI raises content output a hundred or a thousand times over, and the market will only get more flooded. But people's time is fixed. The future won't be short on content, only on good content that carries real emotional value. Which means the winners are never the teams with the strongest AI muscle. They're the teams that understand users, can tell a story, can build IP series, and can deliver that emotional payoff consistently.
Users are paying for a feeling, not for content. Old entertainment logic was simple: games were about winning, shows were about getting lost in them. AI fused winning, immersion, and interaction into one personal, emotionally tuned experience, so what people pay for is a moment of resonance, the sense that this thing gets me. What AI entertainment sells was never content. It's small, personal hits of emotional value. And that's the answer to the pricing question this piece opened with: emotional value is what lets you charge above the marginal cost of compute. Margin comes from meaning, not from tokens.
