English translation of the AWS Technical Blog article published September 29, 2026.
Introduction
To generate choreography from a song, MVNT had to solve two problems: turn dance footage into training data and serve GPU-heavy inference at scale. Dance combines musical timing, body mechanics, genre conventions, and movement through 3D space. A general-purpose model alone cannot reliably capture all of those relationships.
MVNT builds AI dance generation tools for this domain. Its team brings together cofounder and choreographer Young-jun Choi with engineers experienced in motion generation, 3D, computer vision, and humanoid robotics. MVNT offers its technology through a web app, an API, and an Unreal Engine plugin for choreographers, K-pop agencies, AI user-generated content creators, game developers, and virtual production studios.
In this post, we describe how MVNT built a data pipeline and inference serving architecture on AWS. An event-driven workflow converts unstructured dance videos into training data, asynchronous GPU workers handle intensive analysis, and a tiered serving architecture supports concurrent generation requests. The design shows how a small team can move from a proof of concept toward a service used by global customers.
The challenge: Capturing choreography as data
The K-pop industry produces approximately 1,500 choreographies each year. Demand grows further when J-pop, Latin music, Bollywood, pop, and musical theater are included. Dance challenges have become a prominent part of music promotion on social media, while dance items are a significant revenue source for games. Yet creating choreography that resonates with an audience remains difficult, even for professional choreographers.
MVNT began with a question: Could someone provide a song and generate choreography that fits it? Music generation tools have made it easier to create a track, but generating dance presents a different set of challenges. Choreography is more than a sequence of poses. It must account for musical structure, rhythm, body mechanics, genre, interpretation, and movement through three-dimensional space and time.
Training a dance model therefore takes more than collecting videos. MVNT needed high-quality 3D motion capture data, motion extracted at scale through human mesh recovery (HMR), and annotations that express a choreographer's intent in a form a model can learn. It also needed a way to process those datasets and serve the trained model reliably to users and partners worldwide.
This article focuses on the two parts of that system: the data pipeline and the inference serving architecture.
Data pipeline: Turning dance into training data
Over the past five years, MVNT has worked with choreographers to develop workflows for collecting and analyzing dance. A pose tells you where a dancer is. Choreography also depends on timing, energy, which body parts drive a movement, and how that movement should feel. MVNT's data workflow captures that information alongside the movement itself.
Two sources of motion data
Motion capture data
Dance motion analyzed from video
MVNT uses both motion capture and video analysis to collect joint and 3D motion data. This hybrid approach balances quality with scale.
- Quality: MVNT builds its own motion capture dataset to obtain dance data suitable for game engines and visual effects. In those applications, finger movement, a performer's precise motion, and the intent behind it matter. Its motion capture equipment records movement with less than 1 mm of error.
- Scale: Dance videos capture real performers, camera angles, and performance context, and provide far more material than motion capture alone. Extracted motion may lose some of the nuance of captured performances and can contain artifacts such as floating feet. It therefore needs additional processing before it is suitable for training.
Labeling the data
Motion data becomes more useful when it includes context: the intent behind a phrase, its genre, its intensity, and its relationship to the music. Descriptions from dancers and more detailed genre labels help the model learn distinctions that joint coordinates alone cannot express.
On its in-house labeling platform, MVNT works with dancers to build a dance ontology that includes technical descriptions as well as movement texture and emotion. The following abbreviated annotation keeps the source video title in Korean and translates the movement instruction into English:
{
"videoTitle": "너와의 모든 지금 (Every Moment With You)",
"choreographer": "Young-jun Choi",
"annotations": [
{
"name": "description",
"label": "Swing both arms up and to the right at a 45-degree angle. Keep the right arm fully extended and bring the left hand to the solar plexus...",
"start": 33.06,
"end": 35.43
},
{
"name": "emotion",
"label": "playful",
"start": 33.06,
"end": 35.43
},
{
"name": "texture",
"label": "flowing, wavy",
"start": 33.06,
"end": 35.43
}
]
} For the 33.06–35.43-second segment, the annotation pairs the movement instruction with emotion and texture labels.
Converting unstructured video into model-ready data is both repetitive and computationally expensive. The pipeline analyzes human movement frame by frame, reconstructs 3D structure from 2D poses, tracks multiple people, and organizes the results as time-series data. Processing a large video collection this way is difficult to manage in a local environment.
Processing dance videos at scale on AWS
Researchers first authenticate through the application. Each upload then starts a job: Amazon Simple Storage Service (Amazon S3) stores the video, Amazon Simple Queue Service (Amazon SQS) buffers the processing request, and GPU-backed Amazon EC2 workers analyze it asynchronously.
- Access control and routing. When MVNT researchers use the labeling tool, AWS WAF helps protect the application from common web attacks and excessive requests. Amazon Cognito handles authentication and access control. Authenticated requests pass through Amazon API Gateway to the backend, where AWS Lambda routes data retrieval, uploads, and job creation according to the request type.
- Ingestion, triggering, and buffering. Amazon S3 stores the source dance videos, images, and other files to be analyzed. A new upload triggers downstream processing, and Amazon SQS queues the job. Processing time varies with video size and duration, so handling every upload synchronously would be impractical. The queue absorbs upload spikes and lets workers process pending jobs at a sustainable rate.
- Asynchronous GPU inference and failure handling. MVNT's video-based HMR model needs GPUs, and longer videos take more time to process. GPU-backed Amazon EC2 workers pull jobs from SQS and run inference asynchronously. Failures can arise from corrupt files, excessively long videos, missing people in a frame, or resource limits. Failure events pass through Amazon SNS and AWS Lambda and are recorded in Amazon DynamoDB. Operators can inspect the records, split a video, adjust preprocessing parameters, or exclude a sample from the training dataset.
The result is structured, reviewable training data rather than a collection of media files. S3 retains the source videos. The analysis pipeline produces pose, motion, mesh, and visualization outputs, while DynamoDB records each sample's status and metadata. The data team reviews those outputs, adds any needed labels, and incorporates approved samples into the dataset for the music-conditioned choreography model.
Serving architecture: From proof of concept to production
MVNT's initial proof of concept ran the API server and its diffusion-based model on a single EC2 instance. That arrangement helped a small team iterate quickly, but it made the service vulnerable to traffic spikes, GPU contention, cascading failures, and deployment risk as usage and external API integrations grew.
MVNT moved to a tiered architecture on AWS that separates API handling from model inference and allows each tier to scale independently.
Within one Amazon VPC, MVNT separated the system into web, application, and database tiers:
| Tier | Configuration | Responsibility |
|---|---|---|
| Web | CPU-based Amazon EC2 instances in public subnets | Receive and authenticate requests; handle routing and lightweight API operations |
| Application | GPU-based Amazon EC2 Auto Scaling group in private subnets | Run the diffusion model and inference workers |
| Database | Amazon Aurora primary database and read replicas | Store job status, user requests, result metadata, and usage information |
- External traffic enters through an internet gateway and Elastic Load Balancing before reaching the web tier. An internal load balancer routes requests from the web tier to GPU servers in the application tier.
- The application tier runs GPU-based EC2 instances in an Auto Scaling group. MVNT can adjust capacity to match traffic patterns and distribute requests to other instances when an instance or Availability Zone becomes unavailable.
- The database tier uses Aurora with a primary database and read replicas to distribute read traffic and improve data-tier availability.
Results: Concurrent requests and faster generation
Separating the web and GPU tiers lets MVNT scale generation capacity without scaling the API tier in lockstep. At Unreal Fest Chicago in June 2026, approximately 70 participants used its dance AI plugin in a hands-on workshop. The system load-balanced 70 generation requests submitted around the same time.
The August 2026 Unreal Fest Seoul event brought a larger audience: around 50 people attended on site, while roughly 2,000 watched online. Generation requests arrived in a burst. Inference optimization, Auto Scaling, and load balancing helped absorb the load and reduce response delays for on-site users.
Moving inference from g5 to g7e.2xlarge GPU instances improved generation speed. Before the change, a generation took about 3–5 minutes. On g7e.2xlarge, MVNT recorded approximately 35 seconds for 20 seconds of audio, 62 seconds for 30 seconds, and 98 seconds for 40 seconds.
New users received their first result in about 51 seconds after registering, down from roughly 2 minutes—a reduction of approximately 57.5%. The generation completion rate also rose by about 10% compared with July.
Unreal Fest Seoul
Conclusion and next steps
The quality of a music-conditioned choreography model depends on the quality of its training data. MVNT's system connects a labeling pipeline for choreographers' domain knowledge, processing infrastructure for motion data, and an inference architecture that serves the model to users worldwide.
AWS provides the foundation for a small team to scale those parts independently. Amazon S3, AWS Lambda, and Amazon DynamoDB support the pipeline that turns dance videos into structured data. Elastic Load Balancing, Amazon EC2 Auto Scaling, and Amazon Aurora support model serving for users and partners in different regions.
The serving architecture also supported reported growth in August 2026: roughly 3× in new registrations and 2× in generation requests. MVNT plans to train models across more genres, musical structures, and movement styles, making dance creation and use easier in game, entertainment, and creator workflows. A mobile app is also planned to support growing traffic.
Although MVNT works specifically with dance, the architectural pattern applies to other specialized generative AI services: build a domain-specific data pipeline, separate GPU-intensive processing from request handling, and scale the serving tiers independently. Managed AWS services let a small team build those capabilities incrementally.
You can try MVNT's AI dance generation in mvntSTUDIO and find the Unreal Engine plugin on Fab. For similar workloads, the AWS guidance on asynchronous processing with API Gateway and SQS and its multi-tier architecture overview are useful starting points.
About the authors
Seiok Kim, CTO Seiok brings an engineer’s eye and a b-boy’s instincts to MVNT. He leads dance generation research and the infrastructure behind it. He’s also exploring what it would take to teach a humanoid robot to break.
Eunhee Kim, AI Research Engineer Eunhee develops motion generation methods that look convincing and hold up to physics. She’s also interested in the engineering behind the movement: data pipelines and inference optimization that make models practical to train and deploy.
Joon Jung, CEO A professional dancer-turned-founder, Joon is interested in what happens when art and technology share a stage. He’s goal is simple: make people dance more with MVNT, just like Duolingo!
Minwook Yoon, Solutions Architect at AWS Minwook helps startups turn cloud architecture diagrams into secure, efficient, cost-conscious systems. He provides technical guidance, shares AWS best practices, and leads hands-on workshops.