Make Hotel LobbyPricing

Hotel Lobby AI Prompt

Copy, fill in the [brackets] and paste into your AI video tool with one photo of each person.

Classic

The closest to the trend: relaxed, confident, taking turns.

A 10-second vertical 9:16 video, filmed as one continuous take. Two people perform a rap duet side by side inside a small recording booth. The booth walls are bright, glossy orange and the light is soft and even, with no harsh shadows. One studio microphone hangs on a cable between them, level with their mouths.

Person A stands on the left: the person from reference image 1 ([short description, e.g. curly black hair, red hoodie]). Person B stands on the right: the person from reference image 2 ([short description, e.g. short blond hair, white T-shirt]). Each person keeps their own face, hairstyle and clothing from their reference image for the entire video.

Performance: they trade lines. Person A leans in to the microphone and raps first while Person B nods to the beat and reacts. Then Person B takes a turn at the microphone while Person A steps back slightly and hypes them up. For the last line they rap together, both close to the microphone.

Camera: locked-off tripod shot at chest height, facing them straight on. Medium shot that keeps both people visible from the waist up the whole time. The camera does not move.

No AI video tool?Upload one photo of each person and we make the booth video for you.

Skip prompting — make it from two photos

New to the trend? Read the step-by-step guide.

Before you paste

  • Use a tool that takes reference images

    Image-to-video or reference-to-video tools work best. Upload Person A's photo first and Person B's second, and rename “reference image 1/2” to your tool's syntax if it has one.

  • Only one start image allowed?

    Put both people side by side in one frame first (any photo editor works), then describe them as “the person on the left” and “the person on the right”.

  • Fill in the [brackets]

    A few visible details per person — hair, top, glasses — give the model a second way to tell them apart.

  • Set 9:16 and about 10 seconds

    Match your tool's aspect-ratio and duration settings to the prompt. Longer clips drift more.

Prompt variations

Same booth, people and camera as the Classic prompt above — only the performance changes.

More Energy

Bigger gestures and bounce. Energy comes from the people, not the camera.

A 10-second vertical 9:16 video, filmed as one continuous take. Two people perform a rap duet side by side inside a small recording booth. The booth walls are bright, glossy orange and the light is soft and even, with no harsh shadows. One studio microphone hangs on a cable between them, level with their mouths.

Person A stands on the left: the person from reference image 1 ([short description, e.g. curly black hair, red hoodie]). Person B stands on the right: the person from reference image 2 ([short description, e.g. short blond hair, white T-shirt]). Each person keeps their own face, hairstyle and clothing from their reference image for the entire video.

Performance: high energy. Person A raps the first lines with big hand gestures, bouncing on the beat, while Person B shouts ad-libs and points at them. Then they switch: Person B raps at the microphone with sharp head nods and hand chops while Person A bounces and cheers. They finish shoulder to shoulder, both rapping into the microphone.

Camera: locked-off tripod shot at chest height, facing them straight on. Medium shot that keeps both people visible from the waist up the whole time. The camera does not move.

Playful

For friends, siblings and couples who don't take it seriously.

A 10-second vertical 9:16 video, filmed as one continuous take. Two people perform a rap duet side by side inside a small recording booth. The booth walls are bright, glossy orange and the light is soft and even, with no harsh shadows. One studio microphone hangs on a cable between them, level with their mouths.

Person A stands on the left: the person from reference image 1 ([short description, e.g. curly black hair, red hoodie]). Person B stands on the right: the person from reference image 2 ([short description, e.g. short blond hair, white T-shirt]). Each person keeps their own face, hairstyle and clothing from their reference image for the entire video.

Performance: playful and teasing. Person A starts rapping very seriously, then cracks a smile. Person B pretends to be unimpressed, then can't stop grinning and takes the microphone for their turn. They trade exaggerated reactions to each other's lines and give a light shoulder bump. Both laugh on the last line.

Camera: locked-off tripod shot at chest height, facing them straight on. Medium shot that keeps both people visible from the waist up the whole time. The camera does not move.

Fix identity problems

If faces look like a mix of both people, someone's outfit changes or they swap sides, add this paragraph to the end of your prompt.

Identity-Safe add-on

Paste at the end of any version above if faces drift, blend or swap sides.

Identity rules: Person A is only the person in reference image 1 and always stays on the left. Person B is only the person in reference image 2 and always stays on the right. Never swap their positions. Do not blend, average or mix their facial features. Keep each person's skin tone, face shape, hairstyle, facial hair, glasses, jewelry and clothing exactly as in their reference image. Both faces stay clearly visible and mostly turned toward the camera.

What to avoid

No cuts, no zooms, no face swaps, no identity blending, no extra people, no text or logos, no clothing changes. If your tool has a negative-prompt field, paste this there without the word “Avoid”.

Avoid list

Paste into your tool's negative prompt field, or at the end of the prompt.

Avoid: cuts, scene changes, zooms, pans, handheld camera shake, close-ups of one person, face swaps, identity blending, extra people, anyone in the background, text, captions, subtitles, logos, watermarks, clothing changes, hairstyle changes, extra microphones, instruments, distorted hands.

Why the prompts are written this way

Keep Person A and Person B distinct

Name each person once, then tie that name to three things: their reference image, their side of the booth and a short visible description. Use exactly the same name every time after that. Photos with different outfits help too.

Preserve clothing and identity

Many models quietly restyle people to fit the scene. Saying each person keeps their face, hair and clothing “for the entire video” pushes back on that, and waist-up reference photos show the model the outfit it's supposed to keep. If faces still drift, add the Identity-Safe paragraph and use the Classic version — less movement means less drift.

Get an alternating performance, not random motion

Write the performance as a sequence: who raps first, what the other person does meanwhile, who goes next, how it ends. Without an order, models tend to animate both mouths at once or add movement that has nothing to do with rapping.

Keep the camera fixed

The trend is one locked shot. Say it plainly (“locked-off tripod”, “the camera does not move”) and list zooms, pans and cuts in the avoid list. If you want more energy, put it in the performers' movement, never in the camera.

Keep waist-up, medium framing

Name the shot size and say it lasts “the whole time”. In a vertical frame two people side by side get small quickly; waist-up keeps both faces large enough to recognize and leaves room for hand gestures. Waist-up reference photos make this easier.

Why “make them rap” prompts are unstable

A one-line prompt leaves the model to guess the setting, who is who, who raps when and what the camera does — and it guesses differently every run. That's where the random rooms, merged faces, surprise zooms and one-person-hogs-the-mic results come from. Each sentence in the prompts above removes one of those guesses.

No AI video tool?Upload one photo of each person and we make the booth video for you.

Skip prompting — make it from two photos

FAQ

Which AI video tools can I use these prompts with?

+

Any image-to-video or reference-to-video tool that accepts photos of the people. Adjust the reference syntax, aspect ratio and duration to your tool; results vary between tools and between runs.

Is this the exact prompt Make Hotel Lobby uses?

+

No. These are general-purpose prompts written for people using their own tools. Make Hotel Lobby runs its own setup, and you don't write any prompt there — you upload two photos.

Should I name Quavo, Takeoff or COLORS in the prompt?

+

No. Describe the look instead — orange booth, hanging microphone, two people trading lines. Naming real people can pull their likeness into the video, and many tools block it.

Will the prompt add the Hotel Lobby song?

+

No. Video tools don't include the original recording. Add the Hotel Lobby sound from TikTok's or Instagram's sound library when you post.

Updated . Only use photos you have permission to use.