Scaling Audio-Visual Speech Generation enables Versatile Transfer

ICML Anonymous Submission

Audio-Visual Gemma is a family of 3B and 10B visual-conditioned speech generation foundation models.
You can think of it as a Visual Language Model with native “speaking” and audio understanding capabilities.
We carefully study the model’s scaling properties and demonstrate its versatile transfer to novel audio-visual generation tasks.

Table of Contents

Model Image-to-Speech Generation Long-Form Spoken Captioning Video-to-Speech Generation Spoken Visual QA Spoken Instructions Generation Audio-Visual Sounds Prediction Audio-Visual Sounds Spoken Captioning Comparison to Prior Work

Audio-Visual Gemma

Audio-Visual Gemma is developed on three principles: (1) Native: Directly generates audio at pretraining without requiring explicit adaptation or alignment procedures,
(2) Simple: Utilizes an encoder-free and decoder-free architecture for processing and generating audio, (3) Scalable: Scales effectively with both data and model size.



Pretraining via Spoken Captioning

Audio-Visual Gemma is pretrained on 350M web-scale spoken captions, or roughly 285k hours of high-quality synthetic speech. We provide two pretraining variants: (1) joint text-speech, and (2) speech-only. In either case, audio tokens are produced from the Gemma LLM natively without having a speech decoder component.

Pretraining Setup

Versatile Audio-Visual Transfer

Audio-Visual Gemma works with a suite of audio-visual transfer setups, such as video-to-speech, spoken visual QA, and general sounds understanding. Distinctly, it adopts an encoder-free approach to contextual speech and audio understanding -- speech and audio are presented as hard tokens in-context to the model.

Finetuning Setup

Direct Image-to-Speech Generation at Scale

Explore Audio-Visual Gemma's speech generation samples on internet images. For each image, two speech samples are provided: one using an enrolled speaker, and another with a random speaker.

Webli 1

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: native american artifacts and a photo of a child

Ground Truth: native american trunk detail jpg

Webli 2

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: watercolor floral bouquet with blue and white flowers

Ground Truth: elegant gold letters bridal shower by mail invitation

Webli 3

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: clear makeup bag with black trim

Ground Truth: deli clear acrylic business card holder

Webli 4

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: clear acrylic business card holder

Ground Truth: deli clear acrylic business card holder

Webli 5

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: llama llama mad at mama

Ground Truth: screen shot 2013 01 07 at 7.15 25 pm

Webli 6

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: what good is the warmth of summer without the cold of winter image

Ground Truth: what good is the warmth of summer without the cold of winter image

Webli 7

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: destroy or be destroyed i just love that way of life image

Ground Truth: destroy or be destroyed i just love that way of life mike tyson picture quote

Webli 8

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: peony flower illustration drawing design

Ground Truth: color illustration of a lotus flower on a light background vector premium vector

Webli 9

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: portion of cola nuts close up shot on wooden background stock image

Ground Truth: portion of milled cola nut close up shot on wooden background stock image

Webli 10

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: a new kind of torture

Ground Truth: halloween 2018 page 4

Webli 11

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: stepped architrave set

Ground Truth: profile of 3 stepped architrave

Webli 12

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: best of the best silver award

Ground Truth: best of the best award

Webli 13

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: two cats looking away

Ground Truth: two different breeds of cat

Webli 14

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: gold vinyl record on a white background

Ground Truth: vinyl record lp isolated on white background 3d illustration

Webli 15

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: all natural deodorant lavender & tea tree

Ground Truth: all natural deodorant lavender & tea tree

Webli 16

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: seamless pattern with sewing accessories

Ground Truth: seamless pattern with sewing elements

Webli 17

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: 2013 dodge dart 4 door with alloy wheels

Ground Truth: 2013 dodge dart limited

Webli 18

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: burgundy flower girl dress with sunflowers

Ground Truth: burgundy flower girl dress long sleeve

Webli 19

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: abaya with belt

Ground Truth: dresses ladies turkish coat abaya

Webli 20

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: be kind not perfect print

Ground Truth: be real not perfect print calligraphy print calligraphy wall art printable calligraphy print digital download

Webli 21

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: figure 4 resistivity pseudo section along the line aa

Ground Truth: figure 5 line 1

Webli 22

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: set of horizontal banners of grass silhouettes vector

Ground Truth: horizontal banners of meadow silhouettes with short grass vector image

Webli 23

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: we just want to go home i left the stove on misc

Ground Truth: we just want to go home i left the stove on district 9 alien wants to go home

Webli 24

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: tribal arrows clip art

Ground Truth: clip art tribal arrows american indian

Webli 25

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: bearish classic t shirt

Ground Truth: bearish classic t shirt

Webli 26

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: beautiful little girl in a dress with polka dots

Ground Truth: girl in polka dot dress red gloves and bow standing looking awa

Webli 27

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: bearish classic t shirt

Ground Truth: bearish classic t shirt

Webli 28

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: sell your diamond

Ground Truth: certified vintage & antique estate jewelry

Webli 29

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: white and gold wrap dress

Ground Truth: gabri white chain print plunge wrap over mini dress

Webli 30

Speech Generation (Enrolled Speaker)

Speech Generation (Random Speaker)

Transcript: google photos memories

Ground Truth: google photos photo backup smartphone google i o mobile device data encryption android

Back to Top

Long-Form Spoken Captioning

Explore Audio-Visual Gemma's long-form (up to 20 seconds) spoken captioning generation, showcasing its ability to generate visually comprehensive and descriptive speech from images.

Image Paragraph 1

Long-Form Speech Generation

Transcript: A red barn with a silver roof is on a farm. There is a red gate on the side of the barn. There is a wooden fence on the side of the barn

Image Paragraph 2

Long-Form Speech Generation

Transcript: A man in a brown jacket and a black hat is walking on the street. There is a train on the tracks in front of the man. There are flowers on the ground in front of the train.

Image Paragraph 3

Long-Form Speech Generation

Transcript: A man is snowboarding. He is wearing a gray shirt and black and white pants. He is jumping off of a ramp. There is a brick wall behind him.

Image Paragraph 4

Long-Form Speech Generation

Transcript: A plate of food is shown. There are various types of pasta and vegetables on the plate. The pasta is yellow and curly. There are also some green vegetables.

Image Paragraph 5

Long-Form Speech Generation

Transcript: A brown teddy bear is sitting on a black leather couch. There is a mural of a wave on the wall behind the couch. The wave is blue and white. There is a mountain on the other side of the wave. The teddy bear is wearing a red bow tie.

Image Paragraph 6

Long-Form Speech Generation

Transcript: A bunch of teddy bears are sitting in a window. The window is white and has red and white letters on it. There is a sign on the window that says the bear shop.

Image Paragraph 11

Long-Form Speech Generation

Transcript: A silver pole has two red lights on it. There is a sign on the pole that says do not stop on tracks. There are cars on the street behind the pole.

Image Paragraph 13

Long-Form Speech Generation

Transcript: A man is sitting at a table outside. He is wearing a blue cardigan and a green shirt. There is a plate of food in front of him. There is a pizza on a wooden board in front of him.

Image Paragraph 14

Long-Form Speech Generation

Transcript: A man is sitting on a bench. He is wearing a tan hat, a brown vest, and gray pants. He is holding a cell phone in his hands. There is a black bag and a black suitcase on the bench. There are train tracks behind the man. There are trees behind the train tracks.

Image Paragraph 17

Long-Form Speech Generation

Transcript: A white house with black trim has three sheep grazing in the front yard. The sheep are all white with a blue stripe on their head. The house has three windows on the front and a brown door. The grass is green and has rocks in the yard. There is a gravel driveway in front of the house. There are bushes in the yard.

Image Paragraph 18

Long-Form Speech Generation

Transcript: A polar bear is swimming in the water. The bear is white and has black fur on its face. The bear has its mouth open and its eyes are open. The water is dark blue and has a lot of white foam on it.

Image Paragraph 21

Long-Form Speech Generation

Transcript: Two women are playing tennis on a grass court. The court is surrounded by a tall green fence. The fence is made of green fabric. The fabric is covered with a green tent. The women are wearing white shirts and white shorts. The woman on the left is swinging a racket at a yellow tennis ball. The woman on the right is swinging a racket at a white tennis ball.

Image Paragraph 23

Long-Form Speech Generation

Transcript: A woman is walking down the street carrying two baskets of vegetables on her back. She is wearing a plaid shirt and black pants. She has a ponytail in her hair.

Image Paragraph 25

Long-Form Speech Generation

Transcript: A black pole with a white sign on it. The sign says "C have you paid". There are cars parked on the sidewalk. There is a large building on the right.

Image Paragraph 27

Long-Form Speech Generation

Transcript: A miniature train set is on display. The tracks are made of metal and gravel. There are small green plants growing on the tracks. There are two trains on the tracks. One is green and yellow, and the other is blue and red. There are lights on the top of the trains. There is a black train engine on the ground.

Image Paragraph 28

Long-Form Speech Generation

Transcript: There are four small brown bears on the grass. There is a large log on the ground behind them. There is a large rock next to the log.

Direct Video-to-Speech Generation

Explore Audio-Visual Gemma's direct video-to-speech generation, extending its image-conditioned speech generation capability to video content.

Video-to-Speech Generation

Transcript: two deer are grazing in the woods

Video-to-Speech Generation

Transcript: a football team is huddled around a coach who is explaining the play

Video-to-Speech Generation

Transcript: a video of the World Trade Center collapsing

Video-to-Speech Generation

Transcript: a gymnast is doing a backflip off of a high bar

Video-to-Speech Generation

Transcript: two cats are playing with a string

Video-to-Speech Generation

Transcript: a cartoon of a man and a woman talking to each other

Video-to-Speech Generation

Transcript: oranges are being thrown into a pool of water

Video-to-Speech Generation

Transcript: a plane is flying over a city at sunset

Video-to-Speech Generation

Transcript: a man is laughing hysterically

Video-to-Speech Generation

Transcript: a pallet of boxes is being moved down a conveyor belt in a factory

Video-to-Speech Generation

Transcript: two teams of rowers are racing each other in a river

Video-to-Speech Generation

Transcript: a black and white video of a man signing a document

Video-to-Speech Generation

Transcript: a math problem is being worked out on a Blackboard

Video-to-Speech Generation

Transcript: yellow pills are falling onto a white surface

Video-to-Speech Generation

Transcript: a video game of a BMW driving down a racetrack

Video-to-Speech Generation

Transcript: two hippopotamus in an aquarium with their mouths open wide

Video-to-Speech Generation

Transcript: a black and white video of a cowboy riding a horse

Video-to-Speech Generation

Transcript: a person is looking at a website and they are highlighting some text

Spoken Visual Question Answering

Explore Audio-Visual Gemma's spoken visual question answering capability in a speech-in and speech-out setting. Given an image, it "listens" to questions and "reads out" the answers.

Spoken VQA 2

Input Spoken Question

Output Spoken Response

Spoken VQA 3

Input Spoken Question

Output Spoken Response

Spoken VQA 4

Input Spoken Question

Output Spoken Response

Spoken VQA 5

Input Spoken Question

Output Spoken Response

Spoken VQA 6

Input Spoken Question

Output Spoken Response

Spoken VQA 7

Input Spoken Question

Output Spoken Response

Spoken VQA 8

Input Spoken Question

Output Spoken Response

Spoken VQA 9

Input Spoken Question

Output Spoken Response

Spoken VQA 10

Input Spoken Question

Output Spoken Response

Spoken VQA 11

Input Spoken Question

Output Spoken Response

Spoken VQA 12

Input Spoken Question

Output Spoken Response

Spoken VQA 22

Input Spoken Question

Output Spoken Response

Spoken VQA 23

Input Spoken Question

Output Spoken Response

Spoken VQA 24

Input Spoken Question

Output Spoken Response

Spoken VQA 25

Input Spoken Question

Output Spoken Response

Spoken VQA 26

Input Spoken Question

Output Spoken Response

Spoken Instructions Generation from Egocentric Views

Explore how Audio-Visual Gemma generates spoken instructions from egocentric video perspectives, showing its ability to process and understand complex action sequences beyond typical video content.

Spoken Instructions Generation

Transcript: cut leek

Spoken Instructions Generation

Transcript: wash pot

Spoken Instructions Generation

Transcript: drink water

Spoken Instructions Generation

Transcript: pour the water into the pot

Spoken Instructions Generation

Transcript: remove slice

Spoken Instructions Generation

Transcript: dry hands

Spoken Instructions Generation

Transcript: rinse knife

Spoken Instructions Generation

Transcript: wash colander

Spoken Instructions Generation

Transcript: put milk into fridge

Spoken Instructions Generation

Transcript: stir blueberries

Spoken Instructions Generation

Transcript: turn oven on

Spoken Instructions Generation

Transcript: pour juice

Spoken Instructions Generation

Transcript: pour corn in pan

Spoken Instructions Generation

Transcript: turn on coffee machine

Spoken Instructions Generation

Transcript: take sponge

Spoken Instructions Generation

Transcript: get kitchen towel

Open-Vocabulary Audio-Visual Sound Events Prediction

Explore Audio-Visual Gemma's capability to predict sound events using both video and audio streams without relying on a predefined label set. Observe how some examples depend on video inputs, others rely solely on audio, and some combine both.

Audio Source (sounds, music, noise)

Model Predictions: Laughter

Audio Source (sounds, music, noise)

Model Predictions: Wind noise (microphone), Motorboat, Wind, Boat

Audio Source (sounds, music, noise)

Model Predictions: Rock and roll

Audio Source (sounds, music, noise)

Model Predictions: Siren

Audio Source (sounds, music, noise)

Model Predictions: Sewing machine

Audio Source (sounds, music, noise)

Model Predictions: Whimper

Audio Source (sounds, music, noise)

Model Predictions: Tap, Water tap

Audio Source (sounds, music, noise)

Model Predictions: Musical instrument, Tabla

Audio Source (sounds, music, noise)

Model Predictions: Gong

Audio Source (sounds, music, noise)

Model Predictions: Electric shaver

Audio Source (sounds, music, noise)

Model Predictions: Hair dryer

Audio Source (sounds, music, noise)

Model Predictions: Musical instrument, Plucked string instrument, Ukulele

Audio Source (sounds, music, noise)

Model Predictions: Mantra

Audio Source (sounds, music, noise)

Model Predictions: Lullaby

Audio Source (sounds, music, noise)

Model Predictions: Classical music

Audio Source (sounds, music, noise)

Model Predictions: Trombone, Brass instrument

Audio-Visual Sound Events Spoken Captioning

Explore Audio-Visual Gemma's capability to generate spoken captions for sound events by leveraging both video and audio streams.

Audio Source (sounds, music, noise)

Sound Event Spoken Captioning

Transcript: A dog is whimpering and a woman speaks

Audio Source (sounds, music, noise)

Sound Event Spoken Captioning

Transcript: A bird is chirping

Audio Source (sounds, music, noise)

Sound Event Spoken Captioning

Transcript: Waves crashing

Audio Source (sounds, music, noise)

Sound Event Spoken Captioning

Transcript: A man speaks and sprays paint

Audio Source (sounds, music, noise)

Sound Event Spoken Captioning

Transcript: A child speaks and a crowd murmurs

Audio Source (sounds, music, noise)

Sound Event Spoken Captioning

Transcript: A motorboat engine running

Audio Source (sounds, music, noise)

Sound Event Spoken Captioning

Transcript: A person whistling

Audio Source (sounds, music, noise)

Sound Event Spoken Captioning

Transcript: A train horn blows as a train passes by

Audio Source (sounds, music, noise)

Sound Event Spoken Captioning

Transcript: A sheep bleats

Audio Source (sounds, music, noise)

Sound Event Spoken Captioning

Transcript: A horse is trotting

Audio Source (sounds, music, noise)

Sound Event Spoken Captioning

Transcript: Fire truck sirens

Audio Source (sounds, music, noise)

Sound Event Spoken Captioning

Transcript: A person burping continuously

Text-Free Training: Comparison to Prior Work

This section compares Audio-Visual Gemma's image-to-speech generation with Hsu et al.'s Image-to-Speech Synthesis Using Learned Segmental Units (ACL 2020) under a "text-free" training setup. Sampled images are taken from their demo webpage, enabling a fair comparison of generation quality between our model and theirs.

Input Image for Comparison 6515
Input Image for Comparison 4882
Input Image for Comparison 26030
Input Image for Comparison 24752
Input Image for Comparison 1686
Input Image for Comparison 13471
Input Image for Comparison 8309
Input Image for Comparison 3233
Input Image for Comparison 30773
Input Image for Comparison 29172
Input Image for Comparison 27582
Input Image for Comparison 26166
Input Image for Comparison 23024
Input Image for Comparison 21549
Input Image for Comparison 19877
Input Image for Comparison 18302
Input Image for Comparison 16622
Input Image for Comparison 155
Input Image for Comparison 15061
Input Image for Comparison 11809

Ethics Statement

We acknowledge that, like other advanced AI innovations, our technology carries the potential for misuse and unintended harm. Audio-Visual Gemma, in its research application, processes visual inputs to generate visually coherent spoken content. To mitigate risks of misuse, we have decided not to release the model checkpoints for now. However, we describe our approach and findings in detail within this paper to support transparency and collaboration in the research community. While we remain committed to advancing AI and fostering openness, we believe it is essential to balance these values with responsible and ethical stewardship.