Audio-Visual Gemma is a family of 3B and 10B visual-conditioned speech generation
foundation models.
You can think of it as a Visual Language Model with native “speaking” and audio understanding capabilities.
We carefully study the model’s scaling properties and demonstrate its versatile transfer to novel audio-visual
generation tasks.
Audio-Visual Gemma is developed on three principles:
(1) Native: Directly generates audio at pretraining without requiring explicit adaptation or
alignment procedures,
(2) Simple: Utilizes an encoder-free and decoder-free architecture for processing and
generating audio,
(3) Scalable: Scales effectively with both data and model size.
Audio-Visual Gemma is pretrained on 350M web-scale spoken captions, or roughly 285k hours of high-quality synthetic speech. We provide two pretraining variants: (1) joint text-speech, and (2) speech-only. In either case, audio tokens are produced from the Gemma LLM natively without having a speech decoder component.
Audio-Visual Gemma works with a suite of audio-visual transfer setups, such as video-to-speech, spoken visual QA, and general sounds understanding. Distinctly, it adopts an encoder-free approach to contextual speech and audio understanding -- speech and audio are presented as hard tokens in-context to the model.
Explore Audio-Visual Gemma's speech generation samples on internet images. For each image, two speech samples are provided: one using an enrolled speaker, and another with a random speaker.
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: native american artifacts and a photo of a child
Ground Truth: native american trunk detail jpg
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: watercolor floral bouquet with blue and white flowers
Ground Truth: elegant gold letters bridal shower by mail invitation
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: clear makeup bag with black trim
Ground Truth: deli clear acrylic business card holder
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: clear acrylic business card holder
Ground Truth: deli clear acrylic business card holder
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: llama llama mad at mama
Ground Truth: screen shot 2013 01 07 at 7.15 25 pm
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: what good is the warmth of summer without the cold of winter image
Ground Truth: what good is the warmth of summer without the cold of winter image
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: destroy or be destroyed i just love that way of life image
Ground Truth: destroy or be destroyed i just love that way of life mike tyson picture quote
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: peony flower illustration drawing design
Ground Truth: color illustration of a lotus flower on a light background vector premium vector
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: portion of cola nuts close up shot on wooden background stock image
Ground Truth: portion of milled cola nut close up shot on wooden background stock image
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: a new kind of torture
Ground Truth: halloween 2018 page 4
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: stepped architrave set
Ground Truth: profile of 3 stepped architrave
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: best of the best silver award
Ground Truth: best of the best award
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: two cats looking away
Ground Truth: two different breeds of cat
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: gold vinyl record on a white background
Ground Truth: vinyl record lp isolated on white background 3d illustration
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: all natural deodorant lavender & tea tree
Ground Truth: all natural deodorant lavender & tea tree
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: seamless pattern with sewing accessories
Ground Truth: seamless pattern with sewing elements
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: 2013 dodge dart 4 door with alloy wheels
Ground Truth: 2013 dodge dart limited
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: burgundy flower girl dress with sunflowers
Ground Truth: burgundy flower girl dress long sleeve
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: abaya with belt
Ground Truth: dresses ladies turkish coat abaya
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: be kind not perfect print
Ground Truth: be real not perfect print calligraphy print calligraphy wall art printable calligraphy print digital download
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: figure 4 resistivity pseudo section along the line aa
Ground Truth: figure 5 line 1
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: set of horizontal banners of grass silhouettes vector
Ground Truth: horizontal banners of meadow silhouettes with short grass vector image
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: we just want to go home i left the stove on misc
Ground Truth: we just want to go home i left the stove on district 9 alien wants to go home
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: tribal arrows clip art
Ground Truth: clip art tribal arrows american indian
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: bearish classic t shirt
Ground Truth: bearish classic t shirt
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: beautiful little girl in a dress with polka dots
Ground Truth: girl in polka dot dress red gloves and bow standing looking awa
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: bearish classic t shirt
Ground Truth: bearish classic t shirt
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: sell your diamond
Ground Truth: certified vintage & antique estate jewelry
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: white and gold wrap dress
Ground Truth: gabri white chain print plunge wrap over mini dress
Speech Generation (Enrolled Speaker)
Speech Generation (Random Speaker)
Transcript: google photos memories
Ground Truth: google photos photo backup smartphone google i o mobile device data encryption android
Explore Audio-Visual Gemma's long-form (up to 20 seconds) spoken captioning generation, showcasing its ability to generate visually comprehensive and descriptive speech from images.
Long-Form Speech Generation
Transcript: A red barn with a silver roof is on a farm. There is a red gate on the side of the barn. There is a wooden fence on the side of the barn
Long-Form Speech Generation
Transcript: A man in a brown jacket and a black hat is walking on the street. There is a train on the tracks in front of the man. There are flowers on the ground in front of the train.
Long-Form Speech Generation
Transcript: A man is snowboarding. He is wearing a gray shirt and black and white pants. He is jumping off of a ramp. There is a brick wall behind him.
Long-Form Speech Generation
Transcript: A plate of food is shown. There are various types of pasta and vegetables on the plate. The pasta is yellow and curly. There are also some green vegetables.
Long-Form Speech Generation
Transcript: A brown teddy bear is sitting on a black leather couch. There is a mural of a wave on the wall behind the couch. The wave is blue and white. There is a mountain on the other side of the wave. The teddy bear is wearing a red bow tie.
Long-Form Speech Generation
Transcript: A bunch of teddy bears are sitting in a window. The window is white and has red and white letters on it. There is a sign on the window that says the bear shop.
Long-Form Speech Generation
Transcript: A silver pole has two red lights on it. There is a sign on the pole that says do not stop on tracks. There are cars on the street behind the pole.
Long-Form Speech Generation
Transcript: A man is sitting at a table outside. He is wearing a blue cardigan and a green shirt. There is a plate of food in front of him. There is a pizza on a wooden board in front of him.
Long-Form Speech Generation
Transcript: A man is sitting on a bench. He is wearing a tan hat, a brown vest, and gray pants. He is holding a cell phone in his hands. There is a black bag and a black suitcase on the bench. There are train tracks behind the man. There are trees behind the train tracks.
Long-Form Speech Generation
Transcript: A white house with black trim has three sheep grazing in the front yard. The sheep are all white with a blue stripe on their head. The house has three windows on the front and a brown door. The grass is green and has rocks in the yard. There is a gravel driveway in front of the house. There are bushes in the yard.
Long-Form Speech Generation
Transcript: A polar bear is swimming in the water. The bear is white and has black fur on its face. The bear has its mouth open and its eyes are open. The water is dark blue and has a lot of white foam on it.
Long-Form Speech Generation
Transcript: Two women are playing tennis on a grass court. The court is surrounded by a tall green fence. The fence is made of green fabric. The fabric is covered with a green tent. The women are wearing white shirts and white shorts. The woman on the left is swinging a racket at a yellow tennis ball. The woman on the right is swinging a racket at a white tennis ball.
Long-Form Speech Generation
Transcript: A woman is walking down the street carrying two baskets of vegetables on her back. She is wearing a plaid shirt and black pants. She has a ponytail in her hair.
Long-Form Speech Generation
Transcript: A black pole with a white sign on it. The sign says "C have you paid". There are cars parked on the sidewalk. There is a large building on the right.
Long-Form Speech Generation
Transcript: A miniature train set is on display. The tracks are made of metal and gravel. There are small green plants growing on the tracks. There are two trains on the tracks. One is green and yellow, and the other is blue and red. There are lights on the top of the trains. There is a black train engine on the ground.
Long-Form Speech Generation
Transcript: There are four small brown bears on the grass. There is a large log on the ground behind them. There is a large rock next to the log.
Explore Audio-Visual Gemma's direct video-to-speech generation, extending its image-conditioned speech generation capability to video content.
Video-to-Speech Generation
Transcript: two deer are grazing in the woods
Video-to-Speech Generation
Transcript: a football team is huddled around a coach who is explaining the play
Video-to-Speech Generation
Transcript: a video of the World Trade Center collapsing
Video-to-Speech Generation
Transcript: a gymnast is doing a backflip off of a high bar
Video-to-Speech Generation
Transcript: two cats are playing with a string
Video-to-Speech Generation
Transcript: a cartoon of a man and a woman talking to each other
Video-to-Speech Generation
Transcript: oranges are being thrown into a pool of water
Video-to-Speech Generation
Transcript: a plane is flying over a city at sunset
Video-to-Speech Generation
Transcript: a man is laughing hysterically
Video-to-Speech Generation
Transcript: a pallet of boxes is being moved down a conveyor belt in a factory
Video-to-Speech Generation
Transcript: two teams of rowers are racing each other in a river
Video-to-Speech Generation
Transcript: a black and white video of a man signing a document
Video-to-Speech Generation
Transcript: a math problem is being worked out on a Blackboard
Video-to-Speech Generation
Transcript: yellow pills are falling onto a white surface
Video-to-Speech Generation
Transcript: a video game of a BMW driving down a racetrack
Video-to-Speech Generation
Transcript: two hippopotamus in an aquarium with their mouths open wide
Video-to-Speech Generation
Transcript: a black and white video of a cowboy riding a horse
Video-to-Speech Generation
Transcript: a person is looking at a website and they are highlighting some text
Explore Audio-Visual Gemma's spoken visual question answering capability in a speech-in and speech-out setting. Given an image, it "listens" to questions and "reads out" the answers.
Input Spoken Question
Output Spoken Response
Input Spoken Question
Output Spoken Response
Input Spoken Question
Output Spoken Response
Input Spoken Question
Output Spoken Response
Input Spoken Question
Output Spoken Response
Input Spoken Question
Output Spoken Response
Input Spoken Question
Output Spoken Response
Input Spoken Question
Output Spoken Response
Input Spoken Question
Output Spoken Response
Input Spoken Question
Output Spoken Response
Input Spoken Question
Output Spoken Response
Input Spoken Question
Output Spoken Response
Input Spoken Question
Output Spoken Response
Input Spoken Question
Output Spoken Response
Input Spoken Question
Output Spoken Response
Input Spoken Question
Output Spoken Response
Explore how Audio-Visual Gemma generates spoken instructions from egocentric video perspectives, showing its ability to process and understand complex action sequences beyond typical video content.
Spoken Instructions Generation
Transcript: cut leek
Spoken Instructions Generation
Transcript: wash pot
Spoken Instructions Generation
Transcript: drink water
Spoken Instructions Generation
Transcript: pour the water into the pot
Spoken Instructions Generation
Transcript: remove slice
Spoken Instructions Generation
Transcript: dry hands
Spoken Instructions Generation
Transcript: rinse knife
Spoken Instructions Generation
Transcript: wash colander
Spoken Instructions Generation
Transcript: put milk into fridge
Spoken Instructions Generation
Transcript: stir blueberries
Spoken Instructions Generation
Transcript: turn oven on
Spoken Instructions Generation
Transcript: pour juice
Spoken Instructions Generation
Transcript: pour corn in pan
Spoken Instructions Generation
Transcript: turn on coffee machine
Spoken Instructions Generation
Transcript: take sponge
Spoken Instructions Generation
Transcript: get kitchen towel
Explore Audio-Visual Gemma's capability to predict sound events using both video and audio streams without relying on a predefined label set. Observe how some examples depend on video inputs, others rely solely on audio, and some combine both.
Audio Source (sounds, music, noise)
Model Predictions: Laughter
Audio Source (sounds, music, noise)
Model Predictions: Wind noise (microphone), Motorboat, Wind, Boat
Audio Source (sounds, music, noise)
Model Predictions: Rock and roll
Audio Source (sounds, music, noise)
Model Predictions: Siren
Audio Source (sounds, music, noise)
Model Predictions: Sewing machine
Audio Source (sounds, music, noise)
Model Predictions: Whimper
Audio Source (sounds, music, noise)
Model Predictions: Tap, Water tap
Audio Source (sounds, music, noise)
Model Predictions: Musical instrument, Tabla
Audio Source (sounds, music, noise)
Model Predictions: Gong
Audio Source (sounds, music, noise)
Model Predictions: Electric shaver
Audio Source (sounds, music, noise)
Model Predictions: Hair dryer
Audio Source (sounds, music, noise)
Model Predictions: Musical instrument, Plucked string instrument, Ukulele
Audio Source (sounds, music, noise)
Model Predictions: Mantra
Audio Source (sounds, music, noise)
Model Predictions: Lullaby
Audio Source (sounds, music, noise)
Model Predictions: Classical music
Audio Source (sounds, music, noise)
Model Predictions: Trombone, Brass instrument
Explore Audio-Visual Gemma's capability to generate spoken captions for sound events by leveraging both video and audio streams.
Audio Source (sounds, music, noise)
Sound Event Spoken Captioning
Transcript: A dog is whimpering and a woman speaks
Audio Source (sounds, music, noise)
Sound Event Spoken Captioning
Transcript: A bird is chirping
Audio Source (sounds, music, noise)
Sound Event Spoken Captioning
Transcript: Waves crashing
Audio Source (sounds, music, noise)
Sound Event Spoken Captioning
Transcript: A man speaks and sprays paint
Audio Source (sounds, music, noise)
Sound Event Spoken Captioning
Transcript: A child speaks and a crowd murmurs
Audio Source (sounds, music, noise)
Sound Event Spoken Captioning
Transcript: A motorboat engine running
Audio Source (sounds, music, noise)
Sound Event Spoken Captioning
Transcript: A person whistling
Audio Source (sounds, music, noise)
Sound Event Spoken Captioning
Transcript: A train horn blows as a train passes by
Audio Source (sounds, music, noise)
Sound Event Spoken Captioning
Transcript: A sheep bleats
Audio Source (sounds, music, noise)
Sound Event Spoken Captioning
Transcript: A horse is trotting
Audio Source (sounds, music, noise)
Sound Event Spoken Captioning
Transcript: Fire truck sirens
Audio Source (sounds, music, noise)
Sound Event Spoken Captioning
Transcript: A person burping continuously
This section compares Audio-Visual Gemma's image-to-speech generation with Hsu et al.'s Image-to-Speech Synthesis Using Learned Segmental Units (ACL 2020) under a "text-free" training setup. Sampled images are taken from their demo webpage, enabling a fair comparison of generation quality between our model and theirs.
We acknowledge that, like other advanced AI innovations, our technology carries the potential for misuse and unintended harm. Audio-Visual Gemma, in its research application, processes visual inputs to generate visually coherent spoken content. To mitigate risks of misuse, we have decided not to release the model checkpoints for now. However, we describe our approach and findings in detail within this paper to support transparency and collaboration in the research community. While we remain committed to advancing AI and fostering openness, we believe it is essential to balance these values with responsible and ethical stewardship.