BARS system overview

BARS: Beat-Adaptive Rap Synthesis via Flow Matching

Accepted at ISMIR 2026

Taemin Cho, Hector Martel, Enyan Koh
BandLab Technologies

Abstract

We present BARS (Beat-Adaptive Rap Synthesis), a non-autoregressive rapping voice generation model based on Flow Matching, conditioned on a beat track (instrumental backing), lyrics, and a short reference voice for zero-shot timbre control. Unlike prior autoregressive rap synthesis, where rap pace and output duration are strictly determined by the beat, BARS enables independent control over the rap tempo, allowing flexible-length generation without explicit duration prediction or phoneme-level alignment.

We encode lyrics using a pretrained ByT5 encoder with a lightweight Transformer adapter trained jointly with the model, without relying on speech-pretrained text encoders, enabling natural handling of non-standard orthography (e.g., slang, abbreviations) and numerals common in rap lyrics.

To address alignment instabilities, such as token skipping or repetition, under strong beat conditioning, we propose Asymmetric Scaled RoPE for cross-attention, which introduces a positional structure to facilitate sequential lyric-audio alignment learning. Furthermore, the model supports partial re-generation via span masking, enabling lyric-level editing while preserving the surrounding vocal style.

Experiments demonstrate that the proposed model achieves high lyric fidelity, natural rap flow that is rhythmically coherent with the beat, and robust tempo control across a wide range of rapping styles and speeds.

🎤 ...but what if the abstract was rapped?
🎤 Reference Voice
🎹 Input Beat
A free beat from BandLab Beats: "Soul Train" by Fake Flowers.
🎹 Listen to "Soul Train" on BandLab Beats ↗
📝 Lyrics: the abstract above, fed directly into BARS as the rap's lyrics.
🎤 BARS-Generated Rap
🎵 Mixed with Beat
ℹ️ Note: for the demos below, we do not provide the beat track separately, as the license does not permit redistributing the unmodified beat on its own. Only the mix of the generated rap and the given beat is shown.

1. Style Versatility

🎤 Reference Voice
📝 Lyrics
Same lyrics, one beat — bars slow and heavy,
swap the beat out — now the flow is deadly.
The words don't change but the style sure does,
BARS shifts the delivery just because.

Boom bap beat? Watch the syllables sprawl,
trap hi-hats? Bars start to crawl and fall
into pockets you never knew were there —
same sixteen, brand new style in the air.

The lyrics stay locked but the vibe's a shapeshifter,
BARS reads the beat and becomes a flow drifter.
One set of words, infinite ways to ride —
the beat decides the style, BARS steps inside.
🎵 Flow Variations
UK Grime
Trap
Orchestral
Drill
Trip Hop
Slow Hip Hop

2. Zero-Shot Timbre Control

📝 Lyrics
Zero-shot timbre, your voice from a clip, powerful tech but we gotta stay legit.
Built for the artists, the grinders, the craft, not for misuse — we thought about that.
Explicit consent — that's the golden accord, unauthorized synthesis? Put down the sword.
The tech is powerful, use it with care, BARS is a tool — respect whose voice is there.
🎤 Reference Voice 🎵 BARS
Female
Male
Female2
Child (Female)
Donald Trump
Pikachu (Character Voice)

3. Speed Control

Speed is controlled by trimming the length of the input beat track: a shorter beat forces a faster rap flow to fit the same lyrics, while a longer beat produces a slower flow.

📝 Lyrics
First of its kind, to the best of our knowledge
Flow-Matching-based framework that's non-autoregressive
no duration prediction no phoneme alignment
to generate this rap you hear, right this moment
μ (mean speed of training data: 17.9 chars/s, σ = 2.9)
🐢 Slower than μ
μ − σ (15.0 cps)
μ − 2σ (12.1 cps)
μ − 3σ (9.2 cps)
🐇 Faster than μ
μ + σ (20.8 cps)
μ + 2σ (23.7 cps)
μ + 6σ (35.3 cps)
μ + 10σ (46.9 cps)

4. Multilingual Generation

This section is a preliminary showcase and not part of the main paper contributions.

🌍 Language 📝 Lyrics 🎵 Audio
Spanish Subiendo sobre el flow multilingüe sin freno
Brilla ByT5 dentro del modelo BARS pleno
Encoder de texto, rompe toda barrera
Rimas del otro lado del mundo en una sola esfera

Aunque cambie el idioma, el flow es uno
En el mundo de datos el rap es común
BARS + encoder escupe el futuro en acción
Generación multilingüe, esa es la conexión
Russian По мультиязычному флоу мы вверх летим
Внутри BARS-модели ByT5, скрытый ритм
Текстовый энкодер рушит языковую стену
Рифмы с другой стороны мира входят в сцену

Языки разные, но флоу один всегда
В мире данных рэп объединяет слова
BARS плюс энкодер — будущее в спит
Мультиязычная генерация — настоящий хит
Korean 멀티랭귀지 위를 타고 올라가는 flow
BARS 모델 속에 숨은 ByT5 glow
텍스트 인코더, 언어 벽을 깨
지구 반대편 라임도 한 번에 채

언어는 달라도 flow는 하나
데이터 속 세계가 랩으로 하나
Bars + encoder, 미래를 spit
다국어 랩 생성, that's the legit hit
French Dans le flow multilingue on s’élève sans arrêt
BARS ByT5, lumière calibrée
L’encodeur brise les murs, fusion des langues
Des rimes du monde entier qui s’entrechoquent

Les langues changent mais le groove reste un
Dans les datas le rap devient commun
BARS dans le système, futur en direct
Multilingual flow, vibe plus que parfaite
Japanese マルチリンガルの上を駆け上がるフロー
BARSモデルの中に潜むByT5のグロウ
テキストエンコーダー、言語の壁をブレイク
地球の反対側のライムも一瞬でメイク

言語は違ってもフローはひとつ
データの世界でラップはひとつ
BARS + エンコーダー、未来をスピット
多言語ラップ生成、それがリアルヒット
Italian Sul flow multilingue senza paura
BARS è il modello, ByT5 la struttura
Encoder di testo rompe il confine
E rime dal mondo intero arrivano infine

Lingue diverse, ma il flow è uno solo
Nel mondo dei dati il rap è il polo.
BARS + encoder, il futuro è in track.
Generazione globale, questo è un fact
Persian یه روز خوب میاد
که ما هم رو نکشیم
به هم نگاه بد نکنیم
با هم دوست باشیم و
دست بندازیم رو شونه‌های هم
آها، مثل بچگی‌ها تو دبستان
هیج‌کدوممون هم نیستیم بی‌کار
در حال ساخت و ساز ایران
Excerpt from "Ye Rooze Khoob Miad" by Hichkas