YuE2, a new music-generation research system, is combining symbolic composition with audio synthesis in an effort to make AI-created songs more controllable. Instead of moving directly from a text prompt to a finished recording, the project first turns lyrics and style instructions into a musical plan that can be inspected and edited.
That intermediate plan contains elements such as melody, rhythm and chords. A user can change lyrics, tempo, arrangement or melodic choices and then generate a new full song with vocals and accompaniment. The approach is intended to make revisions more deliberate than repeatedly asking an audio model for another opaque output.
The YuE2 team reports a mean score of 6.9632 for its best-of-eight configuration on SongBench, compared with 6.8721 for Suno v5 in the same evaluation. The comparison used WildSongBench, a set of 192 prompts tested across 15 system settings. Best-of-eight means that one result was selected from eight generations using musicality, prompt control and lyric accuracy. These are project-reported automatic evaluation results, and rankings can change with the metric and selection method.
The broader release also describes two music-representation encoders, MERT2-30s and MERT2-FS. Both have 632 million parameters, with training contexts of 30 seconds and 300 seconds respectively. The team says the pair produced the leading result on 14 of 15 metrics in its MARBLE comparison, covering tasks including tagging, key, genre and emotion recognition.
A related transcription model, SheetSage2, is designed to extract multiple forms of musical structure with one system. It handles beats, downbeats, keys, chords, sections and melody. The researchers report leading scores on 10 of 13 benchmark metrics in comparisons that included SheetSage1, Madmom and task-specific systems. Some evaluation conditions differ between models, including the use of cross-validation for ChordFormer on the Chords1217 data set, so the headline counts should be read in that context.
Training-data provenance is a prominent part of the project’s presentation. The team says its models are trained primarily on public-domain CC0 music and synthetic material, with Tokenwave.AI supplying most of the synthetic training data under license.
YuE2’s most distinctive idea is therefore not only its claimed audio quality. It is the separation of composition and rendering, which gives users a visible structure between the prompt and the waveform. If the approach transfers well beyond demonstrations, symbolic planning could make generated music easier to direct, revise and evaluate while preserving the speed of an end-to-end audio system.
The demonstrations span multiple genres and languages, showing how material can be reshaped through deliberate changes to tempo, arrangement, words and melody.



