# Spotify reports gains from a conversational music-recommendation agent

> A Spotify research team describes how it used simulated conversations and automated prompt refinement to build a music agent. In A/B tests, the system increased listening and weekly active users compared with a feature limited to refining a listening session, the team reports.

Oossa · 2026-09-29 · https://oossa.com/en/spotify-reports-gains-from-a-conversational-music-recommendation-agent

Spotify researchers say a conversational recommendation agent helped increase listening and weekly active users in online A/B tests. The system lets people describe what they want in ordinary language—for example, “Italian indie artists I haven’t heard before”—rather than only asking for changes to a current listening session.

The results appear in a paper published on arXiv on Sept. 28, 2026, and listed as a RecSys 2026 paper. The authors report 14% more user listening, a 5% increase in weekly active users, and a 5% reduction in skip rate compared with Spotify’s earlier session-refinement experience.

## How the system was built

The team focused on a challenge that comes before launch: teaching an agent how to plan, including which tools to use and in what order, when there are not yet real user conversations to learn from. Its pipeline turns single-turn prompts into simulated, multi-turn conversations so the team can test the agent in advance.

It then uses an automated improvement loop to find and address mistakes in planning and tool use. The method compares different agent responses and applies iterative edits with help from a coding agent. The researchers say this raised quality by 8% over an already highly optimized manual prompt. They also say the process shortened development cycles, though the paper’s abstract does not give a time estimate.

## What the results show

The reported listening and engagement figures come from online tests against a specific earlier experience, not a comparison with every way people use Spotify. The paper does not say in its abstract how many people took part, how long the tests ran, or whether the agent is available to all users.

The work is mainly about a development method: using generated conversations and automated revisions to improve an agent before and during launch. That approach could help teams test natural-language recommendation tools even when they have little real-world interaction data at the start.

## The facts

- The paper was published on arXiv on Sept. 28, 2026, and lists RecSys 2026 as its journal reference.
- The researchers generated multi-turn conversations from single-turn prompts to test the agent before launch.
- The team reports an 8% quality improvement over a highly optimized manual prompt.
- In A/B tests, the researchers report 14% more listening, 5% more weekly active users, and a 5% lower skip rate than with the prior session-refinement experience.

## Why it matters

For listeners, the approach is meant to make it easier to ask for music using detailed, everyday descriptions. The reported gains are encouraging, but the abstract does not provide test size or duration, and it does not establish that the agent is available to everyone.

## Sources & references

1. [Bootstrapping Conversational Recommendation Agents At Spotify: Synthetic Data Generation and Self-Improvement Loops](https://arxiv.org/abs/2609.30297) – arXiv, 2026-09-28

Last updated: 2026-09-29
