Design a Music Streaming Platform Like Spotify
Spotify delivers high-fidelity audio streaming to over 600 million monthly active users across a catalog of 100+ million songs. Unlike video streaming (where chunk sizes are large and video starts within 1β2 seconds), audio streaming demands near-instantaneous playback startup (), low bandwidth consumption, seamless offline synchronization, and real-time collaborative playlists.
1. Understanding the Problem
Functional Requirements
- Audio Streaming: Stream music with instant playback start and seamless quality switching (96 kbps, 160 kbps, 320 kbps in Ogg Vorbis / AAC).
- Catalog Search & Discovery: Search by track, artist, album, or podcast in .
- Collaborative Playlists: Multiple users can edit, add, delete, and re-order songs in a playlist concurrently with real-time updates.
- Offline Listening: Premium users can download encrypted tracks to local device storage with a 30-day license check.
- Personalized Recommendations: Daily Mixes, Discover Weekly, and real-time radio streams generated via vector embeddings.
Non-Functional Requirements
- Sub-200ms Playback Startup: Clicking "Play" must begin sound output within 200ms.
- Zero Audio Glitches / Jitter: Audio buffer underrun rate of streams.
- High Concurrency & Scale: 600M Monthly Active Users (MAU), 100M+ songs.
- Storage Durability: 100% durability for artist catalogs and audio masters.
Capacity Estimations & Sizing (5 Years)
- Active User Base: 600 Million MAU, 200 Million Daily Active Users (DAU).
- Concurrent Active Streams: Peak 20 Million simultaneous audio listeners.
- Audio Catalog Sizing:
- 100 Million tracks.
- Average song length: 3.5 minutes (210 seconds).
- Bitrates: Standard (96 kbps ), High (160 kbps ), Very High (320 kbps Ogg Vorbis ).
- Storage per track across 3 tiers: .
- Catalog Storage: (excluding raw lossless master backups ).
- Peak Streaming Bandwidth:
- .
2. The Set Up
Defining Core Entities
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β TRACK β
ββββββββββββββββββββ¬βββββββββββββββ¬βββββββββββββββββββββββ€
β track_id β UUID β PRIMARY KEY β
β isrc_code β VARCHAR(32) β Global Industry ID β
β title β VARCHAR(256) β Song Name β
β artist_id β UUID β Primary Artist β
β duration_ms β INT β Milliseconds β
β audio_hash β CHAR(64) β SHA-256 Audio ID β
β explicit β BOOLEAN β Parental Advisory β
ββββββββββββββββββββ΄βββββββββββββββ΄βββββββββββββββββββββββ
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β AUDIO_FILE β
ββββββββββββββββββββ¬βββββββββββββββ¬βββββββββββββββββββββββ€
β file_id β UUID β PRIMARY KEY β
β track_id β UUID β FOREIGN KEY β
β format β ENUM β OGG_VORBIS, AAC, FLACβ
β bitrate_kbps β INT β 96, 160, 320 β
β s3_uri β VARCHAR(512) β Object Storage Path β
β encryption_key_idβ UUID β AES-128 Key ID β
ββββββββββββββββββββ΄βββββββββββββββ΄βββββββββββββββββββββββ
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β PLAYLIST_ENTRY β
ββββββββββββββββββββ¬βββββββββββββββ¬βββββββββββββββββββββββ€
β playlist_id β UUID β Composite PK β
β track_id β UUID β Track Reference β
β added_by_user_id β UUID β Creator of Entry β
β fractional_pos β DOUBLE β Fractional Index β
β added_at β TIMESTAMP β Sort Timestamp β
ββββββββββββββββββββ΄βββββββββββββββ΄βββββββββββββββββββββββ
The API Design
1. Request Audio Stream Chunkβ
GET /api/v1/storage/tracks/{track_id}/stream
Range: bytes=0-163839 // Fetch first 160 KB (first 8-10 seconds)
Authorization: Bearer <user_session_token>
Response (206 Partial Content):
HTTP/1.1 206 Partial Content
Content-Range: bytes 0-163839/4404019
Content-Type: audio/ogg
Cache-Control: public, max-age=31536000, immutable
X-Encryption-IV: 8f4a9b...
<binary audio bytes>
2. Re-Order Song in Collaborative Playlistβ
PATCH /api/v1/playlists/{playlist_id}/tracks/{track_id}/position
Content-Type: application/json
If-Match: "rev_4210"
{
"prev_fractional_pos": 2.0,
"next_fractional_pos": 3.0
}
Response (200 OK):
{
"playlist_id": "7b8e...",
"track_id": "9a12...",
"new_fractional_pos": 2.5,
"revision": "rev_4211"
}
3. High-Level Design
Walkthrough of Core Flows
1. Audio Ingestion & Transcodingβ
- Record labels upload lossless FLAC audio masters to Spotify's Ingestion Gateway via SFTP/S3.
- The ingestion worker parses audio metadata, calculates acoustic loudness (LUFS normalization to ), and computes the Audio Fingerprint (AcoustID) for deduplication.
- Transcoding workers convert the master into Ogg Vorbis (desktop/Android) and AAC (iOS/Web) across 96 kbps, 160 kbps, and 320 kbps.
- Each file is encrypted using AES-128 CTR mode and stored in Amazon S3 / Google Cloud Storage, with metadata indexed in PostgreSQL / Cassandra.
2. The Playback Start Flow (Sub-200ms)β
- When a user clicks a song, the client checks its local persistent disk cache:
- Local Cache Hit (30β40% of plays): Plays immediately from disk ().
- Local Cache Miss: Client sends an HTTP range request
Range: bytes=0-163839to the nearest Cloudflare/Fastly CDN edge.
- The CDN returns the first 160 KB chunk (representing the first 8β10 seconds of the song).
- The client player initializes the audio decoder with the Ogg Vorbis header and begins sound output within 150ms.
- While the first 10 seconds play, the background worker progressively streams the remainder of the file in larger 512 KB chunks over an HTTP/2 or HTTP/3 multiplexed connection.
4. Potential Deep Dives & Bottlenecks
Deep Dive 1: How Spotify Achieves Sub-200ms Latency
Why does Spotify feel instantaneous while YouTube takes 1β2 seconds to start?
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β SUB-200MS AUDIO STARTUP PIPELINE β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β β
β 1. Predictive Pre-Buffering β
β When Track N plays, client pre-fetches the first β
β 160 KB of Track N+1 in the background. β
β β
β 2. Local Disk Cache (LRU on Device) β
β Frequently played tracks (liked songs, playlists) β
β are stored encrypted in local flash storage. β
β β
β 3. Ogg Vorbis Stream Header Separation β
β Codec initialization headers (Ogg Page 0) are β
β embedded in the first 4 KB, allowing audio hardwareβ
β DAC to initialize before the full file loads. β
β β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
- Range Request Chunking: Fetching the entire 5 MB file blocks playback until the whole download finishes. By splitting the request into an urgent initial 160 KB chunk followed by larger background chunks, the network download completes well ahead of real-time playback speed ().
- HTTP/3 Connection Multiplexing: Eliminates TCP head-of-line blocking and saves 1 RTT during connection handshakes.
Deep Dive 2: Collaborative Real-Time Playlists: Fractional Indexing
When multiple friends edit a shared playlist simultaneously, how do we support moving song between song and song without re-indexing all entries?
Traditional Array Problem:
[Song A (Index 0), Song B (Index 1), Song C (Index 2), Song D (Index 3)]
Inserting between B and C requires shifting C and D (O(N) database write cascade).
Concurrent inserts cause identical collision indexes!
Fractional Indexing Solution:
[Song A (Pos: 1.0), Song B (Pos: 2.0), Song C (Pos: 3.0)]
Insert between B (2.0) and C (3.0):
New Pos = (2.0 + 3.0) / 2 = 2.5
β Exactly ONE row updated in the database.
β Commutative and conflict-free for concurrent operations.
- Precision Exhaustion (Rebalancing): Over time, repeated insertions between adjacent items produce tiny fractional deltas (). When floating-point precision drops below , a background worker executes a rebalancing transaction, resetting positions to clean integers ().
Deep Dive 3: Offline Sync Licensing & DRM
How does Spotify prevent users from extracting raw MP3s or keeping songs forever after cancelling Premium?
- Symmetric Key Wrapping: Audio files on device storage are encrypted with unique AES-128 keys.
- Client Key Store: The AES-128 decryption key is never stored in plaintext on disk; it is requested from Spotify's KMS and held only in volatile device RAM or the OS Secure Enclave / KeyStore.
- 30-Day Lease Token: When downloading for offline listening, the server issues a signed cryptographic lease token expiring in 30 days. The mobile app must connect to the Internet at least once every 30 days to renew the lease; if not renewed, the local decryption keys are purged from the device.
Deep Dive 4: Music Recommendation Engine (Two-Tower Model + Annoy Index)
How does Spotify generate personalized recommendations from 100M songs in ?
User Context (History, Time, Likes) βββΊ [User Neural Tower] βββΊ 128-D Vector U
β Dot
Candidate Song (Audio, Genre, BPM) βββΊ [Item Neural Tower] βββΊ 128-D Vector I β Product
βΌ
Cosine Similarity Score
- Offline Training: Two-Tower deep neural networks project users and songs into a shared 128-dimensional latent embedding space.
- Approximate Nearest Neighbor (ANN): Spotify developed and open-sourced Annoy (Approximate Nearest Neighbors Oh Yeah), which creates a forest of random projection trees in memory.
- Real-Time Serving: When a user opens Spotify, their user vector is queried against the in-memory Annoy index. Annoy traverses the binary trees to retrieve the top 100 candidate tracks in , which are then ranked by a final multi-task scoring model.
5. Architectural Trade-Off Matrix
| Design Alternative | Option A | Option B | Selected Choice & Rationale |
|---|---|---|---|
| Audio Format | MP3 (Legacy standard) | Ogg Vorbis / AAC | Ogg Vorbis / AAC: Superior acoustic compression efficiency; 160 kbps Ogg sounds identical to 320 kbps MP3, saving 50% CDN egress bandwidth. |
| Playlist Ordering | Integer Array Indices () | Fractional Indexing () | Fractional Indexing: Enables concurrent insertion and reordering without updating millions of sibling rows. |
| Streaming Protocol | Stateful WebSockets / RTMP | Chunked HTTP/2 Range Requests | HTTP/2 Range Requests: Highly cacheable at standard CDN edge caches; avoids maintaining expensive stateful streaming servers. |
| Recommendation Search | Exact Cosine Matrix Multiplication | Approximate Nearest Neighbors (Annoy) | Annoy (ANN): Trades 1% precision for a speedup, serving recommendations in . |
6. What is Expected at Each Level?
Mid-Level (L4 / IC4)
- Clarifies functional requirements, understanding the difference between audio streaming and progressive downloading.
- Designs basic relational schemas for users, tracks, albums, and playlists.
- Explains why audio files are cached on client devices and CDN edges.
Senior (L5 / IC5)
- Explains the sub-200ms playback startup optimization (first 160 KB chunk range request).
- Analyzes collaborative playlist concurrency using fractional indexing to eliminate write cascades.
- Designs offline playback security with AES-128 encryption and 30-day lease token expirations.
- Details the acoustic normalization and transcoding pipeline.
Staff+ (L6 / Principal)
- Evaluates peer-to-peer (P2P) vs pure CDN economics for audio delivery.
- Formulates Approximate Nearest Neighbor (ANN) vector search topologies using Annoy/HNSW for real-time recommendations.
- Analyzes audio streaming over lossy cellular networks using HTTP/3 QUIC connection migration (switching from Wi-Fi to 5G without stream interruption).
- Designs global multi-region active-active playlist synchronization with CRDT conflict resolution.
