Design a Text Sharing Service Like Pastebin
A Pastebin service allows users to store plain text or source code snippets online and generate a unique, compact URL (e.g. https://pastebin.com/a9X4k2) to share with others. While functionally similar to a URL shortener, Pastebin deals with significantly larger payloads (up to 10 MB per paste), syntax highlighting rendering, configurable Time-to-Live (TTL) expiration schedules, and strict abuse/phishing mitigation.
1. Understanding the Problem
Functional Requirements
- Create Paste: Users can paste plain text or code (up to 10 MB), specify an optional expiration date (e.g., 10 minutes, 1 day, 1 month, never), select a programming language for syntax highlighting, and generate a short URL.
- Access Paste: Given a paste ID, retrieve and render the raw text or syntax-highlighted code.
- Custom URL Alias (Optional): Support custom vanity URLs (e.g.,
pastebin.com/my-script). - Burn After Reading (Optional): The paste is permanently deleted from the database immediately after its first read.
- Password Protection (Optional): Client-side or server-side password-encrypted pastes.
Non-Functional Requirements
- High Read-to-Write Ratio: 100:1 read-to-write ratio (100 reads for every 1 paste uploaded).
- Sub-20ms Read Latency: Viewing a paste should resolve in globally.
- High Availability: availability for reading existing pastes.
- Storage Durability: Pastes configured with "Never expire" must persist reliably without data loss.
- Abuse Prevention: Scan pastes in real time to prevent hosting malicious scripts, malware payloads, or phishing credentials.
Capacity Estimations & Sizing (5 Years)
- New Pastes per Day: 10 Million new pastes/day 115 writes/second average (peaking at 1,000 writes/sec).
- Read Requests: 100:1 ratio 1 Billion reads/day 11,500 reads/second average (peaking at 50,000 reads/sec).
- Average Paste Size: 10 KB.
- Storage Projections (5 Years):
- Daily Ingestion: .
- Annual Ingestion: .
- 5-Year Storage: (excluding expired pastes).
- RAM Caching Sizing (80/20 Rule):
- 20% of the pastes generate 80% of total read traffic.
- Daily read volume: 1 Billion reads/day across the Redis cache cluster.
2. The Set Up
Defining Core Entities
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β PASTE_METADATA β
ββββββββββββββββββββ¬βββββββββββββββ¬βββββββββββββββββββββββ€
β paste_id β VARCHAR(10) β PRIMARY KEY (Base62) β
β title β VARCHAR(128) β Optional Title β
β user_id β UUID β NULLABLE (Anonymous) β
β language β VARCHAR(32) β Syntax (e.g. "java") β
β size_bytes β INT β Payload byte length β
β s3_object_key β VARCHAR(256) β Storage Path / Blob β
β is_burn_after_rd β BOOLEAN β Self-destruct flag β
β password_hash β VARCHAR(128) β Argon2 hash β
β created_at β TIMESTAMP β Creation Time β
β expires_at β TIMESTAMP β NULL = Never Expire β
ββββββββββββββββββββ΄βββββββββββββββ΄βββββββββββββββββββββββ
The API Design
1. Create Pasteβ
POST /api/v1/pastes
Content-Type: application/json
Idempotency-Key: 9481a-8291-419b-a012
{
"content": "public class HelloWorld { public static void main(String[] args) {} }",
"title": "Java Hello",
"language": "java",
"ttl_seconds": 86400, // 24 hours (null = never)
"burn_after_read": false,
"password": "secret_password" // Optional
}
Response (201 Created):
{
"paste_id": "a9X4k2",
"url": "https://pastebin.com/a9X4k2",
"size_bytes": 68,
"expires_at": "2026-10-02T12:00:00Z"
}
2. Get Pasteβ
GET /api/v1/pastes/{paste_id}
X-Paste-Password: secret_password // Optional
Response (200 OK):
{
"paste_id": "a9X4k2",
"title": "Java Hello",
"language": "java",
"content": "public class HelloWorld { public static void main(String[] args) {} }",
"created_at": "2026-10-01T12:00:00Z",
"expires_at": "2026-10-02T12:00:00Z"
}
3. High-Level Design
Walkthrough of Core Flows
1. Create Paste Write Flowβ
- Client submits text payload to the API Gateway.
- Gateway verifies rate limits (e.g. max 20 pastes/hour per IP).
- The Paste Service queries an asynchronous Malware & Phishing Scanner (ClamAV / ML threat model) to verify the text does not contain malicious code or stolen credentials.
- The service fetches a unique 64-bit ID from a Key Generation Service (KGS) and converts it to a 7-character Base62 string (
paste_id). - Storage Separation:
- The large text payload is compressed via Zstandard (Zstd) and written directly to Distributed Object Storage (Amazon S3 / MinIO).
- The metadata (
paste_id,s3_object_key,expires_at,size_bytes) is inserted into PostgreSQL / DynamoDB.
- The paste metadata and text are primed into the Redis Cluster.
- Returns the shortened paste URL to the user.
2. Get Paste Read Flowβ
- User requests
GET /pastes/{paste_id}. - The request hits Cloudflare Edge CDN:
- CDN Cache Hit (60%+): Returns the cached paste text in .
- On CDN miss, the request routes to the Paste Service:
- Queries Redis: If present, returns in .
- If Redis miss, queries the Metadata DB to verify the paste has not expired.
- Streams raw text from Amazon S3, primes Redis, and returns to user.
- Burn After Reading Check: If
is_burn_after_read == true, the service immediately enqueues an asynchronous deletion task to purge the metadata row, Redis key, and S3 object!
4. Potential Deep Dives & Bottlenecks
Deep Dive 1: Storage Tiering: Database BLOB vs Object Storage
Should paste text be stored inside database rows or external object storage?
| Storage Alternative | Database BLOB (Postgres / MySQL) | Distributed Object Storage (S3 / GCS) |
|---|---|---|
| Cost per GB | High (0.25 / GB / month on SSD) | Ultra-low (0.023 / GB / month) |
| Max Payload Size | Large payloads bloat DB buffer pools and trigger row-chaining. | Easily handles multi-megabyte payloads up to 10 MB without degradation. |
| Throughput & IOPS | High write amplification on DB indexes and WAL logs. | Unlimited horizontal scaling and zero database IOPS consumption. |
| Decision | Store only lightweight metadata (). | Store all raw paste text payloads as compressed blobs. |
Deep Dive 2: Expiration & TTL Deletion Mechanics
How do we delete expired pastes without running slow database full-table scans?
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β DUAL-TIER EXPIRATION & DELETION ENGINE β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β β
β Tier 1: Lazy Deletion on Read Path β
β β’ When a user reads a paste: β
β IF (expires_at < current_timestamp): β
β Return HTTP 404 Not Found! β
β Trigger async deletion event. β
β β
β Tier 2: Asynchronous Janitor Sweep β
β β’ Database B+Tree Index on `expires_at`: β
β `CREATE INDEX idx_pastes_expiry ON pastes(expires_at)β
β β’ Background cron runs every 10 minutes: β
β SELECT paste_id, s3_key FROM pastes β
β WHERE expires_at < NOW() LIMIT 5000; β
β β’ Deletes S3 blobs in batch and drops metadata rows. β
β β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
- S3 Object Lifecycle Policies: Alternatively, we set S3 lifecycle expiration rules based on bucket prefixes (e.g.,
s3://pastes/ttl-24h/,s3://pastes/ttl-7d/), allowing AWS to delete underlying storage automatically at zero compute cost!
Deep Dive 3: Unique Short ID Generation (Range-Based KGS)
To ensure zero collisions and sub-millisecond generation:
- An Apache ZooKeeper cluster maintains global integer counters.
- Each paste server requests a block of 1,000,000 IDs at startup.
- The server increments IDs in local memory using
AtomicLongand converts to Base62: - If a server crashes, the remaining unused IDs in its local memory block are simply discarded (with combinations, gaps are completely negligible).
Deep Dive 4: Syntax Highlighting Caching
Computing syntax highlighting (converting raw code to thousands of colored HTML <span> tags via PrismJS or Pygments) is CPU-intensive ( per 100 KB file).
- The Optimization: Compute syntax highlighting once at write time or on the first read.
- Cache the highlighted HTML directly in the CDN and Redis alongside the raw text. Subsequent reads serve the pre-rendered HTML without burning backend CPU cycles.
5. Architectural Trade-Off Matrix
| Design Alternative | Option A | Option B | Selected Choice & Rationale |
|---|---|---|---|
| Text Storage | Relational Database TEXT column | Amazon S3 / MinIO Object Storage | Object Storage: Slashes database storage cost by 90%; protects database buffer pool from memory churn. |
| Expiration Method | Real-time database poll every second | Lazy Read Check + Binned S3 Lifecycle | Lazy + Binned Lifecycle: Eliminates database CPU spikes from constant background polling. |
| ID Generation | Truncated MD5 Hash | Range-allocated ZooKeeper KGS | Range KGS: Guarantees zero hash collisions without expensive uniqueness retry queries. |
| Syntax Highlighting | Dynamic Client-side JavaScript | Pre-rendered & Edge Cached HTML | Pre-rendered & Edge Cached: Eliminates client-side layout shifts and CPU rendering freezes on low-end mobile phones. |
6. What is Expected at Each Level?
Mid-Level (L4 / IC4)
- Identifies the similarity and differences between Pastebin and a URL shortener (larger text payloads vs simple redirects).
- Designs a clean database schema and calculates storage sizing for text files.
- Implements basic expiration logic and Base62 encoding.
Senior (L5 / IC5)
- Separates metadata storage (database) from text payload storage (object storage / S3).
- Implements the dual-tier expiration engine (lazy read checks + asynchronous janitor sweeps).
- Designs read caching using CDN edge nodes and Redis (80/20 rule).
- Handles syntax highlighting caching and abuse scanning integration.
Staff+ (L6 / Principal)
- Evaluates client-side zero-knowledge encryption architecture (encrypting text in browser with AES-GCM before transmission so server cannot read contents).
- Formulates multi-region replication strategies for globally distributed text snippet retrieval.
- Designs defense-in-depth protection against malware payload hosting and crawler scraping.
- Formulates high-throughput "Burn After Reading" race-condition prevention using atomic distributed locks.
