Skip to main content

Elasticsearch & The ELK Stack Overview

Elasticsearch Search Architecture: Inverted Index & Text Analysis
The "Back-of-the-Book" Search Analogy

If you need to find the word "distributed" in a 500-page textbook, reading page-by-page from start to finish is a full scan (slow $O(N)$ relational SQL LIKE %distributed%). Instead, you turn to the Index at the back of the book: find the word alphabetically in seconds, and see the exact page numbers listed. That is exactly what an Inverted Index does.

Select Search Term:
Relational DB (Forward Table Scan)
SELECT * FROM docs WHERE text LIKE '%distributed%'
Doc #1: Building resilient distributed database clusters[MATCH]
Doc #2: High throughput database connection pooling
Doc #3: Real-time full text search engine indexing
Doc #4: Scaling distributed search systems across nodes[MATCH]
⚠️ Traverses all 4 rows sequentially. In a 50M row database, this causes disk saturation and multi-second query timeouts.
Elasticsearch (Inverted Index Lookup)
Dictionary Term βž” Postings List (Document IDs)
"cluster"
Doc #1
"database"
Doc #1Doc #2
"distributed"
Doc #1Doc #4
"search"
Doc #3Doc #4
βœ“ Inverted index locates term in $O(1)$ memory time (FST) and instantly returns exact matching document IDs!

1. Why Elasticsearch? The "Back-of-the-Book" Analogy

When building modern web applications, search is one of the most critical user-facing features. However, traditional Relational Database Management Systems (PostgreSQL, MySQL, Oracle) were never designed for full-text document search.

The Problem with Relational Databases

In a relational database, searching for a keyword in a description column requires a pattern match query:

SELECT * FROM products WHERE description LIKE '%wireless headphones%';
  • The Full Table Scan Trap: Because the query begins with a leading wildcard (%), standard B-Tree indexes cannot be used. The database engine is forced to read every single row sequentially off disk (O(N)O(N) table scan).
  • No Relevance Ranking: A SQL query returns a binary boolean: a row either matches or it does not. It cannot tell you which product is the best match.
  • No Typo Tolerance: If a user types "wireles headfones", SQL returns zero results.
  • No Stemming: Searching for "running" will fail to match a record containing "ran" or "runs".

The Inverted Index Analogy

Imagine looking for the word "distributed" in a 500-page computer science textbook:

  1. The SQL Way (Forward Scan): Read page 1 word-by-word, then page 2, page 3... until page 500. This is an exhaustive, linear scan.
  2. The Elasticsearch Way (Inverted Index): Flip directly to the Index at the back of the book. Words are organized alphabetically in seconds. You look up "distributed" and immediately see: Pages 12, 45, 189, 342.

Elasticsearch operates on this exact principle: it inverts the relationship from Document βž” Words to Word (Term) βž” List of Documents (Postings List).

Relational Model (Forward Index):
Doc #1 ──► "Building distributed database clusters"
Doc #2 ──► "High throughput database connection pooling"
Doc #3 ──► "Distributed search systems"

Elasticsearch Model (Inverted Index):
"building" ──► [Doc #1]
"clusters" ──► [Doc #1]
"connection" ──► [Doc #2]
"database" ──► [Doc #1, Doc #2]
"distributed" ──► [Doc #1, Doc #3] <── Instant O(1) Term Lookup!
"pooling" ──► [Doc #2]
"search" ──► [Doc #3]
"throughput" ──► [Doc #2]

2. Core Search Capabilities Beyond Simple Matching

Elasticsearch is much more than a fast lookup table. It provides five foundational capabilities that power modern search bars:

1. BM25 Relevance Scoring

Unlike databases that return unranked rows, Elasticsearch assigns every hit a relevance score (_score) using the Okapi BM25 algorithm (an advanced evolution of TF-IDF):

  • Term Frequency (TF): How often the search term appears in this document (with diminishing returns to prevent keyword stuffing).
  • Inverse Document Frequency (IDF): How rare the term is across the entire index. Uncommon words (e.g., "Kubernetes") carry far more weight than common words (e.g., "the" or "system").
  • Field-Length Normalization: Matches in shorter fields (e.g. a product title) are ranked higher than matches buried in long product descriptions.

Users constantly misspell words on mobile devices. Elasticsearch implements Fuzzy Matching based on the Damerau-Levenshtein Distance:

  • Calculates the minimum number of single-character operations (insertions, deletions, substitutions, or transpositions) needed to turn the search query into a matching dictionary term.
  • Setting fuzziness: "AUTO" allows 1 edit for words of 3–5 letters, and 2 edits for words longer than 5 letters.
{
"query": {
"fuzzy": {
"product_name": {
"value": "samusng",
"fuzziness": "AUTO"
}
}
}
}

3. Autocomplete & Search-As-You-Type

Powers real-time search bar dropdowns as users type:

  • Edge N-Grams: Splits tokens into sequential prefix character combinations (e.g. "laptop" becomes ["la", "lap", "lapt", "lapto", "laptop"]).
  • Completion Suggesters: In-memory Finite State Transducer (FST) data structures that return top matching terms in sub-5ms latency.

4. Real-Time Highlighting

Returns the matching text snippet with search terms wrapped in HTML tags (e.g., <em>distributed</em>), allowing client UIs to visually highlight why the result matched.

5. Aggregations (Real-Time Analytics & Faceting)

Elasticsearch is also an analytics engine. While traditional SQL requires expensive GROUP BY operations, Elasticsearch uses Doc Values (columnar on-disk storage) to compute aggregations concurrently with search:

  • Metric Aggregations: avg, sum, min, max, cardinality (HyperLogLog++ for counting unique users).
  • Bucket Aggregations: Grouping documents into buckets (e.g. e-commerce category facets, price ranges, or date_histogram time-series).

3. Relational Database vs Elasticsearch Architectural Comparison

Architectural DimensionRelational Database (PostgreSQL / MySQL)Elasticsearch
Primary Data ModelStructured relational tables (Rows & Columns)Schemaless or schema-enforced JSON Documents
Primary Index StructureB-Tree / B+Tree (Optimized for ranges & exact keys)Inverted Index + FST (Optimized for text search)
Search MechanismPattern matching (LIKE %term%) β†’\to Full table scanInverted Index lookup β†’\to O(1)O(1) memory seek
Relevance Ranking❌ None (Boolean match only)βœ… BM25 Relevance Scoring
Typo Tolerance❌ None (Requires strict regex or manual Trigrams)βœ… Native Levenshtein Fuzzy Matching
Data ModificationsIn-place row updates with ACID transactionsImmutable Lucene Segments (append-only + merge)
Transaction GuaranteesACID (Immediate consistency, rollback support)Near-Real-Time (NRT) (1s refresh window, no multi-doc ACID)
Horizontal ScalabilityChallenging (Sharding requires complex application logic)Native (Automatic primary & replica shard rebalancing)

4. The Dual-Storage Architecture (RDBMS + Elasticsearch)

In production architectures, Elasticsearch is almost never used as the single primary source of truth for transactional data. It is paired with an ACID database using the Dual-Storage / CQRS Pattern:

The CQRS / Dual-Storage Pattern:
[ Client App ] ──1. Write Transaction (ACID)──► [ PostgreSQL / MySQL ]
β”‚ (Source of Truth)
β”‚ 2. Change Data Capture (CDC)
β–Ό
[ Debezium / Kafka ]
β”‚
β–Ό 3. Bulk Indexing
[ Elasticsearch Cluster ]
β–² (Read-Optimized Search)
β”‚
[ Web Search Bar / Mobile App ] ──4. Search & Filterβ”€β”€β”€β”€β”€β”€β”˜
  1. RDBMS as System of Record: PostgreSQL handles transactional consistency, foreign keys, order payments, and user accounts.
  2. Asynchronous CDC Synchronization: Debezium captures database write-ahead log (WAL) changes and streams them to Kafka, which pushes updates into Elasticsearch in real time.
  3. Elasticsearch as Read Engine: All full-text searches, filter faceted sidebars, and analytical dashboards query Elasticsearch directly, keeping heavy read load off the primary transactional database.

5. The ELK Stack Ecosystem

The Elastic Stack (ELK + Beats) provides an end-to-end data pipeline:

Pipeline StageComponentLanguage / RuntimeCore Role & Responsibilities
1. CollectBeats (Filebeat, Metricbeat)Go (Ultra-lightweight)Host-level daemon harvesting logs and metrics with negligible CPU footprint.
2. ProcessLogstashJava / JRubyCentralized ETL engine running Grok regex, field mutation, GeoIP enrichment, and queue buffering.
3. Store & SearchElasticsearchJava / LuceneDistributed JSON document store indexing text into inverted indexes with sharding and replication.
4. VisualizeKibanaNode.js / ReactWeb portal for analytics dashboards, aggregations, Discover search, and cluster management.

6. Principal Engineering Interview Questions & Answers

Relational databases use B-Tree indexes, which map row IDs to column values. Searching for text containing a keyword (LIKE '%keyword%') cannot use a B-Tree index because of the leading wildcard, forcing an O(N)O(N) sequential scan across every row on disk. In contrast, Elasticsearch uses an Inverted Index: it parses and stems text during ingestion, mapping each unique word to a list of matching document IDs (Postings List). Searching for a keyword is a sub-millisecond O(1)O(1) memory lookup in Lucene's in-memory Term Dictionary (Finite State Transducer), entirely avoiding full-disk scans.

Q2. What is the difference between a match query and a term query in Elasticsearch?

  • term Query: Executes an exact value match without analyzing the search string. It searches the inverted index for the exact bytes provided. Used for structured data like status codes, IDs, booleans, or keyword fields (e.g., status: "ACTIVE").
  • match Query: Executes full-text search. The query string is passed through the same Analyzer (lowercasing, stemming, tokenizing) as the field itself before querying the inverted index. Used for natural language searches across text fields.

Q3. Why shouldn't Elasticsearch be used as the primary database in a transactional application?

Elasticsearch is optimized for search speed and horizontal analytical scaling, not ACID transactional safety:

  1. No Multi-Document ACID Transactions: Elasticsearch does not support atomic cross-document rollback transactions.
  2. Near-Real-Time (NRT) Search Delay: Documents are written to an in-memory buffer and become searchable only after an index refresh (default every 1 second), meaning reads immediately following a write may return stale results.
  3. Risk of Silent Data Corruption: While the translog provides crash durability, Lucene segment merges and cluster rebalances prioritize availability and eventual consistency over strict serializable isolation.

πŸ“–
Track Page Progress0 / 635 Read
Knowledge Base Completion0%