Skip to main content
πŸ“–DATA STORAGE: FUNDAMENTALS TO PRODUCTION

Data Storage.
Documented for Systems Engineers.

Complete reference curriculum and interactive simulators: Linux, Filesystems, RAID, Ceph, Cloud, and Kubernetes. Zero fluff, production-tested.

Course Syllabus: 18 Modules

From Linux kernel I/O mechanics and RAID to Ceph, Cloud, and Kubernetes storage engineering.

MODULE 01

Intro to Data Storage

Data vs information, storage hierarchy, capacity, IOPS, latency, and Linux benchmark tools.

Read Guide→
MODULE 02

How Storage Actually Works

App to disk I/O path, blocks, sectors, 4K alignment, RMW penalty, and page cache profiling.

Read Guide→
MODULE 03

Storage Media (HDD, SSD, NVMe)

Platters and seek times vs NAND flash physics, FTL wear leveling, garbage collection, and TRIM.

Read Guide→
MODULE 04

Interfaces & Protocols

SATA, SAS, Fibre Channel SAN, iSCSI, and hands-on NFS server configuration.

Read Guide→
MODULE 05

Filesystems Deep-Dive

Inodes, directory trees, ext4 vs XFS vs Btrfs vs ZFS, journaling, and mount tuning.

Read Guide→
MODULE 06

Block Storage, LVM & RAID

Linux device mapper, LVM PV/VG/LV, RAID 0/1/5/6/10 write penalty, and disk failure simulation.

Read Guide→
MODULE 07

Object Storage Internals

Block vs File vs Object, S3 architecture, versioning, multipart uploads, and local MinIO lab.

Read Guide→
MODULE 08

Distributed Storage

Replication models, quorum calculus (R+W>N), sharding, and Reed-Solomon Erasure Coding.

Read Guide→
MODULE 09

Ceph in Practice

RADOS, MON, OSD BlueStore, CRUSH map determinism, RBD block devices, and cluster recovery.

Read Guide→
MODULE 10

Databases & Storage

Slotted pages, B-Tree vs LSM, Write Amplification Factor, ARIES WAL crash recovery.

Read Guide→
MODULE 11

Performance & fio

Workload characterization, Little’s Law, queue depth tuning, and enterprise fio benchmark suites.

Read Guide→
MODULE 12

Data Protection & DR

Backup vs replication, CoW vs RoW snapshots, RPO & RTO formulas, and 3-2-1 backup strategies.

Read Guide→
MODULE 13

Cloud Storage (AWS/Azure/GCP)

EBS gp3/io2 vs Managed Disks vs GCS, object lifecycle tiering, and cloud invoice optimization.

Read Guide→
MODULE 14

Kubernetes Storage & CSI

PV, PVC, StorageClasses, Container Storage Interface plugins, and StatefulSet database lab.

Read Guide→
MODULE 15

Storage Security & Compliance

LUKS block encryption, TLS in-transit, POSIX ACLs, IAM, and NIST SP 800-88 crypto-shredding.

Read Guide→
MODULE 16

Architecture & System Design

6-factor storage tradeoffs, capacity planning, and reference patterns for FinTech & streaming.

Read Guide→
MODULE 17

Troubleshooting Runbook

Diagnostic playbooks: Inode exhaustion, latency spikes, degraded RAID, and slow Ceph OSDs.

Read Guide→
MODULE 18

Final Capstone Project

Design a 100 TB multi-region storage platform: 30% growth, 99.99% SLA, hybrid DB + object.

Read Guide→

Mechanical Sympathy in Real Time

Explore the physical latencies separating registers from persistent flash and spinning disks.

⏱️Numbers Every Storage Engineer Should Know

Hardware access times span 8 orders of magnitude. If an L1 cache hit took 1 second, reading from an NVMe SSD is like waiting 5.5 hours, and an HDD seek is like waiting 7.6 months!

L1 Cache Referencecpu
0.5 ns⏳ 1 seconds
Fetching an instruction or scalar from on-die L1 data cache.CPU Core (L1d)
Branch Mispredictcpu
5 ns⏳ 10 seconds
Pipeline flush & speculative execution rewind.Branch Target Buffer
L2 Cache Referencecpu
7 ns⏳ 14 seconds
Fetching from unified L2 core cache (~512KB - 1MB per core).CPU Core (L2)
Mutex Lock / Unlockcpu
25 ns⏳ 50 seconds
Uncontended atomic compare-and-swap (CAS) operation.CPU Cache Coherency (MESI)
Main Memory (DRAM) Accessmemory
100 ns⏳ 3.3 minutes
DDR4/DDR5 memory controller access (CAS latency + transfer).DIMM / Memory Bus
Compress 1KB with Zstandardcpu
2.0 ¡s⏳ 1.1 hours
Lempel-Ziv + FSE fast compression pass on modern CPU.SIMD / CPU registers
NVMe Gen4 SSD 4KB Random Readstorage
10.0 ¡s⏳ 5.6 hours
PCIe 4.0 x4 bus traversal + NAND Flash die cell read.NVMe controller + 3D TLC
SATA SSD Random 4KB Readstorage
150.0 ¡s⏳ 3.5 days
AHCI protocol queue + SATA 6Gbps bus overhead.SATA III Flash SSD
Datacenter Roundtrip (Same DC)network
500.0 ¡s⏳ 11.6 days
Top-of-rack (ToR) switch hop + NIC interrupt handling.100GbE Optical Fabric
Read 1MB Sequentially (NVMe)storage
250.0 ¡s⏳ 5.8 days
Continuous DMA streaming at ~4,000 MB/s across multi-channel NAND.PCIe NVMe Gen4
HDD Mechanical Seek + Rotationalstorage
10.0 ms⏳ 7.6 months
Arm actuator movement + platter 7200 RPM rotational delay.Magnetic Platter & Actuator
Transatlantic Network Ping (NYC to London)network
70.0 ms⏳ 4.4 years
Speed of light in fiber optics (~200,000 km/s) across Atlantic seabed.Submarine Optical Cable

Deep-Dive Architectural Pillars

Specialized technical reference encyclopedia across 7 core domains.

⚑

1. Hardware & Physical Layer

Memory hierarchies, NAND Flash cell physics (SLC/TLC/QLC), FTL wear leveling, NVMe over PCIe, and mechanical disk seeking.

πŸ–₯️

2. Kernel & OS Subsystem

How the Linux kernel abstracts block devices: VFS dentries/inodes, dirty page writeback, io_uring ring buffers, and O_DIRECT.

🌲

3. Storage Engines & Data Structures

The computational primitives organizing data on disk: B+ Tree slotted pages, LSM SSTables, Bloom filters, and segmented logs.

πŸ—„οΈ

4. Database Storage Architectures

Transactional durability and analytical engines: Row vs Columnar (OLTP vs OLAP), WAL & ARIES recovery, and MVCC snapshot isolation.

🌐

5. Distributed Storage & Consensus

Scaling persistence across fault-prone networks: Consistent hashing rings, tunable quorums (R+W>N), Raft state machines, and S3.

πŸ“¦

6. Formats & Compression

Binary encoding efficiency: Apache Parquet Dremel shredding, Apache Arrow in-memory IPC, and Zstandard / LZ4 compression.

🧠

7. Caching & Memory Management

Algorithms and topologies for high-speed caching: ARC, W-TinyLFU, Cache-Aside vs Write-Behind, and cache stampede solutions.

Engine Decision Engine

Evaluate the RUM conjecture trade-offs for your specific workload.

🧭Storage Engine Architecture Decision Matrix
Interactive Tool

Storage engines make fundamental trade-offs governed by the RUM Conjecture (Read, Update, Memory/Space). Select your workload profile to explore the optimal storage structure:

Log-Structured Merge Tree (LSM)

Append-Only Multi-Level Sorted String Tables (SSTables)
Optimal Fit

Writes are written sequentially to a Write-Ahead Log (WAL) and memory table (MemTable), then flushed to disk as immutable SSTables. Compaction merges runs in the background. Exceptional write throughput.

DimensionCharacteristics & Trade-offImpact
Read AmplificationMedium (Bloom-filtered)Read Path
Write AmplificationLow (Sequential Append)NAND Wear / Throughput
Space AmplificationLow (Compacted)Disk Footprint
Notable Production Systems:
RocksDBApache CassandraLevelDBCockroachDB (Pebble)ScyllaDB